Gradient-Based Prediction Improvement for Video Coding
By rounding displacements to a unified accuracy level, the solution addresses the complexity and power consumption issues in video coding, enabling efficient gradient-based prediction refinement across various inter prediction modes with reduced logic circuitry.
Patent Information
- Application Number
- JP2021566995
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-14
- Filing Date
- 2020-05-15
- Publication Date
- 2025-06-12
- Estimated Expiration
- 2040-05-15
AI Technical Summary
Existing video coding technologies face challenges in efficiently managing different accuracy levels for horizontal and vertical displacements across various inter prediction modes, leading to increased complexity and power consumption due to the need for multiple logic circuits.
The proposed solution involves rounding the horizontal and vertical displacements to a unified accuracy level for different inter prediction modes, allowing the same logic circuit to perform gradient-based prediction refinement, thereby reducing the number of required logic circuits and power consumption.
This approach improves the overall operation of video coders by simplifying the logic circuitry and reducing power consumption, while maintaining effective gradient-based prediction refinement across different inter prediction modes.
Smart Images

Figure 0007691936000057 
Figure 0007691936000058 
Figure 0007691936000059
Abstract
Description
Technical Field
[0001]
[0001] This application claims the benefit of U.S. Provisional Application No. 62 / 849,352, filed May 17, 2019, and claims priority to U.S. Application No. 16 / 874,057, filed May 14, 2020, the entire contents of each of which are hereby incorporated by reference.
[0002]
[0002] This disclosure relates to video encoding and video decoding.
Background Art
[0003]
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radiotelephones, so-called "smartphones," video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions to such standards. By implementing such video coding techniques, video devices can transmit, receive, encode, decode, and / or store digital video information more efficiently.
[0004]
[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in a video sequence. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, sometimes referred to as coding tree units (CTUs), coding units (CUs), and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in adjacent blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in adjacent blocks in the same picture, or temporal prediction with respect to reference samples in other reference pictures. A picture may sometimes be referred to as a frame, and a reference picture may sometimes be referred to as a reference frame.
Summary of the Invention
[0005]
[0005] Generally, in this disclosure, techniques for gradient-based prediction refinement are described. A video coder (e.g., a video encoder or a video decoder) determines one or more prediction blocks for inter-predicting a current block (e.g., based on one or more motion vectors for the current block). In gradient-based prediction refinement, the video coder modifies one or more samples of the prediction block based on various factors such as displacement in the horizontal direction, horizontal gradient, displacement in the vertical direction, and vertical gradient.
[0006]
[0006] For example, a motion vector identifies a prediction block. The displacement in the horizontal direction (also referred to as horizontal displacement) refers to the change (e.g., delta) of the x - coordinate of the motion vector, and the displacement in the vertical direction (also referred to as vertical displacement) refers to the change (e.g., delta) of the y - coordinate. The horizontal gradient refers to the result of applying a filter to the first set of samples in the prediction block, and the vertical gradient refers to the result of applying a filter to the second set of samples in the prediction block.
[0007]
[0007] The exemplary techniques described in the present disclosure provide gradient - based prediction improvement that is unified (e.g., the same) for different prediction modes with different precision levels of displacement (e.g., at least one of horizontal displacement or vertical displacement). For example, for a first prediction mode (e.g., affine mode), the motion vector can be at a first precision level, and for a second prediction mode (e.g., bi - directional optical flow (BDOF)), the motion vector can be at a second precision level. Thus, the vertical and horizontal displacements for the motion vector used for the affine mode and the motion vector used for BDOF can be different. In the present disclosure, the video coder can be configured to round (e.g., round up or round down) the vertical and horizontal displacements for the motion vector so that the precision levels of the displacements are the same regardless of the prediction mode (e.g., the vertical and horizontal displacements for affine mode and BDOF have the same precision level).
[0008]
[0008] By rounding the displacement accuracy level, exemplary techniques can improve the overall operation of a video coder. For example, gradient-based prediction improvement involves multiplication and shift operations. If the displacement accuracy level varies for each mode, different logic circuits may be required to support different accuracy levels (e.g., a logic circuit configured for one accuracy level may not be suitable for other accuracy levels). Since the accuracy level for displacement is the same for different modes, the same logic circuit can be reused for blocks, resulting in a smaller overall logic circuit and reduced power consumption as there is no need to power unused logic circuits.
[0009]
[0009] In some examples, the technique for determining displacement can be based on information already available in a video decoder. For example, the way a video decoder determines a horizontal or vertical displacement can be based on information available to the video decoder for inter-predicting the current block according to an inter-prediction mode. Further, there may be some inter-prediction modes that are prohibited for use for some block types (e.g., based on size). In some examples, these inter-prediction modes that are prohibited for use for some block types can be made available for these block types, but the prediction blocks for such blocks can be modified using the exemplary techniques described in this disclosure.
[0010]
[0010] In one example, the present disclosure describes a method for decoding video data. The method includes determining a prediction block for inter-predicting a current block, determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block, rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a BDOF mode, determining one or more refinement offsets based on the rounded horizontal and vertical displacements, modifying one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block, and reconstructing the current block based on the modified prediction block.
[0011]
[0011] In one example, the present disclosure describes a method for encoding video data. The method includes determining a prediction block for inter-predicting a current block, determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block, rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a BDOF mode, determining one or more refinement offsets based on the rounded horizontal and vertical displacements, modifying one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block, determining a residual value indicating a difference between the current block and the modified prediction block, and signaling information indicating the residual value.
[0012]
[0012] In one example, the present disclosure describes a device for coding video data, the device comprising a memory configured to store one or more samples of a prediction block and processing circuitry. The processing circuitry is configured to determine a prediction block for inter-predicting a current block, determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block, round the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a BDOF mode, determine one or more refined offsets based on the rounded horizontal displacement and the vertical displacement, modify one or more samples of the prediction block based on the determined one or more refined offsets to generate a modified prediction block, and code the current block based on the modified prediction block.
[0013]
[0013] In one example, the present disclosure describes a computer-readable storage medium storing instructions that, when executed, cause one or more processors to determine a prediction block for inter-predicting a current block, determine a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block, round the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a BDOF mode, determine one or more refined offsets based on the rounded horizontal displacement and the vertical displacement, modify one or more samples of the prediction block based on the determined one or more refined offsets to generate a modified prediction block, and code the current block based on the modified prediction block.
[0014]
[0014] In one example, the present disclosure describes a device for coding video data, the device comprising means for determining a prediction block for inter-predicting a current block, means for determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block, means for rounding the horizontal displacement and the vertical displacement to the same level of accuracy for different inter-prediction modes including an affine mode and a BDOF mode, means for determining one or more refinement offsets based on the rounded horizontal displacement and the vertical displacement, means for modifying one or more samples of the prediction block based on the determined one or more refinement offsets to generate a modified prediction block, and means for coding the current block based on the modified prediction block.
[0015]
[0015] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will become apparent from the description, drawings, and claims.
Brief Description of the Drawings
[0016]
Figure 1
[0016] Block diagram showing an exemplary video encoding and decoding system that can execute the techniques of the present disclosure.
Figure 2A
[0017] Conceptual diagram showing an exemplary quadtree binary tree (QTBT) structure.
Figure 2B
Figure 3
[0018] Block diagram showing an exemplary video encoder that can execute the techniques of the present disclosure.
Figure 4
[0019] Block diagram showing an exemplary video decoder that can execute the techniques of the present disclosure.
Figure 5
[0020] Conceptual diagram showing an extended coding unit (CU) region used in bidirectional optical flow (BDOF).
Figure 6
[0021] Conceptual diagram showing an example of sub-block motion vector (MV) selection.
Figure 7
[0022] Flowchart showing an exemplary method of coding video data. [Embodiments for Carrying Out the Invention]
[0017]
[0023] The present disclosure relates to gradient-based prediction improvement. In gradient-based prediction improvement, a video coder (e.g., a video encoder or a video decoder) determines a prediction block for a current block based on motion vectors as part of inter prediction, and modifies (e.g., improves) samples of the prediction block to generate modified prediction samples (e.g., improved prediction samples). The video encoder signals a residual value indicating the difference between the modified prediction samples and the current block. The video decoder performs the same operations that the video encoder performed to modify the samples of the prediction block to generate the modified prediction samples. The video decoder adds the residual value to the modified prediction samples to reconstruct the current block.
[0018]
[0024] One exemplary method of modifying samples of a prediction block is for the video coder to determine one or more improvement offsets and add the samples of the prediction block to the improvement offsets. One exemplary method of generating the improvement offsets is based on a gradient and a motion vector displacement. The gradient can be determined from a gradient filter applied to the samples of the prediction block.
[0019]
[0025] Examples of motion vector displacements include a horizontal displacement with respect to the motion vector and a vertical displacement with respect to the motion vector. The horizontal displacement can be a value added to or subtracted from the x coordinate of the motion vector, and the vertical displacement can be a value added to or subtracted from the y coordinate of the motion vector. For example, the horizontal displacement may be called Δv x where v x is the x coordinate of the motion vector, and the vertical displacement may be called Δv y where v y is the y coordinate of the motion vector.
[0020]
[0026] The accuracy level of the motion vector of the current block can vary for each inter prediction mode. For example, the coordinates of the motion vector (e.g., the x coordinate or the y coordinate) can include an integer part and may include a fractional part. Since the integer part of the motion vector identifies the actual pixel in the reference picture containing the predicted block, the fractional part is called the sub-pel portion of the motion vector, and the sub-pel portion of the motion vector adjusts the motion vector to identify a location between pixels in the reference picture.
[0021]
[0027] The accuracy level of the motion vector is based on the sub-pel portion of the motion vector and indicates the granularity of the movement of the motion vector from the actual pixel in the reference picture. As an example, when the sub-pel portion of the x coordinate is 0.5, the motion vector is in the middle between two horizontal pixels in the reference picture. When the sub-pel portion of the x coordinate is 0.25, the motion vector is at a quarter between two horizontal pixels, and so on. In these examples, the accuracy level of the motion vector can be equal to the sub-pel portion (e.g., the accuracy level is 0.5, 0.25, etc.).
[0022]
[0028] In some examples, the accuracy levels of the horizontal displacement and the vertical displacement may be based on the accuracy level of the motion vector or the way the motion vector is generated. For example, in some examples such as the merge mode which is a form of the inter prediction mode, the sub-pel parts of the x coordinate and the y coordinate of the motion vector may be the horizontal displacement and the vertical displacement respectively. As another example, in the case of the affine mode which is a form of the inter prediction, the motion vector may be based on the motion vectors of the endpoints, and the horizontal displacement and the vertical displacement may be determined based on the motion vectors of the endpoints.
[0023]
[0029] The accuracy levels of the horizontal displacement and the vertical displacement may vary for each inter prediction mode. For example, in the case of some inter prediction modes, the horizontal displacement and the vertical displacement may be more accurate (e.g., the accuracy level is 1 / 128 for the first prediction mode) compared to other inter prediction modes (e.g., the accuracy level is 1 / 16 for the second prediction mode).
[0024]
[0030] In an implementation, the video coder may need to include different logic circuits to handle different accuracy levels. Performing gradient-based prediction refinement involves multiplication, shift operations, addition, and other arithmetic operations. The logic circuit configured for one accuracy level for the horizontal or vertical displacement may not be able to process the horizontal and vertical displacements with a higher accuracy level. Thus, some video coders include one set of logic circuits for performing gradient-based prediction refinement for one inter prediction mode where the horizontal and vertical displacements have a first accuracy level and a different set of logic circuits for performing gradient-based prediction refinement for another inter prediction mode where the horizontal and vertical displacements have a second accuracy level.
[0025]
[0031] However, having different logic circuits for performing gradient-based prediction improvements for different inter prediction modes results in additional logic circuits that consume additional power in addition to increasing the size of the video coder. For example, if the current block is inter predicted in the first mode, a first set of logic circuits for gradient-based prediction improvement is used. However, a second set of logic circuits for gradient-based prediction improvement for different inter prediction modes is still receiving power.
[0026]
[0032] In the present disclosure, examples of techniques for rounding the accuracy levels for horizontal displacement and vertical displacement to the same accuracy level for different inter prediction modes will be described. For example, a video coder may round a first displacement (e.g., a first horizontal displacement or a first vertical displacement) having a first accuracy level for a first block inter predicted in a first inter prediction mode to a set accuracy level, and may round a second displacement (e.g., a second horizontal displacement or a second vertical displacement) having a second accuracy level for a second block inter predicted in a second inter prediction mode to the same set accuracy level. In other words, a video coder may round at least one of a horizontal displacement and a vertical displacement to an accuracy level that is the same for different inter prediction modes. As an example, the first inter prediction mode may be an affine mode, and the second inter prediction mode may be a bi-directional optical flow (BDOF).
[0027]
[0033] In this way, instead of having different logic circuits for different inter prediction modes, the same logic circuit may be used for gradient-based prediction improvement for different inter prediction modes. For example, the logic circuit of a video coder may be configured to perform gradient-based prediction improvement for horizontal and vertical displacements having a set accuracy level. The video coder may round the horizontal and vertical displacements such that the accuracy levels of the rounded horizontal and vertical displacements are equal to the set accuracy level, enabling the same logic circuit to perform gradient-based prediction improvement for different inter prediction modes.
[0028]
[0034] FIG. 1 is a block diagram showing an exemplary video encoding and decoding system 100 that can execute the techniques of the present disclosure. The techniques of the present disclosure generally target encoding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata such as signaling data.
[0029]
[0035] As shown in FIG. 1, system 100 includes, in this example, a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. In particular, source device 102 provides video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, set-top boxes, etc. In some cases, source device 102 and destination device 116 can be equipped for wireless communication and can thus be referred to as wireless communication devices.
[0030]
[0036] In the example of FIG. 1, the source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. The destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to the present disclosure, the video encoder 200 of the source device 102 and the video decoder 300 of the destination device 116 may be configured to apply techniques for improving gradient-based prediction. Thus, the source device 102 represents an example of a video encoding device, and the destination device 116 represents an example of a video decoding device. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device 102 may receive video data from an external video source such as an external camera. Similarly, the destination device 116 may interface with an external display device rather than including an integrated display device.
[0031]
[0037] The system 100 shown in FIG. 1 is merely an example. In general, any digital video encoding and / or decoding device may perform techniques for improving gradient-based prediction. The source device 102 and the destination device 116 are merely examples of coding devices that generate coded video data for transmission from the source device 102 to the destination device 116. In the present disclosure, a "coding" device is referred to as a device that performs coding (encoding and / or decoding) of data. Thus, the video encoder 200 and the video decoder 300 represent examples of coding devices, particularly a video encoder and a video decoder, respectively. In some examples, the devices 102, 116 may operate substantially symmetrically such that each of the devices 102, 116 includes a video encoding component and a video decoding component. Thus, the system 100 may support one-way or two-way video transmission between the video device 102 and the video device 116, for example, for video streaming, video playback, video broadcast, or video telephony.
[0032]
[0038] Generally, video source 104 represents a source of video data (i.e., raw, unencoded video data), provides a continuous series of pictures (also referred to as "frames") of video data to video encoder 200, and video encoder 200 encodes the data for the pictures. The video source 104 of source device 102 can include a video capture device such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, video source 104 can generate computer graphics-based data, or a combination of live video, archived video, and computer-generated video, as the source video. In each case, video encoder 200 encodes the captured video data, pre-captured video data, or computer-generated video data. Video encoder 200 can reorder the pictures from the received order (sometimes referred to as the "display order") to a coding order for coding. Video encoder 200 can generate a bitstream including the encoded video data. Source device 102 can then output the encoded video data to computer-readable medium 110 via output interface 108 for reception and / or retrieval, for example, by input interface 122 of destination device 116.
[0033]
[0039] Memory 106 of source device 102 and memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106, 120 may store raw video data, such as raw video from video source 104 and raw, decoded video data from video decoder 300. Additionally or alternatively, memories 106, 120 may store software instructions executable by video encoder 200 and video decoder 300, respectively. Although video encoder 200 and video decoder 300 are shown separately in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally equivalent or equivalent purposes. Further, memories 106, 120 may store encoded video data, such as the output from video encoder 200 and the input to video decoder 300. In some examples, portions of memories 106, 120 may be allocated as one or more video buffers to store, for example, raw decoded and / or encoded video data.
[0034]
[0040] Computer-readable medium 110 may represent any type of medium or device capable of transferring encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. Output interface 108 may modulate a transmission signal including the encoded video data, and input interface 122 may demodulate the received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other device that may be useful for facilitating communication from source device 102 to destination device 116.
[0035]
[0041] In some examples, source device 102 may output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as a hard drive, a Blu-ray Disc®, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data.
[0036]
[0042] In some examples, the source device 102 may output the encoded video data to a file server 114 or another intermediate storage device that may store the encoded video generated by the source device 102. The destination device 116 may access the video data stored from the file server 114 via streaming or downloading. The file server 114 may be any type of server device that can store the encoded video data and transmit the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a content delivery network device, or a Network Attached Storage (NAS) device. The destination device 116 may access the encoded video data from the file server 114 through any standard data connection including an Internet connection. This may include a wireless channel (e.g., Wi-Fi® connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored in the file server 114. The file server 114 and the input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.
[0037]
[0043] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet® card), a wireless communication component operating according to any of various IEEE 802.11 standards, or other physical components. In examples where the output interface 108 and the input interface 122 include wireless components, the output interface 108 and the input interface 122 may be configured to transfer data such as video data encoded according to cellular communication standards such as 4G, 4G-LTE® (Long Term Evolution), LTE-Advanced, 5G, etc. In some examples where the output interface 108 includes a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transfer data such as video data encoded according to other wireless standards such as the IEEE 802.11 specifications, the IEEE 802.15 specifications (e.g., ZigBee®), the Bluetooth® standards, etc. In some examples, the source device 102 and / or the destination device 116 may include respective system-on-chip (SoC) devices. For example, the source device 102 may include an SoC device for performing functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include an SoC device for performing functions attributed to the video decoder 300 and / or the input interface 122.
[0038]
[0044] The techniques of the present disclosure may be applied to video coding that supports any of various multimedia applications such as over-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.
[0039]
[0045] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The computer-readable medium 110 of the encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements having values that describe the characteristics and / or processing of video blocks or other coded units (e.g., slices, pictures, groups of pictures, sequences, etc.), which are also used by the video decoder 300. The display device 118 displays the decoded pictures of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0040]
[0046] Although not shown in FIG. 1, in some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or an audio decoder and may include an appropriate MUX-DEMUX unit, or other hardware and / or software, to handle a multiplexed stream that includes both audio and video in a common data stream. When applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol, or other protocols such as the User Datagram Protocol (UDP).
[0041]
[0047] The video encoder 200 and the video decoder 300 can each be implemented as any of a variety of suitable encoder circuits and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques herein are implemented in part in software, the device can store software instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Each of the video encoder 200 and the video decoder 300 can be included in one or more encoders or decoders, and any of them can be integrated as part of a combined encoder / decoder (CODEC) in their respective devices. Devices including the video encoder 200 and / or the video decoder 300 can include integrated circuits, microprocessors, and / or wireless communication devices such as cellular telephones.
[0042]
[0048] The video encoder 200 and the video decoder 300 may operate according to a video coding standard such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or an extension thereof such as multi-view and / or scalable video coding extensions. Alternatively, the video encoder 200 and the video decoder 300 may operate according to other proprietary or industry standards, such as ITU-T H.266, also known as Versatile Video Coding (VVC). The most recent draft of the VVC standard is described in Bross et al., "Versatile Video Coding (Draft 4)", Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 13th meeting: Marrakech, MA, January 9-18, 2019, JVET-M1001-v5 (hereinafter, "VVC Draft 4"). The more recent draft of the VVC standard is described in Bross et al., "Versatile Video Coding (Draft 8)", Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 17th meeting: Brussels, BE, January 7-17, 2020, JVET-Q2001-vD (hereinafter, "VVC Draft 8"). However, the techniques of the present disclosure are not limited to any particular coding standard.
[0043]
[0049] Generally, video encoder 200 and video decoder 300 may perform block - based coding of pictures. The term "block" generally refers to a structure that contains data to be processed (e.g., to be encoded, decoded, or used in other ways in the encoding and / or decoding process). For example, a block may include a two - dimensional matrix of samples of luminance and / or chrominance data. Generally, video encoder 200 and video decoder 300 may code video data represented in the YUV (e.g., Y, Cb, Cr) format. That is, instead of coding red, green, and blue (RGB) data for the samples of a picture, video encoder 200 and video decoder 300 may code a luminance component and a chrominance component, where the chrominance component may include both a red - phase and a blue - phase chrominance component. In some examples, video encoder 200 converts the received RGB - format data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre - processing and post - processing units (not shown) may perform these conversions.
[0044]
[0050] This disclosure may generally refer to the coding (e.g., encoding and decoding) of pictures, including the process of encoding or decoding picture data. Similarly, this disclosure may refer to the coding of picture blocks, including the process of encoding or decoding data for a block, e.g., prediction and / or residual coding processes. An encoded video bitstream generally includes a series of values of syntax elements that represent coding decisions (e.g., coding modes) and the partitioning of a picture into blocks. Thus, a reference to coding a picture or a block should generally be understood as coding the values of the syntax elements that form the picture or block.
[0045]
[0051] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) divides a coding tree unit (CTU) into CUs according to a quadtree structure. That is, the video coder divides the CTU and CU into four equal and non-overlapping squares, and each node of the quadtree has either zero or four child nodes. A node without child nodes may be called a "leaf node", and the CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further divide the PU and TU. For example, in HEVC, a residual quadtree (RQT) represents the division of the TU. In HEVC, the PU represents inter-prediction data, while the TU represents residual values. An intra-predicted CU includes intra-prediction information such as an intra-mode indication.
[0046]
[0052] As another example, video encoder 200 and video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as video encoder 200) divides a picture into a plurality of coding tree units (CTUs). Video encoder 200 may divide the CTU according to a tree structure, such as a quadtree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple division types, such as the separation between the CU, PU, and TU of HEVC. The QTBT structure includes two levels: a first level divided according to a quadtree division and a second level divided according to a binary tree division. The root node of the QTBT structure corresponds to the CTU. The leaf node of the binary tree corresponds to the coding unit (CU).
[0047]
[0053] In an MTT partitioning structure, a block can be partitioned using a quadtree (QT) partition, a binary tree (BT) partition, and one or more types of ternary tree (TT) partitions. A ternary tree partition is a partition in which a block is divided into three sub-blocks. In some examples, the ternary tree partition divides the block into three sub-blocks without dividing the original block through its center. The partitioning types (e.g., QT, BT, and TT) in MTT can be symmetric or asymmetric.
[0048]
[0054] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components. In other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for each chrominance component).
[0049]
[0055] Video encoder 200 and video decoder 300 can be configured to use a quadtree partition, a QTBT partition, an MTT partition, or other partitioning structures according to HEVC. For the purpose of explanation, the description of the techniques of the present disclosure is presented with respect to the QTBT partition. However, it should be understood that the techniques of the present disclosure can also be applied to video coders configured to use a quadtree partition or similarly other types of partitions.
[0050]
[0056] The present disclosure may use "N×N" and "N by N" to interchangeably refer to the sample dimensions of blocks (such as CUs or other video blocks) with respect to vertical and horizontal dimensions, e.g., 16×16 samples or 16 by 16 samples. Generally, a 16×16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non - negative integer value. The samples in a CU can be arranged in rows and columns. Further, a CU does not necessarily have to have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU can comprise N×M samples, where M does not necessarily equal N.
[0051]
[0057] Video encoder 200 encodes video data for a CU representing prediction and / or residual information and other information. Prediction information indicates how a CU should be predicted to form a prediction block for the CU. Residual information generally represents the sample - by - sample differences between the samples of a CU before encoding and the prediction block.
[0052]
[0058] To predict a CU, the video encoder 200 can generally form a prediction block for the CU through inter prediction or intra prediction. Inter prediction generally refers to predicting a CU from the data of a previously coded picture, while intra prediction generally refers to predicting a CU from the previously coded data of the same picture. To perform inter prediction, the video encoder 200 can use one or more motion vectors to generate a prediction block. The video encoder 200 can generally perform a motion search to identify a reference block that exactly matches the CU, for example, with respect to the difference between the CU and the reference block. The video encoder 200 can calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block exactly matches the current CU. In some examples, the video encoder 200 can predict the current CU using uni-directional prediction or bi-directional prediction.
[0053]
[0059] Some examples of VVC also provide an affine motion compensation mode that can be regarded as an inter prediction mode. In the affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion, such as zoom in or out, rotation, perspective motion, or other irregular motion types.
[0054]
[0060] To perform intra prediction, the video encoder 200 may select an intra prediction mode to generate a prediction block. Some examples of VVC provide 67 intra prediction modes including various direction modes as well as planar mode and DC mode. Generally, the video encoder 200 selects an intra prediction mode that describes adjacent samples for the current block (e.g., a block of a CU) from which the samples of the current block are to be predicted. Such samples are generally above, above and to the left, or to the left of the current block in the same picture as the current block, assuming that the video encoder 200 codes CTUs and CUs in raster scan order (from left to right, from top to bottom).
[0055]
[0061] The video encoder 200 encodes data representing the prediction mode for the current block. For example, in the inter prediction mode, the video encoder 200 may encode which of the various available inter prediction modes is used, as well as data representing the motion information of the corresponding mode. For example, in uni - directional or bi - directional inter prediction, the video encoder 200 may use advanced motion vector prediction (AMVP) or merge mode to encode the motion vector. The video encoder 200 may use a similar mode to encode the motion vector of the affine motion compensation mode.
[0056]
[0062] Following prediction such as intra prediction or inter prediction of a block, video encoder 200 may calculate a residual value for the block. A residual value, such as a residual block, represents the per-sample difference between a block formed using a corresponding prediction mode and a prediction block for the block. Video encoder 200 may apply one or more transforms to the residual block to generate transformed data in a transform domain rather than a sample domain. For example, video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Further, video encoder 200 may apply a second transform following a first transform, such as a mode-dependent non-separable second-order transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT). Video encoder 200 generates transform coefficients following the application of one or more transforms.
[0057]
[0063] As described above, following any transform for generating transform coefficients, video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to the process by which transform coefficients are quantized to reduce, as much as possible, the amount of data used to represent the coefficients, resulting in further compression. By performing the quantization process, video encoder 200 may reduce the bit depth associated with some or all of the coefficients. For example, video encoder 200 may round an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, video encoder 200 may perform a bitwise right shift of the value to be quantized.
[0058]
[0064] Following quantization, video encoder 200 may scan the transform coefficients to generate a one-dimensional vector from the two-dimensional matrix containing the quantized transform coefficients. The scan may be designed to place coefficients with higher energy (and thus lower frequency) at the front of the vector and transform coefficients with lower energy (and thus higher frequency) at the back of the vector. In some examples, video encoder 200 may utilize a predefined scan order to scan the quantized transform coefficients to generate a serialized vector and then entropy encode the quantized transform coefficients of the vector. In other examples, video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, video encoder 200 may entropy encode the one-dimensional vector, for example, according to context adaptive binary arithmetic coding (CABAC). Video encoder 200 may also entropy encode values for syntax elements that describe metadata associated with the encoded video data for use by video decoder 300 when decoding the video data.
[0059]
[0065] To perform CABAC, video encoder 200 may assign a context within a context model to the symbol to be transmitted. The context may be related to, for example, whether the neighboring values of the symbol are zero values. The probability determination may be based on the context assigned to the symbol.
[0060]
[0066] The video encoder 200 may further generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, in other syntax data, such as a picture header, a block header, a slice header, or a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS), for example, so that the video decoder 300 can generate it. The video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data.
[0061]
[0067] In this way, the video encoder 200 may generate an encoded video data bitstream, for example, including picture partitioning into blocks (e.g., CUs) and syntax elements describing block prediction and / or residual information. Finally, the video decoder 300 may receive the bitstream and decode the encoded video data.
[0062]
[0068] Generally, the video decoder 300 performs a process opposite to that performed by the video encoder 200 to decode the encoded video data in the bitstream. For example, the video decoder 300 may use CABAC in a manner substantially similar to, but opposite to, the CABAC encoding process of the video encoder 200 to decode the values of the syntax elements in the bitstream. The syntax elements may define the picture partitioning information for the CTU into CUs and the partitioning of each CTU according to a corresponding partitioning structure, such as a QTBT structure. The syntax elements may further define prediction and residual information for blocks (e.g., CUs) of the video data.
[0063]
[0069] Residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inverse quantize and inverse transform the quantized transform coefficients of a block in order to reproduce the residual block of the block. The video decoder 300 uses the signaling prediction mode (intra or inter prediction) and related prediction information (for example, motion information for inter prediction) to form the prediction block of the block. The video decoder 300 can then combine (for each sample) the prediction block and the residual block to reproduce the original block. The video decoder 300 can perform additional processing such as performing a deblocking process to reduce visual artifacts along the boundaries of the block.
[0064]
[0070] In the present disclosure, generally, there may be references to "signaling" certain information such as syntax elements. The term "signaling" generally may refer to the communication of value syntax elements and / or other data used to decode encoded video data. That is, the video encoder 200 can signal the value of a syntax element in the bitstream. Generally, signaling refers to generating a value in the bitstream. As described above, the source device 102 can transfer the bitstream to the destination device 116 substantially in real time or transfer the bitstream to the destination device 116 non-real time such that it can occur when the syntax elements are stored in the storage device 112 for later retrieval by the destination device 116.
[0065]
[0071] According to the techniques of the present disclosure, the video encoder 200 and the video decoder 300 may be configured to perform gradient-based prediction improvement. As described above, as part of inter-predicting the current block, the video encoder 200 and the video decoder 300 may determine one or more prediction blocks for the current block (e.g., based on one or more motion vectors). In gradient-based prediction improvement, the video encoder 200 and the video decoder 300 modify one or more samples of the prediction block (e.g., including all samples).
[0066]
[0072] For example, in gradient-based prediction improvement, an inter-prediction sample at location (i, j) (e.g., a sample of the prediction block) is improved by an offset ΔI(i, j) derived from the displacement in the horizontal direction, the horizontal gradient, the displacement in the vertical direction, and the vertical gradient at location (i, j). In one example, the prediction improvement is ΔI(i, j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y described as (i,j), where g x (i,j) is the horizontal gradient, and g y (i,j) is the vertical gradient, and Δv x (i,j) is the displacement in the horizontal direction, and Δv y (i,j) is the displacement in the vertical direction.
[0067]
[0073] The gradient of an image is a measure of the directional change of intensity or color in the image. For example, the gradient value is based on the rate of change of color or intensity in the direction with the largest change in color or intensity based on adjacent samples. As an example, the gradient value is larger when the rate of change is relatively high than when the rate of change is relatively low.
[0068]
[0074] Furthermore, the predicted block for the current block can be in a reference picture different from the current picture containing the current block. The video encoder 200 and the video decoder 300 can determine an offset (e.g., ΔI(i,j)) based on sample values in the reference picture (e.g., the gradient is determined based on sample values in the reference picture). In some examples, the values used to determine the gradient can be values within the predicted block itself or values generated based on the values of the predicted block (e.g., interpolation values, rounded values, etc. generated from values within the predicted block). Also, in some examples, the values used to determine the gradient are outside the predicted block, can be within the reference picture, or can be generated from samples within the reference picture outside the predicted block (e.g., can be interpolated, rounded, etc.).
[0069]
[0075] However, in some examples, the video encoder 200 and the video decoder 300 can determine the offset based on sample values in the current picture. In some examples such as intra block copy, the current picture and the reference picture are the same picture.
[0070]
[0076] The displacement (e.g., vertical and / or horizontal displacement) can be determined based on the inter prediction mode. In some examples, the displacement is determined based on motion parameters. As will be described in more detail, in the case of the motion refinement mode on the decoder side, the displacement can be based on samples in the reference picture. In the case of other inter prediction modes, the displacement may not be based on samples in the reference picture, but the exemplary techniques are not so limited and samples in the reference picture can be used to determine the displacement. There can be various methods for determining the vertical and / or horizontal displacement, and the technique is not limited to a specific method for determining the vertical and / or horizontal displacement.
[0071]
[0077] The following describes an exemplary method for performing gradient calculation. For example, in the case of a gradient filter, in one example, a Sobel filter is used for gradient calculation. The gradient is g x (i,j)=I(i+1,j-1)-I(i-1,j-1)+2*I(i+1,j)-2*I(i-1,j)+I(i+1,j+1)-I(i-1,j+1) and g y (i,j) is calculated as I(i-1,j+1)-I(i-1,j-1)+2*I(i,j+1)-2*I(i,j-1)+I(i+1,j+1)-I(i+1,j-1).
[0072]
[0078] In some examples, a [1,0,-1] filter is applied. The gradient can be calculated as g x (i,j)=I(i+1,j)-I(i-1,j) and g y (i,j)=I(i,j+1)-I(i,j-1). In some examples, some other gradient filters (e.g., Canny filter) can be applied. The exemplary techniques described in this disclosure are not limited to any particular gradient filter.
[0073]
[0079] In the case of gradient normalization, the calculated gradient can be normalized before being used in the derivation of the improved offset (e.g., before calculating ΔI), or the normalization can be performed after the derivation of the improved offset. A rounding process can be applied during normalization. For example, when a [1,0,-1] filter is applied, the normalization is performed by adding 1 to the input value and then right-shifting by 1. When the input is scaled by a power of 2, the normalization is performed by adding 1<<N and then right-shifting by (N+1).
[0074]
[0080] For the case of the gradient at the boundary, the gradient at the boundary of the prediction block can be calculated by extending the prediction block by S / 2 at each boundary, where S is the filter processing step for gradient calculation. In one example, the extended prediction samples are generated by using the same motion vector as the prediction block for inter prediction (motion compensation). In some examples, the extended prediction samples use the same motion vector but are generated by using a shorter filter for the interpolation process in motion compensation. In some examples, the extended prediction samples are generated by using a rounded motion vector for integer motion compensation. In some examples, the extended prediction samples are generated by padding, where the padding is performed by copying the boundary samples. In some examples, when the prediction block is generated by sub-block based motion compensation, the extended prediction samples are generated by using the motion vector of the nearest sub-block. In some examples, when the prediction block is generated by sub-block based motion compensation, the extended prediction samples are generated by using one representative motion vector. In one example, the representative motion vector can be the motion vector at the center of the prediction block. In one example, the representative motion vector can be derived by averaging the motion vectors of the boundary sub-blocks.
[0075]
[0081] Sub-block based gradient derivation can be applied to facilitate parallel processing in hardware or a pipeline-friendly design. The width and height of the sub-blocks, denoted as sbW and sbH, can be determined as sbW = min(blkW, SB_WIDTH) and sbH = min(blkH, SB_HEIGHT). In this formula, blkW and blkH are the width and height of the prediction block respectively. SB_WIDTH and SB_HEIGHT are two predetermined variables. In one example, both SB_WIDTH and SB_HEIGHT are equal to 16.
[0076]
[0082] In the case of horizontal displacement and vertical displacement, the horizontal displacement and vertical displacement Δv used in the improved derivation x (i,j) and Δv y (i,j) can be determined according to the inter prediction mode in some examples. However, the exemplary techniques are not limited to determining the horizontal displacement and the vertical displacement based on the inter prediction mode.
[0077]
[0083] In the case of an inter mode with a small block size (for example, a small-sized block to be inter predicted), in order to reduce the worst memory bandwidth, the inter prediction mode for small blocks can be prohibited or restricted. For example, inter prediction for blocks of 4×4 or less is prohibited, and bidirectional prediction for 4×8, 8×4, 4×16, and 16×4 can be prohibited. The memory bandwidth can be increased by the interpolation process for those small blocks. Integer motion compensation without interpolation can still be applied to those small blocks without increasing the worst memory bandwidth.
[0078]
[0084] In one or more exemplary techniques, inter prediction can be made available for some or all of those small blocks, but using integer motion compensation and gradient-based prediction improvement. The motion vector can first be rounded to an integer motion vector for motion compensation. Then, the remainder of the rounding, i.e., the sub-pel portion of the motion vector, is used as Δv x (i,j) and Δv y for (i,j). For example, if the motion vector for a small block is (2.25, 5.75), the integer motion vector used for motion compensation will be (2, 6), the horizontal displacement (for example, Δv x (i,j)) will be 0.25, and the vertical displacement (for example, Δv y (i,j)) will be 0.75. In this example, the accuracy level of the horizontal displacement and the vertical displacement is 0.25 (or 1 / 4). For example, the horizontal displacement and the vertical displacement can be incremented in steps of 0.25.
[0079]
[0085] In some examples, for small block size inter modes, gradient-based prediction refinement may be available only when small sized blocks are inter predicted in merge mode. Examples of merge mode are described below. In some examples, for small size inter modes, gradient-based prediction refinement may be prohibited for blocks with integer motion modes. In an integer motion mode, one or more motion vectors (e.g., the signaled motion vectors) are integers. In some examples, even for larger sized blocks, gradient-based prediction refinement may be prohibited for blocks that are inter predicted in integer motion mode.
[0080]
[0086] In the case of the normal merge mode, which is an example of an inter prediction mode, when motion information is derived from spatially or temporally adjacent coded blocks, Δv x (i,j) and Δv y (i,j) can be the remainder of a motion vector rounding process (similar to the above example of the motion vector (2.25, 5.75)). In one example, a temporal motion vector predictor is derived by scaling the motion vectors in a temporal motion buffer according to the difference in picture order count between the current picture and a reference picture. A rounding process can be performed to round the motion vectors scaled to a certain precision. The remainder can be used as Δv x (i,j) and Δv y (i,j). The remaining precision (i.e., the precision level of the horizontal displacement and the vertical displacement) can be predefined and can be higher than the precision of the motion vector prediction. For example, when the motion vector precision is 1 / 16, the remaining precision is 1 / (16*MaxBlkSize), where MaxBlkSize is the maximum block size. In other words, the precision level for the horizontal displacement and the vertical displacement (e.g., Δv x and Δv y and) is 1 / (16*MaxBlkSize).
[0081]
[0087] In the case of merge using the motion vector difference (MMVD) mode, which is an example of the inter prediction mode, the motion vector difference is signaled together with the merge index to represent motion information. In some techniques, the motion vector difference (e.g., the difference between the actual motion vector and the motion vector predictor) has a motion vector of the same accuracy. In one or more examples described in the present disclosure, the motion vector difference can be enabled to have higher accuracy. The signaled motion vector difference is first rounded to the motion vector accuracy, and the motion vector indicated by the merge index is added to generate the final motion vector for motion compensation. In one or more examples, the remaining portion after rounding (e.g., the difference between the rounded value of the motion vector difference and the original value of the motion vector difference) can be used as the horizontal displacement and the vertical displacement for gradient-based prediction improvement (e.g., Δv x (i,j) and Δv y (i,j) can be used). In some examples, Δv x (i,j) and Δv y (i,j) can be signaled as candidates for the motion vector difference.
[0082]
[0088] In the case of the motion vector improvement mode on the decoder side, motion compensation using the original motion vector is performed to generate the original bi-predicted block, and the difference between the list 0 and list 1 predictions, denoted as DistOrig, is calculated. List 0 refers to the first reference picture list (RefPicList0) that contains a list of reference pictures that can potentially be used for inter prediction. List 1 refers to the second reference picture list (RefPicList1) that contains a list of reference pictures that can potentially be used for inter prediction. Then, the motion vectors of list 0 and list 1 are rounded to the nearest integer positions. That is, the motion vector pointing to the picture in list 0 is rounded to the nearest integer position, and the motion vector pointing to the picture in list 1 is rounded to the nearest integer position. A search algorithm is used to search within the range of integer displacements to find a pair of displacements with the minimum distortion DistNew between the block of the picture identified in the list 0 prediction and the block of the picture identified in list 1 using the new integer motion vectors for motion compensation. If DistNew is smaller than DistOrig, Δv x (i,j) and Δv y (i,j) are derived by supplying the new integer motion vectors to the bidirectional optical flow (BDOF). Otherwise, BDOF is performed on the original list 0 and list 1 predictions for prediction improvement.
[0083]
[0089] In the case of affine mode, the motion field can be derived for each pixel (e.g., the motion vector can be determined for each pixel). However, a 4×4-based motion field is used for affine motion compensation to reduce complexity and memory bandwidth. For example, instead of determining the motion vector for each pixel, the motion vector is determined for sub-blocks, where one sub-block is, for example, 4×4. Some other sub-block sizes, such as 4×2, 2×4, or 2×2, can also be used. In one or more examples, gradient-based prediction refinement can be used to improve affine motion compensation. The gradient of the block can be calculated as described above. Assuming an affine motion model,
[0084]
Number
[0085] where a, b, c, d, e, and f are values determined by the video encoder 200 and the video decoder 300 based on, for example, the control point motion vectors and the length and width of the block. The values of a, b, c, d, e, and f can be signaled in some examples.
[0086]
[0090] Several example methods for determining a, b, c, d, e, and f are described below. In a video encoder (e.g., video encoder 200 or video decoder 300), a picture is partitioned into sub-blocks for block-based coding in affine mode. The affine motion model for a block also requires three motion vectors (MV: motion vector) at three different locations that are not on the same line
[0087]
Number
[0088] ,
[0089]
Number
[0090] 、and
[0091]
Number
[0092] can be described by. The three locations are usually called control points, and the three motion vectors are called control-point motion vectors (CPMVs). When the three control points are at the three corners of the block, the affine motion is
[0093]
Number
[0094] can be described as, where blkW and blkH are the width and height of the block.
[0095]
[0091] In the case of the affine mode, the video encoder 200 and the video decoder 300 can determine the motion vector for each sub-block using the representative coordinates of the sub-block (e.g., the center position of the sub-block). In one example, the block is divided into non-overlapping sub-blocks. The width of the block is blkW, the height of the block is blkH, the width of the sub-block is sbW, the height of the sub-block is sbH, and there are blkH / sbH rows of sub-blocks, and blkW / sbW sub-blocks in each row. For the 6-parameter affine motion model, the motion vector for the sub-block (referred to as the sub-block MV) in the i-th row (0 <= i < blkW / sbW) and the j-th column (0 <= j < blkH / sbH) is
[0096]
Number
[0097] is derived as
[0098]
[0092] From the above equation, the variables a, b, c, d, e, and f can be defined as follows.
[0099]
Equation
[0100]
[0093] In the case of the affine mode which is an example of the inter prediction mode, the video encoder 200 and the video decoder 300 can determine a displacement (for example, a horizontal displacement or a vertical displacement) by at least one of the following methods. The following is an example and should not be regarded as a limitation. There may be other methods by which the video encoder 200 and the video decoder 300 can determine a displacement (for example, a horizontal displacement or a vertical displacement) for the affine mode.
[0101]
[0094] In the case of 4×4 sub-block based affine motion compensation, for the 2×2 based displacement derivation, the displacement within each 2×2 sub-block is the same. Within each 4×4 sub-block, Δ(i,j) for the four 2×2 sub-blocks within the 4×4 is calculated as follows.
[0102]
Equation
[0103]
[0095] In the case of 1×1 displacement derivation, the displacement is derived for each sample. The coordinates of the top-left sample within the 4×4 can be (0,0), in which case Δv(i,j) is derived as follows.
[0104]
Equation
[0105]
[0096] In some examples, the division by 2 implemented as a right shift operation can be moved to the improved offset calculation. For example, instead of performing the division by 2 operation when deriving the horizontal displacement and the vertical displacement (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform the division by 2 operation as part of determining ΔI (e.g., the improved offset).
[0106]
[0097] In the case of 4×2 sub-block based affine motion compensation, the motion field for storing the motion vectors is still 4×4, but the affine motion compensation is 4×2. The motion vector (MV) for a 4×4 sub-block can be (v x ,v y ), in which case the MV for the left 4×2 motion compensation is (v x -a,v y -c), and the MV for the right 4×2 motion compensation is (v x +a,v y +c).
[0107]
[0098] In the case of 2×2 based displacement derivation, in 2×2 based displacement derivation, the displacements within each 2×2 sub-block are the same. Within each 4×2 sub-block, the Δ(i,j) for two 2×2 sub-blocks within 4×4 is calculated as follows.
[0108]
Number
[0109]
[0099] In the case of 1×1 displacement derivation, the displacement is derived for each sample. Assuming the coordinates of the top-left sample in 4×2 are (0,0), Δv(i,j) is derived as follows.
[0110]
Number
[0111]
[0100] Division by 2, which can be implemented as a right shift operation, can be moved to the improved offset calculation. For example, instead of performing a division operation by 2 when deriving a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform a division operation by 2 as part of determining ΔI (e.g., the improved offset).
[0112]
[0101] In the case of 2×4 sub-block based affine motion compensation, the motion field for storing the motion vectors is still 4×4, but the affine motion compensation is 2×4. The MV for a 4×4 sub-block can be (v x , v y ), in which case the MV for the left 4×2 motion compensation is (v x -b, v y -d), and the MV for the right 4×2 motion compensation is (v x +b, v y +d).
[0113]
[0102] In the case of 2×2 based displacement derivation, the displacements within each 2×2 sub-block are the same. Within each 2×4 sub-block, the Δ(i,j) for two 2×2 sub-blocks within the 2×4 is calculated as follows.
[0114]
Equation
[0115]
[0103] In the case of 1×1 displacement derivation, in 1×1 based displacement derivation, the displacement is derived for each sample. The coordinates of the top left sample within 2×4 can be (0,0), in which case the Δv(i,j) is derived as follows.
[0116]
Equation
[0117]
[0104] Division by 2 that can be implemented as a right shift operation can be moved to the improved offset calculation. For example, instead of performing a division operation by 2 when deriving a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y ), the video encoder 200 and the video decoder 300 can perform a division operation by 2 as part of determining ΔI (e.g., the improved offset).
[0118]
[0105] The following describes the prediction improvement for the affine mode. After sub-block based affine motion compensation is performed, the prediction signal can be improved by adding an offset derived based on the motion in pixel units and the gradient of the prediction signal. The offset at location (m, n) can be calculated as follows.
[0119]
Equation
[0120]
[0106] Here, g x (m, n) is the horizontal gradient of the prediction signal, and g y (m, n) is the vertical gradient. Δv x (m, n) and Δv y (m, n) are the differences in the x and y components between the motion vector calculated at the location pixel location (m, n) and the sub-block MV. Assuming that the coordinates of the upper left sample of the sub-block are (0, 0), the center of the sub-block is
[0121]
Equation
[0122] is. Given the affine motion parameters a, b, c, and d, Δv x (m, n) and Δv y(m,n) can be derived as follows.
[0123]
Number
[0124]
[0107] In the control point-based affine motion model, the affine motion parameters a, b, c, and d are calculated from the CPMV as follows.
[0125]
Number
[0126]
[0108] The following describes the bidirectional optical flow (BDOF). The bidirectional optical flow (BDOF) tool is included in VTM4. BDOF was previously called BIO. BDOF can be used to improve the dual prediction signal of a coding unit (CU) at the 4×4 sub-block level. The BDOF mode is based on the optical flow concept, which assumes that the motion of the object is smooth. For each 4×4 sub-block, the motion improvement (v x ,v y ) is calculated by minimizing the difference between the L0 prediction sample and the L1 prediction sample (e.g., the prediction sample from the reference picture in the first reference picture list L0 and the prediction sample from the reference picture in the second reference picture list L1). The motion improvement is then used to adjust the dual prediction sample values in the 4×4 sub-block. In the BDOF process, the following steps are applied.
[0127]
[0109] First, the horizontal and vertical gradients of the two prediction signals
[0128]
Number
[0129] and
[0130]
Number
[0131] , k = 0, 1 are calculated by directly calculating the difference between two adjacent samples, that is, as follows.
[0132]
Number
[0133]
[0110] Here, I (k) (i, j) is the sample value at the coordinates (i, j) of the predicted signal in list k, where k = 0, 1.
[0134]
[0111] Then, the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated for autocorrelation and cross - correlation as follows.
[0135]
Number
[0136]
[0112] Here,
[0137]
Number
[0138] is.
[0139]
[0113] Here, Ω is a 6×6 window around a 4×4 sub - block.
[0140]
[0114] Motion improvement (v x , v y) is then derived using the cross - correlation terms and auto - correlation terms as follows.
[0141]
Number
[0142]
[0115] Here,
[0143]
Number
[0144] is the floor function.
[0145]
[0116] Based on motion improvement and gradient, the following adjustments are calculated for each sample in the 4×4 sub - block.
[0146]
Number
[0147]
[0117] Finally, the CU's BDOF samples are calculated by adjusting the dual - prediction samples as follows.
[0148]
Number
[0149]
[0118] These values are selected so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit - width of the intermediate parameters in the BDOF process is kept within 32 bits.
[0150]
[0119] To derive the gradient value, some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k)(i,j) needs to be generated. As shown in Figure 5, the BDOF in VTM4 uses one row / column extended around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly taking the reference samples at nearby integer positions without interpolation (by using the floor() operation on the coordinates), and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculation. In the remaining steps during the BDOF process, if any samples and gradient values outside the CU boundary are required, such samples are padded (i.e., repeated) from their nearest neighbors.
[0151]
[0120] The following explains the accuracy of displacement and gradient. In some examples, the same accuracy can be used in all modes for both horizontal and vertical displacements. The accuracy can be predefined or signaled in the high-level syntax. Thus, if the horizontal and vertical displacements are derived from different modes with different accuracies, the horizontal and vertical displacements are rounded to the predefined accuracy. Examples of predefined accuracy are 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.
[0152]
[0121] As explained above, the accuracy, also called the accuracy level, is for the horizontal and vertical displacements (e.g., Δv x and Δv y) can indicate how accurate it is, where the horizontal displacement and the vertical displacement can be determined using one or more of the examples described above or using some other techniques. Generally, the accuracy level is defined as a decimal (e.g., 0.25, 0.125, 0.0625, 0.03125, 0.015625, 0.0078125, etc.) or a fraction (e.g., 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64, 1 / 128, etc.). For example, for an accuracy level of 1 / 4, the horizontal displacement or the vertical displacement can be represented in increments of 0.25 (e.g., 0.25, 0.5, or 0.75). For an accuracy level of 1 / 8, the horizontal displacement and the vertical displacement can be represented in increments of 0.125 (e.g., 0.125, 0.25, 0.325, 0.5, 0.625, 0.75, or 0.825). As can be seen, the smaller the value of the accuracy level (e.g., 1 / 8 is smaller than 1 / 4), the larger the certain granularity for the increments, and the more accurate the values that can be presented (e.g., for an accuracy level of 1 / 4, the displacement is rounded to the nearest 1 / 4, but for an accuracy level of 1 / 8, the displacement is rounded to the nearest 1 / 8).
[0153]
[0122] Since the horizontal displacement and the vertical displacement can have different accuracy levels for different inter prediction modes, the video encoder 200 and the video decoder 300 can be configured to include different logic circuits to perform gradient-based prediction improvements for different inter prediction modes. As described above, to perform gradient-based prediction improvements, the video encoder 200 and the video decoder 300 perform g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) operations, where g x and g y are a first gradient based on a first set of samples among the samples of the prediction block and a second gradient based on a second set of samples among the samples of the prediction block, respectively, and Δv x and Δv yare a horizontal displacement and a vertical displacement, respectively. As can be seen, for gradient-based prediction improvement, the video encoder 200 and the video decoder 300 need to perform multiplication and addition operations, and may utilize memory to store temporary results used in the calculations.
[0154]
[0123] However, the capabilities of logic circuits that perform mathematical operations (e.g., multiplication circuits, adder circuits, memory registers) can be limited by the accuracy level at which the logic circuits are configured. For example, a logic circuit configured for a first accuracy level may not be able to perform operations required for gradient prediction improvement where the horizontal or vertical displacement is at a more accurate second accuracy level.
[0155]
[0124] Accordingly, some techniques utilize different sets of logic circuits configured for different accuracy levels to perform gradient-based prediction improvement for different inter prediction modes. For example, a first set of logic circuits may be configured to perform gradient-based prediction improvement for an inter prediction mode where the horizontal and / or vertical displacement is 0.25, and a second set of logic circuits may be configured to perform gradient-based prediction improvement for an inter prediction mode where the horizontal and / or vertical displacement is 0.125. Having these different sets of circuits increases the overall size of the video encoder 200 and the video decoder 300 as well as potentially wasted power.
[0156]
[0125] In some examples described in the present disclosure, the same gradient calculation process can be used for all inter prediction modes. In other words, the same logic circuits can be used to perform gradient-based prediction improvement for different inter prediction modes. For example, the accuracy of the gradient can be kept the same for prediction improvement in all inter prediction modes. In some examples, with respect to displacement and gradient accuracy, exemplary techniques can ensure that the same (or unified) prediction improvement process can be applied to different modes, and the same prediction improvement module can be applied to different modes.
[0157]
[0126] As an example, the video encoder 200 and the video decoder 300 may be configured to round at least one of a horizontal displacement and a vertical displacement to the same accuracy level for different inter prediction modes (e.g., the same for an affine mode and BDOF). For example, when the accuracy level to which the horizontal displacement and the vertical displacement are rounded is 0.015625 (1 / 64), and when the accuracy level of the horizontal displacement and / or the vertical displacement is 1 / 4 for one inter prediction mode, the accuracy level of the horizontal displacement and / or the vertical displacement is rounded to 1 / 64. When the accuracy level of the horizontal displacement and / or the vertical displacement is 1 / 128, the accuracy level of the horizontal displacement and / or the vertical displacement is rounded to 1 / 64.
[0158]
[0127] In this way, a logic circuit for gradient-based prediction improvement can be reused for different inter prediction modes. For example, in the above example, the video encoder 200 and the video decoder 300 may include a logic circuit for an accuracy level of 0.125, and since the accuracy level of the horizontal displacement and / or the vertical displacement is rounded to 0.125, this logic circuit can be reused for different inter prediction modes.
[0159]
[0128] In some examples, when rounding is not performed according to the techniques described in this disclosure, when the logic circuit is designed to have a relatively high level of accuracy, a logic circuit for multiplication and accumulation type operations can be reused (e.g., a logic circuit designed for a specific accuracy level for multiplication can handle multiplication operations for lower accuracy level values). However, in the case of a shift operation, a logic circuit designed for a specific accuracy may not be able to handle shift operations for lower accuracy level values. Using the exemplary techniques described in this disclosure, it may be possible to reuse the logic circuits included for shift operations for different inter prediction modes using the rounding techniques described.
[0160]
[0129] In one example, the prediction improvement offset is derived as follows.
[0161]
Number
[0162]
[0130] In the above equation, the offset is equal to 1 << (shift - 1), the shift is determined by the predefined precision of the displacement and gradient, and is fixed for different modes. In some examples, the offset is equal to 0.
[0163]
[0131] In some examples, the mode may include one or more of the modes described above with respect to horizontal displacement and vertical displacement, such as inter - mode with a small block size, normal merge mode, merge with motion vector difference, motion vector improvement mode on the decoder side, and affine mode. The mode may also include the bidirectional optical flow (BDOF) described above.
[0164]
[0132] There may be separate improvements for each prediction direction. For example, in the case of bidirectional prediction, the prediction improvement may be performed separately for each prediction direction. The result of the improvement may be clipped to a certain range to ensure the same bit - width as the prediction without improvement. For example, the improvement result is clipped to a 16 - bit range. As described above, exemplary techniques may also be applied to BDOF, where displacements in two different directions are considered to be on the same motion trajectory.
[0165]
[0133] The following describes the multiplication constraint for N bits (e.g., 16 bits). To reduce the complexity of gradient - based prediction improvement, the multiplication can be kept within N bits (e.g., 16 bits). The gradient and displacement must be able to be represented by only 16 bits in this example. Otherwise, the gradient or displacement is quantized to be within 16 bits in this example. For example, a right - shift can be applied to maintain a 16 - bit representation.
[0166]
[0134] The following describes the clipping of the improved offset ΔI(i,j) and the result of the improvement. The improved offset ΔI(i,j) is clipped within a certain range. In one example, the range is determined by the range of the original prediction signal. The range of ΔI(i,j) can be the same as the range of the original prediction signal, or the range can be a scaled range. The scale can be 1 / 2, 1 / 4, 1 / 8, etc. The improvement result is clipped to have the same range as the original prediction signal (e.g., the range of samples in the prediction block). The equation for performing the clipping is as follows.
[0167]
Number
[0168]
[0135] Here, predSamplesL0 and predSamplesL1 are the prediction samples in each single prediction direction. bdofOffset is the improved offset derived by BDOF. Offset4 = 1<<(shift4 - 1), and Clip3(min, max, x) is a function that clips the value of x to be within the range including the minimum and maximum.
[0169]
[0136] In this way, the video encoder 200 and the video decoder 300 can be configured to determine a prediction block for inter-predicting the current block. For example, the video encoder 200 and the video decoder 300 can determine a motion vector or a block vector that points to the prediction block (e.g., for the intra-block copy mode).
[0170]
[0137] The video encoder 200 and the video decoder 300 can determine at least one of a horizontal displacement or a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block. An example of the horizontal displacement is Δv x and an example of the vertical displacement is Δv yIt is. In some examples, the video encoder 200 and the video decoder 300 may determine at least one of a horizontal displacement or a vertical displacement for gradient-based prediction improvement of one or more samples of a prediction block based on an inter prediction mode (for example, as two examples, using the above exemplary techniques for the affine mode, Δv x and Δv y can be determined, or using the above exemplary techniques for the merge mode, Δv x and Δv y can be determined).
[0171]
[0138] According to one or more examples, the video encoder 200 and the video decoder 300 may round at least one of a horizontal displacement and a vertical displacement to the same accuracy level for different inter prediction modes. Examples of different inter prediction modes include the affine mode and BDOF. For example, the accuracy level for a first horizontal displacement or vertical displacement for gradient-based prediction improvement for a first block inter predicted in a first inter prediction mode may be at a first accuracy level, and the accuracy level for a second horizontal displacement or vertical displacement for gradient-based prediction improvement for a second block inter predicted in a second inter prediction mode may be at a second accuracy level. The video encoder 200 and the video decoder 300 may be configured to round the first accuracy level for the first horizontal displacement or vertical displacement to the accuracy level and round the second accuracy level for the first horizontal displacement or vertical displacement to the same accuracy level.
[0172]
[0139] In some examples, the accuracy level may be predefined (for example, may be pre-stored on the video encoder 200 and the video decoder 300) or may be signaled (for example, defined by the video encoder 200 and signaled to the video decoder 300). In some examples, the accuracy level may be 1 / 64.
[0173]
[0140] The video encoder 200 and the video decoder 300 may be configured to determine one or more improved offsets based on at least one of a rounded horizontal displacement or a rounded vertical displacement. For example, the video encoder 200 and the video decoder 300 may determine ΔI(i,j) for each sample of the prediction block using at least one of the rounded horizontal displacement or the rounded vertical displacement of each. That is, the video encoder 200 and the video decoder 300 may determine an improved offset for each sample of the prediction block. In some examples, the video encoder 200 and the video decoder 300 may utilize the rounded horizontal displacement and the rounded vertical displacement to determine an improved offset (e.g., ΔI).
[0174]
[0141] As described, to perform gradient-based prediction refinement, the video encoder 200 and the video decoder 300 may determine a first gradient based on a first set of samples of one or more samples of the prediction block (e.g., may determine g x (i,j), where the first set of samples is the samples used to determine g x (i,j)), and may determine a second gradient based on a second set of samples of one or more samples of the prediction block (e.g., may determine g y (i,j), where the second set of samples is the samples used to determine g y (i,j)). The video encoder 200 and the video decoder 300 may determine an improved offset based on the rounded horizontal displacement, the rounded vertical displacement, the first gradient, and the second gradient.
[0175]
[0142] The video encoder 200 and the video decoder 300 may modify one or more samples of a prediction block based on one or more determined improved offsets to generate a modified prediction block (e.g., one or more modified samples that form the modified prediction block). For example, the video encoder 200 and the video decoder 300 may add or subtract ΔI(i,j) from I(i,j), where I(i,j) refers to a sample in the prediction block located at position (i,j). In some examples, the video encoder 200 and the video decoder 300 may clip one or more improved offsets (e.g., may clip ΔI(i,j)). The video encoder 200 and the video decoder 300 may modify one or more samples of the prediction block based on the one or more clipped improved offsets.
[0176]
[0143] In the case of encoding, the video encoder 200 may determine a residual value (e.g., of a residual block) indicating the difference between the current block and the modified prediction block (e.g., based on the modified samples of the modified prediction block), and may signal information indicating the residual value. In the case of decoding, the video decoder 300 may receive the information indicating the residual value and may reconstruct the current block based on the modified prediction block (e.g., the modified samples of the modified prediction block) and the residual value (e.g., by adding the residual value to the modified samples).
[0177]
[0144] FIGS. 2A and 2B are conceptual diagrams showing an exemplary quadtree binary tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quadtree divisions, and dotted lines represent binary tree divisions. At each division (i.e., non-leaf) node of the binary tree, one flag is signaled to indicate which division type (i.e., horizontal or vertical) is used. Here, in this example, 0 indicates a horizontal division and 1 indicates a vertical division. In the case of a quadtree division, since the quadtree node divides the block horizontally and vertically into four sub-blocks of equal size, there is no need to indicate the division type. Thus, the syntax elements (such as division information) for the region tree level (i.e., solid lines) of the QTBT structure 130 and the syntax elements (such as division information) for the prediction tree level (i.e., dashed lines) of the QTBT structure 130 can be encoded by the video encoder 200 and decoded by the video decoder 300. The video encoder 200 can encode video data such as prediction and transform data for the CU represented by the terminal leaf node of the QTBT structure 130, and the video decoder 300 can decode it.
[0178]
[0145] Generally, the CTU 132 of FIG. 2B can be associated with parameters that define the sizes of the blocks corresponding to the nodes of the QTBT structure 130 at the first and second levels. These parameters can include the CTU size (representing the size of the CTU 132 in sample units), the minimum quadtree size (MinQTSize, representing the minimum allowable quadtree leaf node size), the maximum binary tree size (MaxBTSize, representing the maximum allowable binary tree root node size), the maximum binary tree depth (MaxBTDepth, representing the maximum allowable binary tree depth), and the minimum binary tree size (MinBTSize, representing the minimum allowable binary tree leaf node size).
[0179]
[0146] The root node of the QTBT structure corresponding to the CTU may have four child nodes at the first level of the QTBT structure, each of which may be partitioned according to a quadtree partition. That is, the nodes at the first level are either leaf nodes (having no child nodes) or have four child nodes. An example of the QTBT structure 130 represents nodes such as having a parent node and child nodes with solid lines for branches. If the nodes at the first level are not larger than the maximum allowable binary tree root node size (MaxBTSize), the nodes may be further partitioned by respective binary trees. The binary tree splitting of one node may be repeated until the nodes obtained from the splitting reach the minimum allowable binary tree leaf node size (MinBTSize) or the maximum allowable binary tree depth (MaxBTDepth). An example of the QTBT structure 130 represents nodes such as having a dashed line for branches. The binary tree leaf nodes are called coding units (CUs), and the coding units (CUs) are used for prediction (e.g., intra-picture prediction or inter-picture prediction) and transformation without further partitioning. As discussed above, the CU may also be called a "video block" or a "block".
[0180]
[0147] In an example of the QTBT partitioning structure, the CTU size is set to 128×128 (luma samples and two corresponding 64×64 chroma samples), MinQTSize is set to 16×16, MaxBTSize is set to 64×64, MinBTSize is set to 4 (for both width and height), and MaxBTDepth is set to 4. To generate the quadtree leaf nodes, first, the quadtree partitioning is applied to the CTU. The quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the quadtree tree node is 128×128, since its size exceeds MaxBTSize (i.e., 64×64 in this example), it is not further split by the binary tree. Otherwise, the leaf quadtree node is further partitioned by the binary tree. Thus, the quadtree leaf node is also the root node for the binary tree and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (4 in this example), no further splitting is allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), that implies no further horizontal splitting is allowed. Similarly, a binary tree node with a height equal to MinBTSize implies no further vertical splitting is allowed for that binary tree node. As described above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further partitioning.
[0181]
[0148] FIG. 3 is a block diagram showing an exemplary video encoder 200 that can execute the techniques of the present disclosure. FIG. 3 is provided for illustration purposes and should not be regarded as limiting the techniques widely exemplified and described in the present disclosure. For the purpose of explanation, in the present disclosure, the video encoder 200 is described in the context of video coding standards such as the HEVC video coding standard and the developing H.266 video coding standard. However, the techniques of the present disclosure are not limited to these video coding standards and are generally applicable to video encoding and decoding.
[0182]
[0149] In the example of FIG. 3, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy encoding unit 220. Any or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the conversion processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse conversion processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy encoding unit 220 may be implemented in one or more processors or in a processing circuit. Moreover, the video encoder 200 may include additional or alternative processors or processing circuits for performing these and other functions.
[0183]
[0150] The video data memory 230 may store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive, for example, video data stored in the video data memory 230 from the video source 104 (FIG. 1). The DPB 218 may act as a reference picture memory that stores reference video data for use in the prediction of subsequent video data by the video encoder 200. The video data memory 230 and the DPB 218 may be formed by any of various memory devices, such as a dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM (registered trademark)), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or by separate memory devices. In various examples, the video data memory 230 may be on-chip with or off-chip from the other components of the video encoder 200, as illustrated.
[0184]
[0151] In the present disclosure, references to the video data memory 230 should not be construed as being limited to the memory internal to the video encoder 200, unless specifically stated otherwise, or as being limited to the memory external to the video encoder 200, unless specifically stated otherwise. Rather, references to the video data memory 230 are intended to be understood as a reference memory that stores video data received by the video encoder 200 for encoding (e.g., the video data of the current block to be encoded). The memory 106 of FIG. 1 may also provide temporary storage of outputs from various units of the video encoder 200.
[0185]
[0152] The various units of FIG. 3 are shown to assist in understanding the operations performed by the video encoder 200. The units may be implemented as fixed-function circuitry, programmable circuitry, or a combination thereof. Fixed-function circuitry refers to circuitry that provides a specific function and is pre-set with respect to the operations that can be performed. Programmable circuitry refers to circuitry that is programmed to perform various tasks and to provide a flexible function in the operations that can be performed. For example, the programmable circuitry may execute software or firmware that causes the programmable circuitry to operate in a manner defined by software or firmware instructions. Fixed-function circuitry may execute software instructions (e.g., to receive or output parameters), but the type of operations performed by the fixed-function circuitry is generally immutable. In some examples, one or more of the units may be separate circuit blocks (fixed-function or programmable), and in some examples, one or more units may be integrated circuits.
[0186]
[0153] The video encoder 200 may include an arithmetic logic unit (ALU) formed from a programmable circuit, an elementary function unit (EFU), a digital circuit, an analog circuit, and / or a programmable core. In an example where the operation of the video encoder 200 is performed using software executed by a programmable circuit, the memory 106 (FIG. 1) may store the object code of the software that the video encoder 200 receives and executes, or another memory (not shown) within the video encoder 200 may store such instructions.
[0187]
[0154] The video data memory 230 is configured to store received video data. The video encoder 200 may retrieve a picture of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be raw video data to be encoded.
[0188]
[0155] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, an intra prediction unit 226, and a gradient-based prediction refinement (GBPR) unit 227. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. By way of example, the mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, and the like.
[0189]
[0156] Although the GBPR unit 227 is shown as being separate from the motion estimation unit 222 and the motion compensation unit 224, in some examples, the GBPR unit 227 can be part of the motion estimation unit 222 and / or the motion compensation unit 224. The GBPR unit 227 is shown separately from the motion estimation unit 222 and the motion compensation unit 224 for ease of understanding and should not be considered limiting.
[0190]
[0157] The mode selection unit 202 generally coordinates multiple encoding paths to test combinations of encoding parameters and generates rate distortion values for such combinations. The encoding parameters can include the partitioning of CTUs into CUs, the prediction mode for the CU, the transform type for the residual values of the CU, the quantization parameters for the residual values of the CU, and the like. The mode selection unit 202 can ultimately select a combination of encoding parameters that has a better rate distortion value than other tested combinations.
[0191]
[0158] The video encoder 200 can partition a picture fetched from the video data memory 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can partition the CTUs of the picture according to a tree structure such as the HEVC QTBT structure or the quadtree structure described above. As described above, the video encoder 200 can form one or more CUs from partitioning the CTUs according to the tree structure. Such CUs are sometimes also generally referred to as "video blocks" or "blocks".
[0192]
[0159] Generally, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, the intra prediction unit 226, and the GBPR unit 227) to generate a prediction block for the current block (e.g., the current CU, or in HEVC, the overlapping part of the PU and TU). For the inter prediction of the current block, the motion estimation unit 222 may perform a motion search to identify one or more exactly matching reference blocks in one or more reference pictures (e.g., one or more previous coded pictures stored in the DPB 218). In particular, the motion estimation unit 222 may calculate a value representing how similar a potential reference block is to the current block, for example, according to the sum of absolute differences (SAD), the sum of squared differences (SSD), the mean absolute difference (MAD), the mean squared difference (MSD), etc. The motion estimation unit 222 may generally perform these calculations using the sample-by-sample differences between the current block and the reference block under consideration. The motion estimation unit 222 may identify the reference block having the lowest value obtained from these calculations, which indicates the reference block that most closely matches the current block.
[0193]
[0160] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in a current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, in uni-directional inter prediction, the motion estimation unit 222 may provide a single motion vector, while in bi-directional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then generate a prediction block using the motion vectors. For example, the motion compensation unit 224 may retrieve data of a reference block using the motion vectors. As another example, when the motion vectors have sub-sample accuracy, the motion compensation unit 224 may interpolate the values of the prediction block according to one or more interpolation filters. Moreover, in the case of bi-directional inter prediction, the motion compensation unit 224 may retrieve data for two reference blocks specified by respective motion vectors and combine the retrieved data, for example, through sample-by-sample averaging or weighted averaging.
[0194]
[0161] As another example, for intra prediction, or intra prediction coding, the intra prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, in a directional mode, the intra prediction unit 226 may generally mathematically combine the values of adjacent samples and populate these calculated values in a defined direction across the current block to generate a prediction block. As another example, in a DC mode, the intra prediction unit 226 may calculate the average of adjacent samples for the current block and generate a prediction block that includes this obtained average for each sample of the prediction block.
[0195]
[0162] The GBPR unit 227 can be configured to perform the exemplary techniques described in this disclosure for gradient-based prediction improvement. For example, the GBPR unit 227, together with the motion compensation unit 224, can determine a prediction block for inter-predicting a current block (e.g., based on the motion vectors determined by the motion estimation unit 222). The GBPR unit 227 can determine a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y and) for gradient-based prediction improvement of one or more samples of the prediction block. As an example, the GBPR unit 227 can determine an inter-prediction mode based on the determination made by the mode selection unit 202 for inter-predicting the current block. In some examples, the GBPR unit 227 can determine the horizontal displacement and the vertical displacement based on the determined inter-prediction mode.
[0196]
[0163] The GBPR unit 227 can round the horizontal displacement and the vertical displacement to the same level of accuracy for different inter-prediction modes. For example, the current block can be a first current block, the prediction block can be a first prediction block, the horizontal displacement and the vertical displacement can be a first horizontal displacement and vertical displacement, and the rounded horizontal displacement and vertical displacement can be a first rounded horizontal displacement and vertical displacement. In some examples, the GBPR unit 227 can determine a second prediction block for inter-predicting a second current block and determine a second horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block. The GBPR unit 227 can round the second horizontal displacement and the vertical displacement to the same level of accuracy at which the first horizontal displacement and the vertical displacement were rounded to generate the second rounded horizontal displacement and vertical displacement.
[0197]
[0164] In some cases, the inter prediction mode for inter predicting the first current block may be different from the inter prediction mode for the second current block. For example, the first mode of different inter prediction modes is the affine mode, and the second mode of different inter prediction modes is the bidirectional optical flow (BDOF) mode.
[0198]
[0165] The accuracy level at which the horizontal displacement and the vertical displacement are rounded can be predefined and stored for use by the GBPR unit 227, or the GBPR unit 227 can determine the accuracy level, and the video encoder 200 can signal the accuracy level. As an example, the accuracy level is 1 / 64.
[0199]
[0166] The GBPR unit 227 can determine one or more improved offsets based on the rounded horizontal displacement and vertical displacement. For example, the GBPR unit 227 determines a first gradient based on a first set of samples among one or more samples of the prediction block (for example, using the samples of the prediction block described above to determine g x (i,j)), and may determine a second gradient based on a second set of samples among one or more samples of the prediction block (for example, using the samples of the prediction block described above to determine g y (i,j)). The GBPR unit 227 can determine one or more improved offsets based on the rounded horizontal displacement and vertical displacement, the first gradient, and the second gradient. In some examples, if the value of one or more improved offsets is too high (for example, greater than a threshold), the GBPR unit 227 can clip one or more improved offsets.
[0200]
[0167] The GBPR unit 227 may modify one or more samples of a prediction block based on one or more determined improved offsets or one or more clipped improved offsets to generate a modified prediction block (e.g., one or more modified samples that form the modified prediction block). For example, the GBPR unit 227 may determine g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among one or more samples located at (i,j), and Δv x (i,j) is the rounded horizontal displacement for the sample among one or more samples located at (i,j), and g y (i,j) is the second gradient for the sample among one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among one or more samples located at (i,j). In some examples, Δv x and Δv y may be the same for each sample (i,j) of the prediction block.
[0201]
[0168] The obtained corrected sample may form a prediction block (e.g., a corrected prediction block) during gradient-based prediction refinement. That is, the corrected prediction block is used as a prediction block during gradient-based prediction refinement. The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives the raw, unencoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The obtained sample-by-sample difference defines the residual block for the current block. In some examples, the residual generation unit 204 may also determine the difference between sample values in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.
[0202]
[0169] In an example where the mode selection unit 202 partitions a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs of various sizes. As described above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2N×2N, the video encoder 200 may support PU sizes of 2N×2N or N×N for intra prediction and PU sizes of 2N×2N, 2N×N, N×2N, N×N, or similar symmetric PU sizes for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitions for PU sizes of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter prediction.
[0203]
[0170] In an example where the mode selection unit 202 does not further divide the CU into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As described above, the size of the CU may refer to the size of the luma coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2N×2N, 2N×N, or N×2N.
[0204]
[0171] As some examples, in the case of other video coding techniques such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, the mode selection unit 202 generates a prediction block for the current block being coded via respective units related to the coding technique. In some examples such as palette mode coding, the mode selection unit 202 may not generate a prediction block, and instead, may generate a syntax element indicating the manner in which the block should be reconstructed based on the selected palette. In such a mode, the mode selection unit 202 may provide these syntax elements to the entropy coding unit 220 for encoding.
[0205]
[0172] As described above, the residual generation unit 204 receives the video data for the current block and the corresponding prediction block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0206]
[0173] The transformation processing unit 206 applies one or more transformations to the residual block to generate a block of transformation coefficients (referred to herein as the "transformation coefficient block"). The transformation processing unit 206 may apply various transformations to the residual block to form the transformation coefficient block. For example, the transformation processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen - Loève transform (KLT), or a conceptually similar transform to the residual block. In some examples, the transformation processing unit 206 may perform multiple transformations on the residual block, such as linear and quadratic transformations like a rotation transform. In some examples, the transformation processing unit 206 does not apply a transformation to the residual block.
[0207]
[0174] The quantization unit 208 may quantize the transformation coefficients in the transformation coefficient block to generate a quantized transformation coefficient block. The quantization unit 208 may quantize the transformation coefficients of the transformation coefficient block according to the quantization parameter (QP) value associated with the current block. The video encoder 200 may adjust the degree of quantization applied to the transformation coefficient block associated with the current block by adjusting the QP value associated with the CU (e.g., via the mode selection unit 202). Quantization can result in loss of information, and thus the quantized transformation coefficients may have lower precision than the original transformation coefficients generated by the transformation processing unit 206.
[0208]
[0175] The inverse quantization unit 210 and the inverse transformation processing unit 212 may apply inverse quantization and inverse transformation, respectively, to the quantized transformation coefficient block to reconstruct the residual block from the transformation coefficient block. The reconstruction unit 214 may generate a reconstructed block corresponding to the current block (potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add the samples of the reconstructed residual block to the corresponding samples from the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0209]
[0176] The filter unit 216 may perform one or more filter operations on the reconstructed block. For example, the filter unit 216 may perform a deblocking operation to reduce blockiness artifacts along the edges of the CU. The operation of the filter unit 216 may be skipped in some examples.
[0210]
[0177] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 may store the reconstructed block in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 may store the filtered reconstructed block in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve a reference picture formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to inter-predict blocks of a picture to be encoded later. Additionally, the intra prediction unit 226 may use the reconstructed blocks in the DPB 218 of the current picture to intra-predict other blocks in the current picture.
[0211]
[0178] Generally, the entropy encoding unit 220 may entropy-encode syntax elements received from other functional components of the video encoder 200. For example, the entropy encoding unit 220 may entropy-encode the quantized transform coefficient blocks from the quantization unit 208. As another example, the entropy encoding unit 220 may entropy-encode prediction syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from the mode selection unit 202. The entropy encoding unit 220 may perform one or more entropy encoding operations on syntax elements, which are another example of video data, to generate entropy-encoded data. For example, the entropy encoding unit 220 may perform context-adaptive variable-length coding (CAVLC) operations, CABAC operations, variable-to-variable (V2V) length coding operations, syntax-based context-adaptive binary arithmetic coding (SBAC) operations, probability interval partitioning entropy (PIPE) coding operations, exponential Golomb coding operations, or another type of entropy encoding operation on the data. In some examples, the entropy encoding unit 220 may operate in a bypass mode where the syntax elements are not entropy-encoded.
[0212]
[0179] The video encoder 200 may output a bitstream including entropy-encoded syntax elements required to reconstruct blocks of a slice or a picture. In particular, the entropy encoding unit 220 may output the bitstream.
[0213]
[0180] Regarding the operations described above, an explanation will be given in terms of blocks. Such an explanation should be understood as being for operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and chroma coding blocks are the luma and chroma components of the CU. In some examples, the luma coding blocks and chroma coding blocks are the luma and chroma components of the PU.
[0214]
[0181] In some examples, the operations performed on the luma coding blocks need not be repeated for the chroma coding blocks. As an example, the operation of identifying the motion vector (MV) and the reference picture for the luma coding blocks need not be repeated to identify the MV and the reference picture for the chroma blocks. Rather, the MV for the luma coding blocks can be scaled to determine the MV for the chroma blocks, and the reference picture can be the same. As another example, the intra prediction process can be the same for the luma coding blocks and the chroma coding blocks.
[0215]
[0182] FIG. 4 is a block diagram showing an exemplary video decoder 300 that can execute the techniques of the present disclosure. FIG. 4 is provided for illustrative purposes and does not limit the techniques widely exemplified and described in the present disclosure. For the purpose of explanation, the present disclosure describes that the video decoder 300 will be described according to the techniques of VVC and HEVC. However, the techniques of the present disclosure can be executed by video coding devices configured according to other video coding standards.
[0216]
[0183] In the example of FIG. 4, the video decoder 300 includes a coded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or in a processing circuit. Moreover, the video decoder 300 may include additional or alternative processors or processing circuits for performing these and other functions.
[0217]
[0184] The prediction processing unit 304 includes a motion compensation unit 316, an intra prediction unit 318, and a gradient-based prediction refinement (GBPR) unit 319. The prediction processing unit 304 may include additional units for performing predictions according to other prediction modes. By way of example, the prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, and the like. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0218]
[0185] Although the GBPR unit 319 is shown as being separate from the motion compensation unit 316, in some examples, the GBPR unit 319 may be part of the motion compensation unit 316. The GBPR unit 319 is shown separately from the motion compensation unit 316 for ease of understanding and should not be considered limiting.
[0219]
[0186] The CPB memory 320 can store video data such as an encoded video bitstream to be decoded by components of the video decoder 300. The video data stored in the CPB memory 320 can be obtained, for example, from the computer-readable medium 110 (FIG. 1). The CPB memory 320 can include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Also, the CPB memory 320 can store video data other than syntax elements of coded pictures, such as temporary data representing outputs from various units of the video decoder 300. The DPB 314 generally stores decoded pictures that can be output and / or used as reference video data when the video decoder 300 decodes subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 can be formed by any of various memory devices, such as dynamic random access memory (DRAM) including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The CPB memory 320 and the DPB 314 can be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 can be on-chip with other components of the video decoder 300 or off-chip with respect to those components.
[0220]
[0187] Additionally or alternatively, in some examples, the video decoder 300 can retrieve coded video data from the memory 120 (FIG. 1). That is, the memory 120 can store the data discussed above using the CPB memory 320. Similarly, the memory 120 can store instructions to be executed by the video decoder 300 when some or all of the functionality of the video decoder 300 is implemented in software because it is executed by the processing circuitry of the video decoder 300.
[0221]
[0188] The various units shown in FIG. 4 are illustrated to assist in understanding the operations performed by video decoder 300. The units may be implemented as fixed function circuitry, programmable circuitry, or a combination thereof. Similar to FIG. 3, fixed function circuitry refers to circuitry that provides a specific function and is preset to the operations that can be performed. Programmable circuitry refers to circuitry that is programmed to perform various tasks and to provide a flexible function in the operations that can be performed. For example, programmable circuitry may execute software or firmware that operates the programmable circuitry in a manner defined by software or firmware instructions. Fixed function circuitry may execute software instructions (e.g., to receive or output parameters), but the type of operations performed by fixed function circuitry is generally immutable. In some examples, one or more of the units may be separate circuit blocks (fixed function or programmable), and in some examples, one or more units may be an integrated circuit.
[0222]
[0189] Video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or programmable core formed from programmable circuitry. In an example where the operations of video decoder 300 are performed by software executed on programmable circuitry, on-chip or off-chip memory may store the software instructions (e.g., object code) that video decoder 300 receives and executes.
[0223]
[0190] Entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to reproduce the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 may generate decoded video data based on the syntax elements extracted from the bitstream.
[0224]
[0191] Generally, video decoder 300 reconstructs pictures block by block. Video decoder 300 may perform a reconstruction operation individually for each block (where the block currently being reconstructed, i.e., the block currently being decoded, may be referred to as the "current block").
[0225]
[0192] Entropy decoding unit 302 may entropy decode syntax elements defining the quantized transform coefficients of the quantized transform coefficient block as well as transform information such as quantization parameter (QP) and / or transform mode indication. Inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the degree of quantization and, similarly, to determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 may perform, for example, a bitwise left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.
[0226]
[0193] After inverse quantization unit 306 forms the transform coefficient block, inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, inverse transform processing unit 308 may apply an inverse DCT, inverse integer transform, inverse Karunen-Loeve transform (KLT), inverse rotation transform, inverse directional transform, or another inverse transform to the coefficient block.
[0227]
[0194] Further, the prediction processing unit 304 generates a prediction block according to the prediction information syntax element entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is to be inter-predicted, the motion compensation unit 316 may generate the prediction block. In this case, the prediction information syntax element may indicate a motion vector that specifies the reference picture in the DPB 314 from which the reference block is taken, and the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 may generally perform the inter-prediction process in a manner substantially similar to that described with respect to the motion compensation unit 224 (FIG. 3).
[0228]
[0195] As another example, if the prediction information syntax element indicates that the current block is to be intra-predicted, the intra-prediction unit 318 may generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Also in this case, the intra-prediction unit 318 may generally perform the intra-prediction process in a manner substantially the same as the manner described with respect to the intra-prediction unit 226 (FIG. 3). The intra-prediction unit 318 may retrieve data of adjacent samples for the current block from the DPB 314.
[0229]
[0196] As another example, if the prediction information syntax element indicates that gradient-based prediction refinement is available, the GBPR unit 319 may modify the samples of the prediction block to generate a modified prediction block (e.g., generate modified samples that form the modified prediction block) for use in reconstructing the current block.
[0230]
[0197] The GBPR unit 319 may be configured to perform the exemplary techniques described in this disclosure for gradient-based prediction improvement. For example, the GBPR unit 319, together with the motion compensation unit 316, may determine a prediction block for inter-predicting a current block (e.g., based on the motion vectors determined by the prediction processing unit 304). The GBPR unit 319 may determine a horizontal displacement and a vertical displacement (e.g., Δv x and Δv y and) for gradient-based prediction improvement of one or more samples of the prediction block. As an example, the GBPR unit 319 may determine an inter-prediction mode for inter-predicting the current block based on prediction information syntax elements. In some examples, the GBPR unit 319 may determine the horizontal displacement and the vertical displacement based on the determined inter-prediction mode.
[0231]
[0198] The GBPR unit 319 may round the horizontal displacement and the vertical displacement to the same level of accuracy for different inter-prediction modes. For example, the current block may be a first current block, the prediction block may be a first prediction block, the horizontal displacement and the vertical displacement may be a first horizontal displacement and vertical displacement, and the rounded horizontal displacement and vertical displacement may be a first rounded horizontal displacement and vertical displacement. In some examples, the GBPR unit 319 may determine a second prediction block for inter-predicting a second current block and determine a second horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block. The GBPR unit 319 may round the second horizontal displacement and the vertical displacement to the same level of accuracy to which the first horizontal displacement and the vertical displacement were rounded to generate a second rounded horizontal displacement and vertical displacement.
[0232]
[0199] In some cases, the inter prediction mode for inter predicting the first current block may be different from the inter prediction mode for the second current block. For example, the first mode of different inter prediction modes is the affine mode, and the second mode of different inter prediction modes is the bi-directional optical flow (BDOF) mode.
[0233]
[0200] The accuracy level at which the horizontal displacement and the vertical displacement are rounded is predefined and can be stored for use by the GBPR unit 319, or the GBPR unit 319 can receive information indicating the accuracy level in the signaled information (for example, the accuracy level is signaled). As an example, the accuracy level is 1 / 64.
[0234]
[0201] The GBPR unit 319 can determine one or more refined offsets based on the rounded horizontal displacement and vertical displacement. For example, the GBPR unit 319 determines a first gradient based on a first set of samples of one or more samples of the prediction block (for example, using the samples of the prediction block described above to determine g x (i,j)), and can determine a second gradient based on a second set of samples of one or more samples of the prediction block (for example, using the samples of the prediction block described above to determine g y (i,j)). The GBPR unit 319 can determine one or more refined offsets based on the rounded horizontal displacement and vertical displacement, the first gradient, and the second gradient. In some examples, if the value of one or more refined offsets is too high (for example, greater than a threshold), the GBPR unit 319 can clip one or more refined offsets.
[0235]
[0202] The GBPR unit 319 may modify one or more samples of the prediction block based on one or more determined improved offsets or one or more clipped improved offsets to generate a modified prediction block (e.g., one or more modified samples that form the modified prediction block). For example, the GBPR unit 319 may determine g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among one or more samples located at (i,j), and Δv x (i,j) is the rounded horizontal displacement for the sample among one or more samples located at (i,j), and g y (i,j) is the second gradient for the sample among one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among one or more samples located at (i,j). In some examples, Δv x and Δv y may be the same for each sample (i,j) of the prediction block.
[0236]
[0203] The obtained modified samples may form a modified prediction block during gradient-based prediction refinement. That is, the modified prediction block may be used as the prediction block during gradient-based prediction refinement. The reconstruction unit 310 may reconstruct the current block using the prediction block and the residual block. For example, the reconstruction unit 310 may add the samples of the residual block to the corresponding samples of the prediction block to reconstruct the current block.
[0237]
[0204] The filter unit 312 may perform one or more filtering operations on the reconstructed block. For example, the filter unit 312 may perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed block. The operations of the filter unit 312 are not necessarily performed in all cases.
[0238]
[0205] The video decoder 300 may store the reconstructed blocks in the DPB 314. As discussed above, the DPB 314 may provide reference information, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation, to the prediction processing unit 304. Further, the video decoder 300 may output the decoded pictures from the DPB for subsequent presentation on a display device, such as the display device 118 of FIG. 1.
[0239]
[0206] According to a first technique of the present disclosure, a video coder (e.g., the video encoder 200 and / or the video decoder 300) may derive the differences between the x and y components of the motion vector calculated at the location pixel location (m, n) based on the sub-block MV and the sub-block MV (i.e., Δv x (m,n) and Δv y (m,n)). For example, if the affine motion parameters a, b, c, d, e, and f in the derivation of Δv x (m,n) and Δv y (m,n) are calculated from the CPMVs (control point motion vectors), the CPMVs of each block may need to be stored in the motion buffer. Since the CPMV has three MVs for each prediction direction instead of one MV as in the normal inter mode, this storage of the CPMVs of each block can significantly increase the buffer size. Therefore, the present disclosure describes that the video coder performs the derivation of Δv x (m,n) and Δv y (m,n) based on the sub-block MV.
[0240]
[0207] In the case of a 6-parameter affine model, three different sub-block MVs that are not all in the same row or column of the same sub-block can be selected. In a 4-parameter affine model, two different sub-block MVs are selected. In some examples, the selected sub-block MVs are similar to the CPMV described above
[0241]
Number
[0242] , can be used as i = 0, 1, 2, where
[0243]
Number
[0244] and
[0245]
Number
[0246] are in the same column of the same sub-block,
[0247]
Number
[0248] and
[0249]
Number
[0250] are in the same row of the same sub-block. Then, the parameter a in
[0251]
Number
[0252] is calculated as, and parameter b is,
[0253] [Number]
[0254] is calculated as, and parameter c is,
[0255] [Number]
[0256] is calculated as, and parameter d is,
[0257] [Number]
[0258] is calculated. In the case of the four-parameter affine mode, parameter a is,
[0259] [Number]
[0260] is calculated as, and parameter c is,
[0261] [Number]
[0262] is calculated as, parameter b is set equal to -c, and parameter d is set equal to a. W is,
[0263] [Number]
[0264] and
[0265]
Number
[0266] is the distance between and, and H is
[0267]
Number
[0268] and
[0269]
Number
[0270] is the distance between and. However, in some examples, three sub-block MVs are selected regardless of whether a 6-parameter affine model or a 4-parameter affine model is used.
[0271]
[0208] The video coder selects sub-block MVs such that W is equal to blkW / 2 and H is equal to blkH / 2. In one example, as shown in FIG. 6,
[0272]
Number
[0273] is the sub-block MV of the upper left sub-block at location (0,0),
[0274]
Number
[0275] is the sub-block MV of the upper middle sub-block at location (blkW / 2,0),
[0276] [Number]
[0277] is the sub-block MV of the sub-block MV of the sub-block at the center left of the location (0, blkH / 2). In another example,
[0278] [Number]
[0279] is the sub-block MV of the sub-block at the center top of the location (blkW / 2 - sbW, 0),
[0280] [Number]
[0281] is the sub-block MV of the sub-block at the upper right of the location (blkW - sbW, 0),
[0282] [Number]
[0283] is the sub-block MV of the sub-block at the center center of the location (blkW / 2 - sbW, blkH / 2).
[0284]
[0209] According to the second technique of the present disclosure, the video coder may perform clipping of Δv x (m,n) and Δv y (m,n). The calculation of the gradient-based improved offset may assume that Δv x (m,n) and Δv y (m,n) are small. In this technique, the video coder clips Δv x (m,n) and Δv y so that their absolute values are less than or equal to a predefined threshold ΔTH.(m, n) can be clipped.
[0285]
[0210] As an example, a predefined threshold is Δv in offset calculation x (m, n) / Δv y (m, n) and gradient Δg x (m, n) / Δg y (m, n) and the multiplication between them can be set so as not to cause buffer overflow. For example, if the allocation for the multiplication result is 16 bits, the maximum absolute value is 1 << 15 (1 bit for the sign), and ΔTH * g x (m, n) or ΔTH * g y (m, n) must not exceed 1 << 15. If the gradient is represented by k bits, ΔTH is set equal to 1 << (15 - k).
[0286]
[0211] As another example, a predefined threshold (e.g., ΔTH) can be set equal to the same value as in the case of bidirectional optical flow (BDOF), i.e., th’ BIO and be set equal to. th’ BIO can represent half a pixel. Δv x (m, n) and Δv y If the basic unit for (m, n) is 1 / q pixels, th’ BIO is q / 2.
[0287]
[0212] As another example, a predefined threshold (e.g., ΔTH) can be set equal to the minimum value between th’ BIO and 1 << (15 - k).
[0288]
[0213] According to the third technique of the present disclosure, the video coder can set the accuracy of Δv x (m, n) and Δv y (m, n) to the same accuracy as in the case of BDOF. In one example, the accuracy of Δv x (m, n) and Δv y (m, n) is determined by shfit1 in Section 1.3. Therefore, Δv x (m, n) or Δvx (m,n) one unit is 1 / (1<<shift1) pixels. In one example, shift is set equal to 6. In another example, shift1 is set equal to max(2,14-bitDepth), where bitDepth is the internal bit depth of the video signal for encoding / decoding.
[0289]
[0214] According to the fourth technique of the present disclosure, the video coder may perform gradient calculation for affine mode prediction improvement using the same process as in the case of BDOF. Therefore, the same module of the video coder may be used for both BDOF calculation and gradient calculation for affine mode prediction improvement. However, the video coder may use different padding methods for prediction samples in the extended area.
[0290]
[0215] As an example, the video coder may generate prediction samples in the extended area (white positions) by directly taking reference samples at nearby integer positions (by using the floor() operation on the coordinates) without interpolation.
[0291]
[0216] As another example, the video coder may generate prediction samples in the extended area (white positions) by directly taking reference samples at the nearest integer positions (by using the round() operation on the coordinates) without interpolation.
[0292]
[0217] As another example, when any sample values outside the sub-block boundary are required, the video coder may pad (i.e., repeat) the required samples from their nearest neighbors. This may also be applied to the gradient calculation of BDOF.
[0293] According to the fifth technique of the present disclosure, the video coder may perform clipping of the improvement result. In inter prediction, the motion compensation prediction signal of a block is usually clipped to the same range as the original signal of the block. However, in bidirectional motion compensation, the motion compensation prediction signals for each direction are kept at intermediate precision and range to improve accuracy. After the weighted averaging process of bidirectional motion compensation, the result is rounded and clipped to the same range and precision as the original signal of the block. In this fifth technique, in the case of bidirectional prediction, the video coder may clip the result of prediction improvement to have the same intermediate precision and range as in the case of normal motion compensation. For example, the number of bits for intermediate precision is 14, in which case the video coder may clip the result of prediction improvement to the range from -(1<<14) to (1<<14).
[0294]
[0219] FIG. 7 is a flowchart showing an exemplary method for coding video data. The current block may include the current CU. For the example of FIG. 7, the processing circuit will be described. Examples of the processing circuit include fixed function and / or programmable circuits for video encoder 200 such as GBPR unit 227 and video decoder 300 such as GBPR unit 319.
[0295]
[0220] In one or more examples, the memory may be configured to store samples of the predicted block. For example, DPB 218 or DPB 314 may be configured to store samples of the predicted block used for inter prediction. A copy of an intra block may be regarded as an exemplary inter prediction mode, in which case the block vector used for the copy of the intra block is an example of a motion vector.
[0296]
[0221] The processing circuit may determine (350) a predicted block stored in the memory for inter predicting the current block. The processing circuit may perform horizontal and vertical displacements for gradient-based prediction improvement of one or more samples of the predicted block (e.g., Δv x and Δvy (352) can determine the horizontal displacement and vertical displacement. As an example, the processing circuit can determine an inter prediction mode for inter predicting the current block. In some examples, the processing circuit can determine a horizontal displacement and a vertical displacement based on the determined inter prediction mode.
[0297]
[0222] The processing circuit can round the horizontal displacement and the vertical displacement to the same accuracy level for different inter prediction modes (354). For example, the current block can be the first current block, the predicted block can be the first predicted block, the horizontal displacement and the vertical displacement can be the first horizontal displacement and vertical displacement, and the rounded horizontal displacement and vertical displacement can be the first rounded horizontal displacement and vertical displacement. In some examples, the processing circuit can determine a second predicted block for inter predicting the second current block and determine a second horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the second predicted block. The processing circuit can round the second horizontal displacement and the vertical displacement to the same accuracy level at which the first horizontal displacement and the vertical displacement are rounded to generate the second rounded horizontal displacement and the vertical displacement.
[0298]
[0223] In some cases, the inter prediction mode for inter predicting the first current block can be different from the inter prediction mode for the second current block. For example, the first mode of the different inter prediction modes is the affine mode, and the second mode of the different inter prediction modes is the bidirectional optical flow (BDOF) mode.
[0299]
[0224] The accuracy level to which the horizontal displacement and the vertical displacement are rounded can be predefined or signaled. As an example, the accuracy level is 1 / 64.
[0300]
[0225] The processing circuit may determine one or more refined offsets based on the rounded horizontal displacement and vertical displacement (356). For example, the processing circuit may determine a first gradient based on a first set of samples among one or more samples of a prediction block (e.g., determining g x (i,j) using the samples of the prediction block described above), and may determine a second gradient based on a second set of samples among one or more samples of the prediction block (e.g., determining g y (i,j) using the samples of the prediction block described above). The processing circuit may determine one or more refined offsets based on the rounded horizontal displacement and vertical displacement, and the first and second gradients. In some examples, if the value of one or more refined offsets is too high (e.g., greater than a threshold), the processing circuit may clip the one or more refined offsets.
[0301]
[0226] The processing circuit may modify one or more samples of the prediction block based on the determined one or more refined offsets or the clipped one or more refined offsets to generate a modified prediction block (e.g., one or more modified samples that form the modified prediction block) (358). For example, the processing circuit may determine g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among one or more samples located at (i,j), and Δv y(i,j) is the rounded vertical displacement for a sample among one or more samples located at (i,j). In some examples, Δv x and Δv y can be the same for each sample (i,j) of the prediction block.
[0302]
[0227] The processing circuit may code (e.g., encode or decode) the current block based on the modified prediction block (e.g., one or more modified samples of the modified prediction block) (360). For example, in video decoding, the processing circuit (e.g., video decoder 300) may reconstruct the current block based on the modified prediction block (e.g., by adding one or more modified samples to the received residual values). In video encoding, the processing circuit (e.g., video encoder 200) may determine the residual value (e.g., of the residual block) between the current block and the modified prediction block (e.g., one or more modified samples of the modified prediction block), and may signal information indicating the residual value.
[0303]
[0228] A non-limiting exemplary list of examples of the present disclosure is described below.
[0304]
[0229] Example 1. A method of coding video data, comprising performing sub-block based affine motion compensation to obtain a prediction signal for a current block of the video data, and improving the prediction signal by adding an offset at least to pixel locations within the prediction signal, wherein the value of the offset is derived based on values of a plurality of sub-block motion vectors (MVs) for sub-blocks of the current block that are not all in the same row or column of the same sub-block of the current block. A method comprising.
[0305]
[0230] Example 2. Δv x (m,n) and Δv yThe method according to Example 1, further comprising determining an offset value at the location (m, n) of the prediction signal based on (m, n).
[0306]
[0231] Example 3. Determining an offset value at the location (m, n) of the prediction signal comprises determining the offset value according to the following formula:
[0307]
Equation
[0308] where Δg x (m, n) is the horizontal gradient of the prediction signal, and g y (m, n) is the vertical gradient of the prediction signal, the method according to Example 2.
[0309]
[0232] Example 4. Further comprising determining the values of Δv x (m, n) and Δv y (m, n) according to the method according to either Example 2 or 3.
[0310]
[0233] Example 5. Determining the values of Δv x (m, n) and Δv y (m, n) comprises determining the values of Δv x (m, n) and Δv y (m, n) according to the method according to Example 4.
[0311]
[0234] Example 6. Deriving Δv x (m, n) and Δv y (m, n) comprises deriving Δv x (m, n) and Δv y (m, n) according to the following formula:
[0312]
Equation
[0313] Here, sbW represents the width of the sub-sub-block of the current block, sbH represents the height of the sub-sub-block of the current block, and (m, n) represents the pixel location within the current block, as described in Example 5.
[0314]
[0235] Example 7. Deriving the a parameter, the b parameter, the c parameter, and the d parameter according to the following formula, where the affine motion model is represented by six parameters,
[0315]
Equation
[0316] Here, W represents the distance between the first MV of the affine motion model and the second MV of the affine motion model, and H represents the distance between the first MV of the affine motion model and the third MV of the affine motion model. The method according to any one of Example 5 or 6, further comprising
[0317]
[0236] Example 8. Deriving the a parameter, the b parameter, the c parameter, and the d parameter according to the following formula, where the affine motion model is represented by four parameters,
[0318]
Equation
[0319] Here, W represents the distance between the first MV of the affine motion model and the second MV of the affine motion model. The method according to any one of Example 5 to 7, further comprising
[0320]
[0237] Example 9. The first MV of the affine motion model is
[0321] [Number]
[0322] and the second MV of the affine motion model is
[0323] [Number]
[0324] and the third MV of the affine motion model is
[0325] [Number]
[0326] The method according to Example 7 or Example 8, which is
[0327]
[0238] Example 10. Δv having an absolute value less than or equal to a predefined threshold x (m,n) and Δv y (m,n) is further clipped, and the method according to any one of Examples 2 to 9.
[0328]
[0239] Example 11. The method according to any one of Examples 2 to 10, further comprising performing a bidirectional optical flow (BDOF) improvement on the prediction signal for the current block.
[0329]
[0240] Example 12. Δv having the same accuracy used for performing the BDOF improvement x (m,n) and Δv y (m,n) is further memorized, and the method according to Example 11.
[0330]
[0241] Example 13. Performing the BDOF improvement comprises performing a gradient calculation, and the method according to either Example 11 or Example 12.
[0331]
[0242] Example 14. The method according to Example 13, wherein performing gradient calculation to perform BDOF improvement uses the same process as calculating the horizontal and / or vertical gradients of the prediction signal.
[0332]
[0243] Example 15. The method according to any one of Examples 1 to 14, further comprising clipping the prediction signal improved to have the same intermediate accuracy as in the case of non-affine motion compensation.
[0333]
[0244] Example 16. The method according to Example 15, wherein when the number of bits for intermediate accuracy is n, clipping the improved prediction signal comprises clipping the improved prediction signal in the range from -(1<<n) to (1<<n).
[0334]
[0245] Example 17. The method according to any one of Examples 1 to 16, wherein coding comprises decoding.
[0335]
[0246] Example 18. The method according to any one of Examples 1 to 17, wherein coding comprises encoding.
[0336]
[0247] Example 19. A device for coding video data, comprising one or more means for performing the method according to any one of Examples 1 to 18.
[0337]
[0248] Example 20. The device according to Example 19, wherein the one or more means comprise one or more processors implemented in a circuit.
[0338]
[0249] Example 21. The device according to any one of Examples 19 and 20, further comprising a memory for storing video data.
[0339]
[0250] Device according to any of Examples 19 - 21, further comprising a display configured to display the decoded video data.
[0340]
[0251] Example 23. A device according to any of Examples 19 - 22, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set - top box.
[0341]
[0252] Example 24. A device according to any of Examples 19 - 23, wherein the device comprises a video decoder.
[0342]
[0253] Example 25. A device according to any of Examples 19 - 24, wherein the device comprises a video encoder.
[0343]
[0254] Example 26. A computer - readable storage medium storing instructions that, when executed, cause one or more processors to execute the method according to any of Examples 1 - 18.
[0344]
[0255] It should be recognized that, depending on the example, some of the acts or events of any of the techniques described herein may be executed in a different sequence, may be added, merged, or completely excluded (e.g., not all of the acts or events described are necessarily required for the practice of the technique). Further, in some examples, the acts or events may not be executed sequentially, but may be executed simultaneously, for example, through multi - threading, interrupt processing, or multiple processors.
[0345] In one or more examples, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or may include a communication medium including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this way, the computer-readable medium generally may correspond to (1) a non-transitory tangible computer-readable storage media, or (2) a communication medium such as a signal or carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0346] By way of example and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM (registered trademark), CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection can be properly called a computer-readable medium. For example, if the instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage media and data storage media are directed to non-transitory, tangible storage media and do not include connections, carrier waves, signals, or other transitory media. As used herein, disk and disc include compact disc (CD), laser disc (registered trademark), optical disc, digital versatile disc (DVD), floppy (registered trademark) disk, and Blu-ray (registered trademark) disc, where disk typically magnetically reproduces data and disc optically reproduces data with a laser. Combinations of the above should also be included within the scope of computer-readable media.
[0347]
[0258] The commands can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Thus, the terms “processor” and “processing circuitry” as used herein can refer to either the foregoing structures, or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a combined codec. Also, the techniques can be fully implemented in one or more circuits or logic elements.
[0348]
[0259] The techniques of the present disclosure can be implemented in a variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chip sets). In the present disclosure, various components, modules, or units are described so as to emphasize the functional aspects of a device configured to execute the disclosed techniques, but implementation by different hardware units is not necessarily required. Instead, as described above, the various units can be combined in a codec hardware unit, including one or more of the processors described above, along with suitable software and / or firmware, or provided by a set of interoperable hardware units.
[0349]
[0260] Various examples have been described. These and other examples fall within the scope of the following claims. The invention described in the claims of the present application at the time of initial filing is appended below. [C1] A method for decoding video data, comprising: determining a prediction block for inter-predicting a current block; determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a bidirectional optical flow (BDOF) mode; determining one or more refined offsets based on the rounded horizontal displacement and vertical displacement; modifying the one or more samples of the prediction block based on the determined one or more refined offsets to generate a modified prediction block, and reconstructing the current block based on the modified prediction block. A method comprising the above. [C2] Clipping the one or more refined offsets, wherein modifying the one or more samples of the prediction block comprises modifying the one or more samples of the prediction block based on the clipped one or more refined offsets. The method according to C1, further comprising the above. [C3] Determining an inter-prediction mode for inter-predicting the current block, wherein determining the horizontal displacement and the vertical displacement comprises determining the horizontal displacement and the vertical displacement based on the determined inter-prediction mode. The method according to C1, further comprising the above. [C4] The method according to C1, wherein the accuracy level is 1 / 64. [C5] Determining a first gradient based on a first set of samples among the one or more samples of the prediction block; determining a second gradient based on a second set of samples among the one or more samples of the prediction block; wherein determining the one or more refined offsets comprises determining the one or more refined offsets based on the rounded horizontal displacement and vertical displacement, the first gradient, and the second gradient. The method according to C1, further comprising the above. [C6] Determining the one or more improved offsets comprises determining g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), and Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), and g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j), the method according to C5. [C7] The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a vertical displacement, the one or more improved offsets are a first one or more improved offsets, the rounded horizontal displacement and the vertical displacement are a first rounded horizontal displacement and a vertical displacement, the corrected prediction block is a first corrected prediction block, and the method comprises determining a second prediction block for inter-predicting a second current block; determining a second horizontal displacement and a second vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block; rounding the second horizontal displacement and the second vertical displacement to the same accuracy level at which the first horizontal displacement and the vertical displacement were rounded to generate a second rounded horizontal displacement and a second vertical displacement, and determining a second one or more improved offsets based on the second rounded horizontal displacement and the second vertical displacement; correcting the one or more samples of the second prediction block based on the determined second one or more improved offsets to generate a second corrected prediction block; reconstructing the second current block based on the second corrected prediction block and further comprising the method according to C1. [C8] A method for encoding video data, comprising determining a prediction block for inter-predicting a current block; Determining horizontal and vertical displacements for gradient-based prediction improvement of one or more samples of the prediction block; Rounding the horizontal and vertical displacements to the same level of accuracy that is the same for different inter-prediction modes including an affine mode and a bidirectional optical flow (BDOF) mode; Determining one or more improved offsets based on the rounded horizontal and vertical displacements; Modifying the one or more samples of the prediction block based on the determined one or more improved offsets to generate a modified prediction block, and determining a residual value indicating a difference between the current block and the modified prediction block; Signaling information indicating the residual value; A method comprising the above steps. [C9] Clipping the one or more improved offsets, wherein modifying the one or more samples of the prediction block comprises modifying the one or more samples of the prediction block based on the clipped one or more improved offsets. The method according to C8, further comprising the above steps. [C10] Determining an inter-prediction mode for inter-predicting the current block, wherein determining the horizontal and vertical displacements comprises determining the horizontal and vertical displacements based on the determined inter-prediction mode. The method according to C8, further comprising the above steps. [C11] The method according to C8, wherein the level of accuracy is 1 / 64. [C12] Determining a first gradient based on a first set of samples among the one or more samples of the prediction block; Determining a second gradient based on a second set of samples among the one or more samples of the prediction block; wherein determining the one or more improved offsets comprises determining the one or more improved offsets based on the rounded horizontal and vertical displacements, the first gradient, and the second gradient. The method according to C8, further comprising the above steps. [C13] Determining the one or more improved offsets comprises x g x (i,j)*Δv y (i,j)+g y (i,j)*Δv (i,j), where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j), the method according to C12. [C14] The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a vertical displacement, the one or more refined offsets are a first one or more refined offsets, the rounded horizontal displacement and the vertical displacement are a first rounded horizontal displacement and a vertical displacement, the corrected prediction block is a first corrected prediction block, the residual value comprises a first residual value, and the method comprises determining a second prediction block for inter-predicting a second current block; determining a second horizontal displacement and a second vertical displacement for gradient-based prediction refinement of one or more samples of the second prediction block; rounding the second horizontal displacement and the second vertical displacement to the same accuracy level at which the first horizontal displacement and the vertical displacement were rounded to generate a second rounded horizontal displacement and a second rounded vertical displacement, and determining a second one or more refined offsets based on the second rounded horizontal displacement and the second rounded vertical displacement; correcting the one or more samples of the second prediction block based on the determined second one or more refined offsets to generate a second corrected prediction block; determining a second residual value indicating the difference between the second current block and the second corrected prediction block; signaling information indicating the second residual value comprising the method according to C8. [C15] A device for coding video data, comprising a memory configured to store one or more samples of a prediction block; a processing circuit and the processing circuit is Determining the prediction block for inter-predicting the current block; Determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of the one or more samples of the prediction block; Rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a bidirectional optical flow (BDOF) mode; Determining one or more refined offsets based on the rounded horizontal displacement and vertical displacement; Modifying the one or more samples of the prediction block based on the determined one or more refined offsets to generate a modified prediction block; Coding the current block based on the modified prediction block; A device configured to perform the above. [C16] For coding the current block, the processing circuit is the device according to C15, configured to reconstruct the current block based on the modified prediction block. [C17] For coding the current block, the processing circuit Determining a residual value indicating a difference between the current block and the modified prediction block; Signaling information indicating the residual value; The device according to C15, configured to perform the above. [C18] The processing circuit Clipping the one or more refined offsets, where, for modifying the one or more samples of the prediction block, the processing circuit is configured to modify the one or more samples of the prediction block based on the clipped one or more refined offsets; The device according to C15, configured to perform the above. [C19] The processing circuit Determining an inter-prediction mode for inter-predicting the current block, where, for determining the horizontal displacement and the vertical displacement, the processing circuit is configured to determine the horizontal displacement and the vertical displacement based on the determined inter-prediction mode; The device according to C15, configured to perform the above. [C20] The accuracy level is 1 / 64, the device according to C15. [C21] The processing circuit Determining a first gradient based on a first set of samples among the one or more samples of the prediction block; Determining a second gradient based on a second set of samples among the one or more samples of the prediction block; Here, in order to determine the one or more improved offsets, the processing circuit is configured to determine the one or more improved offsets based on the rounded horizontal displacement and vertical displacement, and the first gradient and second gradient; A device according to C15, configured to perform. [C22] In order to determine the one or more improved offsets, the processing circuit is g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) is configured to determine, where g x (i,j) is the first gradient for the sample among the one or more samples located at (i,j), and Δv x (i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), and g y (i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δv y (i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j), a device according to C21. [C23] The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and vertical displacement are a first horizontal displacement and vertical displacement, the one or more improved offsets are a first one or more improved offsets, the rounded horizontal displacement and vertical displacement are a first rounded horizontal displacement and vertical displacement, and the corrected prediction block is a first corrected prediction block. Here, the processing circuit is Determining a second prediction block for inter-predicting a second current block; Determining a second horizontal displacement and vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block; Rounding the second horizontal displacement and vertical displacement to the same accuracy level as the first horizontal displacement and vertical displacement were rounded to generate the second rounded horizontal displacement and vertical displacement, and determining one or more improved offsets based on the second rounded horizontal displacement and vertical displacement; Modifying one or more samples of the second prediction block based on the determined one or more improved offsets to generate a second corrected prediction block; Coding the second current block based on the second corrected prediction block; The device according to C15, configured to perform the above. [C24] The device according to C15, further comprising a display configured to display the decoded video data. [C25] The device according to C15, further comprising a camera configured to capture the video data to be encoded. [C26] The device according to C15, wherein the device comprises one or more of a camera, a computer, a wireless communication device, a broadcast receiver device, or a set-top box. [C27] A computer-readable storage medium storing instructions that, when executed, cause one or more processors to Determine a prediction block for inter-predicting a current block; Determine a horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block; Round the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a bidirectional optical flow (BDOF) mode; Determine one or more improved offsets based on the rounded horizontal displacement and vertical displacement; Modify one or more samples of the prediction block based on the determined one or more improved offsets to generate a corrected prediction block, and code the current block based on the corrected prediction block. A computer-readable storage medium. [C28] A device for coding video data, comprising Means for determining a prediction block for inter-predicting a current block; Means for determining a horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block; Means for rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including an affine mode and a bidirectional optical flow (BDOF) mode; Means for determining one or more improved offsets based on the rounded horizontal displacement and vertical displacement; Means for modifying the one or more samples of the prediction block based on the determined one or more improved offsets to generate a modified prediction block; Means for coding the current block based on the modified prediction block A device comprising the above.
Claims
1. A method for decoding video data, comprising: determining a prediction block for inter-predicting a current block; determining a horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block; determining an inter-prediction mode for inter-predicting the current block; wherein an inter-prediction mode for inter-predicting a first current block is different from an inter-prediction mode for a second current block, a first mode of the different inter-prediction modes is an affine mode, a second mode of the different inter-prediction modes is a bidirectional optical flow (BDOF) mode, and determining the horizontal displacement and the vertical displacement comprises determining the horizontal displacement and the vertical displacement based on the determined inter-prediction mode; rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including the affine mode and the bidirectional optical flow (BDOF) mode; determining a first gradient based on a first set of samples among the one or more samples of the prediction block; determining a second gradient based on a second set of samples among the one or more samples of the prediction block; Determining one or more improved offsets based on the rounded horizontal and vertical displacements, the first gradient, and the second gradient, where determining the one or more improved offsets is g x (i, j)*Δv x (i, j) + g y (i, j) *Δv y comprising determining (i, j), where g x (i, j) is the first gradient for the sample among the one or more samples located at (i, j), and Δv x (i, j) is among the one or more samples located at (i, j) The rounded horizontal displacement for the sample, g y (i, j) is at (i, j) The second gradient for the sample among the one or more samples located, Δv y (i, j) is among the one or more samples located at (i, j) wherein the rounded vertical displacement for the samples; performing gradient-based prediction improvement by modifying the one or more samples of the prediction block based at least on the determined one or more improvement offsets to generate a modified prediction block; reconstructing the current block based on the modified prediction block The method comprising.
2. clipping the one or more improvement offsets, wherein modifying the one or more samples of the prediction block comprises modifying the one or more samples of the prediction block based on the clipped one or more improvement offsets; The method according to claim 1, further comprising.
3. The method according to claim 1, wherein the accuracy level is 1 / 64.
4. The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a vertical displacement, the one or more improved offsets are a first one or more improved offsets, the rounded horizontal displacement and the vertical displacement are a first rounded horizontal displacement and a vertical displacement, the corrected prediction block is a first corrected prediction block, and the method determining a second prediction block for inter-predicting a second current block; determining a second horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block; rounding the second horizontal displacement and the vertical displacement to the same accuracy level to which the first horizontal displacement and the vertical displacement were rounded to generate a second rounded horizontal displacement and a vertical displacement; determining a second one or more improved offsets based on the second rounded horizontal displacement and the vertical displacement; correcting the one or more samples of the second prediction block based on the determined second one or more improved offsets to generate a second corrected prediction block; reconstructing the second current block based on the second corrected prediction block The method according to claim 1, further comprising. **Claim 5** A method for encoding video data, comprising: determining a prediction block for inter-predicting a current block; determining a horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block; determining an inter-prediction mode for inter-predicting the current block; wherein an inter-prediction mode for inter-predicting a first current block is different from an inter-prediction mode for a second current block, a first mode of the different inter-prediction modes is an affine mode, a second mode of the different inter-prediction modes is a bidirectional optical flow (BDOF) mode, and determining the horizontal displacement and the vertical displacement comprises determining the horizontal displacement and the vertical displacement based on the determined inter-prediction mode. Rounding the horizontal displacement and the vertical displacement to the same accuracy level that is the same for different inter-prediction modes including the affine mode and the bidirectional optical flow (BDOF) mode, Determining a first gradient based on a first set of samples among the one or more samples of the prediction block, Determining a second gradient based on a second set of samples among the one or more samples of the prediction block, Determining one or more improved offsets based on the rounded horizontal displacement, the rounded vertical displacement, the first gradient, and the second gradient, wherein determining the one or more improved offsets comprises g x (i,j)*Δv x (i,j)+g y (i,j) *Δv y comprising determining (i, j), where g x (i, j) is the first gradient for the sample among the one or more samples located at (i, j), and Δv x (i, j) is among the one or more samples located at (i, j) The rounded horizontal displacement for the sample, g y (i, j) is at (i, j) The second gradient for the sample among the one or more samples located, Δv y (i, j) is among the one or more samples located at (i, j) which is the rounded vertical displacement for the samples, Performing gradient-based prediction improvement by modifying the one or more samples of the prediction block based at least on the determined one or more improved offsets to generate a modified prediction block, Determining a residual value indicating a difference between the current block and the modified prediction block, Signaling information indicating the residual value A method comprising. **Claim 6** Clipping the one or more improved offsets, wherein modifying the one or more samples of the prediction block comprises modifying the one or more samples of the prediction block based on the clipped one or more improved offsets, The method according to claim 5, further comprising. **Claim 7** The method according to claim 5, wherein the accuracy level is 1 / 64. **Claim 8** The prediction block is a first prediction block, the current block is a first current block, the horizontal displacement and the vertical displacement are a first horizontal displacement and a vertical displacement, the one or more improved offsets are a first one or more improved offsets, the rounded horizontal displacement and the vertical displacement are a first rounded horizontal displacement and a vertical displacement, the modified prediction block is a first modified prediction block, the residual value comprises a first residual value, and the method comprises Determining a second prediction block for inter-predicting a second current block, Determining a second horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the second prediction block, Rounding the second horizontal displacement and the vertical displacement to the same accuracy level to which the first horizontal displacement and the vertical displacement were rounded to generate a second rounded horizontal displacement and a vertical displacement, Determining one or more refined offsets of a second based on the second rounded horizontal displacement and vertical displacement; Modifying one or more samples of the second prediction block based on the determined one or more refined offsets of the second to generate a second modified prediction block; Determining a second residual value indicative of a difference between the second current block and the second modified prediction block; Signaling information indicative of the second residual value; The method according to claim 5, further comprising. **Claim 9** A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform the method according to any one of claims 1 to 8. **Claim 10** A device for decoding video data, means for determining a prediction block for inter-predicting a current block; means for determining a horizontal displacement and a vertical displacement for gradient-based prediction refinement of one or more samples of the prediction block; means for determining an inter-prediction mode for inter-predicting the current block; wherein an inter-prediction mode for inter-predicting a first current block is different from an inter-prediction mode for a second current block, a first mode of the different inter-prediction modes is an affine mode, a second mode of the different inter-prediction modes is a bidirectional optical flow (BDOF) mode, and determining the horizontal displacement and the vertical displacement comprises determining the horizontal displacement and the vertical displacement based on the determined inter-prediction mode; means for rounding the horizontal displacement and the vertical displacement to the same level of accuracy for different inter-prediction modes including the affine mode and the bidirectional optical flow (BDOF) mode; means for determining a first gradient based on a first set of samples of the one or more samples of the prediction block; means for determining a second gradient based on a second set of samples of the one or more samples of the prediction block; Means for determining one or more improved offsets based on the rounded horizontal and vertical displacements, the first gradient, and the second gradient, where the means for determining the one or more improved offsets is g x (i, j) * Δv x (i, j) + g y (i, j) * Δv y Comprising means for determining (i, j), where g x (i, j) is the first gradient for the sample among the one or more samples located at (i, j), Δv x (i, j) is the one or more located at (i, j) is the rounded horizontal displacement for the sample among the samples, g y (i, j ) is the second gradient for the sample among the one or more samples located at (i, j), Δv y (i, j) is the one or more located at (i, j) wherein the rounded vertical displacement for the sample among the number of samples; Means for improving gradient-based prediction, and here, the means for improving gradient-based prediction includes means for modifying one or more samples of the prediction block based on the determined one or more improvement offsets to generate a modified prediction block. A device comprising means for reconstructing the current block based on the modified prediction block.
11. A device for encoding video data, means for determining a prediction block for inter-predicting a current block, means for determining a horizontal displacement and a vertical displacement for gradient-based prediction improvement of one or more samples of the prediction block, means for determining an inter-prediction mode for inter-predicting the current block, wherein the inter-prediction mode for inter-predicting a first current block is different from the inter-prediction mode for a second current block, the first mode of the different inter-prediction modes is an affine mode, the second mode of the different inter-prediction modes is a bidirectional optical flow (BDOF) mode, and determining the horizontal displacement and the vertical displacement comprises determining the horizontal displacement and the vertical displacement based on the determined inter-prediction mode. means for rounding the horizontal displacement and the vertical displacement to the same accuracy level for different inter-prediction modes including the affine mode and the bidirectional optical flow (BDOF) mode, means for determining a first gradient based on a first set of samples among the one or more samples of the prediction block, means for determining a second gradient based on a second set of samples among the one or more samples of the prediction block. means for determining one or more improved offsets based on the rounded horizontal and vertical displacements, the first gradient, and the second gradient, wherein the means for determining the one or more improved offsets comprises means for determining gx(i,j)*Δvx(i,j) + gy(i,j)*Δvy(i,j), where gx(i,j) is the first gradient for the sample among the one or more samples located at (i,j), Δvx(i,j) is the rounded horizontal displacement for the sample among the one or more samples located at (i,j), gy(i,j) is the second gradient for the sample among the one or more samples located at (i,j), and Δvy(i,j) is the rounded vertical displacement for the sample among the one or more samples located at (i,j), means for performing gradient-based prediction improvement, wherein the means for performing gradient-based prediction improvement comprises means for modifying one or more samples of the prediction block based on the determined one or more improved offsets to generate a modified prediction block, means for determining a residual value indicative of the difference between the current block and the modified prediction block means for signaling information indicative of the residual value A device comprising the above.
12. The device according to claim 10, further comprising means for performing the method according to any one of claims 2 to 4.
13. The device according to claim 11, further comprising means for performing the method according to any one of claims 6 to 8.
Citation Information
Patent Citations
intra bc and inter integration
JP2017535163A
Using syntax for affine modes with adaptive motion vector resolution.
JP2022500909A
Syntax reuse for affine mode with adaptive motion vector resolution
WO2020058890A1