Bi-directional optical flow refinement for affine motion compensation
By introducing bidirectional optical flow tools and affine motion models into video coding, and combining spatial gradient and temporal optical flow models for motion compensation, the problem of low motion compensation efficiency in inter-frame prediction is solved, thereby improving coding efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-23
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video coding techniques suffer from low motion compensation efficiency in inter-frame prediction, especially when small motions are observed between intra-block samples, leading to reduced coding efficiency.
Motion refinement is achieved using a bidirectional optical flow (BDOF) tool. Motion compensation is performed on a sampling basis in dual prediction mode using an optical flow model. Furthermore, prediction adjustments are made using a spatial gradient and temporal optical flow model in conjunction with an affine motion model.
It improves the coding efficiency of video encoding, reduces motion residuals between samples within blocks, and enhances coding quality and compression performance.
Smart Images

Figure CN114073090B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present embodiments relate generally to a method and apparatus for bi-directional refinement with optical flow in video coding or decoding. BACKGROUND
[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transform to exploit spatial and temporal redundancy in video content. Generally, intra- or inter-prediction is used to exploit intra- or inter-picture correlation, and then the difference (often denoted as prediction error or prediction residual) between the original block and the predicted block is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction. SUMMARY
[0003] According to an embodiment, a method of video coding or decoding is provided, comprising: for a block to be coded in bi-prediction mode, obtaining first and second prediction blocks from first and second reference pictures, respectively; obtaining a first extended region comprising the first prediction block and a plurality of samples surrounding the first prediction block, wherein each of the plurality of samples surrounding the first prediction block is copied from an integer sample position closest to a corresponding fractional sample position in the first reference picture; obtaining a second extended region comprising the second prediction block and a plurality of samples surrounding the second prediction block, wherein each of the plurality of samples surrounding the second prediction block is copied from an integer sample position closest to a corresponding fractional sample position in the second reference picture; obtaining a spatial gradient for each sample in the first and second prediction blocks in response to the first and second extended regions, respectively; obtaining a motion refinement for a sample in the block based on the first and second prediction blocks and the spatial gradients; obtaining a prediction adjustment for the block based on the motion refinement and the spatial gradients; and coding or decoding the block in response to the prediction adjustment, the first prediction block, and the second prediction block.
[0004] According to another embodiment, a method of video coding or decoding is provided, comprising: accessing a block to be coded or decoded in affine mode, the block comprising a plurality of sub-blocks; for one sub-block of the plurality of sub-blocks, obtaining first and second prediction blocks from first and second reference pictures, respectively, using prediction based on sub-block affine motion compensation, the one sub-block being in bi-prediction mode; obtaining a motion refinement for the one sub-block based on the first and second prediction blocks and spatial gradients in the first and second prediction blocks according to a temporal optical flow model and a spatial optical flow model; obtaining a prediction adjustment for the one sub-block based on the motion refinement and the spatial gradients; and coding or decoding the one sub-block in response to the prediction adjustment, the first prediction block, and the second prediction block.
[0005] According to another embodiment, an apparatus for video encoding or decoding is provided, comprising: one or more processors, wherein the one or more processors are configured to: obtain, for a block to be encoded in bi-prediction mode, first and second prediction blocks from first and second reference pictures, respectively; obtain a first extended region comprising the first prediction block and a plurality of samples surrounding the first prediction block, wherein each of the plurality of samples surrounding the first prediction block is copied from an integer sample position closest to a corresponding fractional sample position in the first reference picture; obtain a second extended region comprising the second prediction block and a plurality of samples surrounding the second prediction block, wherein each of the plurality of samples surrounding the second prediction block is copied from an integer sample position closest to a corresponding fractional sample position in the second reference picture; obtain a spatial gradient for each sample in the first and second prediction blocks in response to the first and second extended regions, respectively; obtain a motion refinement for a sample in the block based on the first and second prediction blocks and the spatial gradients; obtain a prediction adjustment for the block based on the motion refinement and the spatial gradients; and encode or decode the block in response to the prediction adjustment, the first prediction block, and the second prediction block.
[0006] According to another embodiment, an apparatus for video encoding or decoding is provided, comprising: one or more processors, wherein the one or more processors are configured to: access a block to be encoded or decoded in affine mode, the block comprising a plurality of sub-blocks; obtain, for one sub-block of the plurality of sub-blocks, first and second prediction blocks from first and second reference pictures using prediction based on sub-block affine motion compensation, the one sub-block being in bi-prediction mode; obtain a motion refinement for the one sub-block based on the first and second prediction blocks and spatial gradients in the first and second prediction blocks according to a temporal and spatial optical flow model; obtain a prediction adjustment for the one sub-block based on the motion refinement and the spatial gradients; and encode or decode the one sub-block in response to the prediction adjustment, the first prediction block, and the second prediction block.
[0007] One or more embodiments also provide a computer program comprising instructions, which, when executed by one or more processors, cause the one or more processors to perform the encoding method or the decoding method according to any of the above embodiments. One or more of the embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to the above methods. One or more embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above methods. One or more embodiments also provide a method and apparatus for transmitting or receiving a bitstream generated according to the above methods. BRIEF DESCRIPTION OF DRAWINGS
[0008] Figure 1 A block diagram of a system in which aspects of the present embodiments can be implemented is shown.
[0009] Figure 2 A block diagram of an embodiment of a video encoder is shown.
[0010] Figure 3 A block diagram of an embodiment of a video decoder is shown.
[0011] Figure 4 A block diagram showing the concept of coding tree units and coding trees representing a compressed picture.
[0012] Figure 5 An affine motion model used in VVC (Versatile Video Coding) is shown.
[0013] Figure 6 An example of an optical flow trajectory is shown.
[0014] Figure 7 An extended CU region used in BDOF (Bi-directional Optical Flow) in VTM (VVC Test Model) 3.0 is shown.
[0015] Figure 8 The difference between a sub-block MV and a pixel MV v(x, y) is shown. .
[0016] Figure 9 A recent integer padding used in PROF is shown.
[0017] Figure 10 A method of applying BDOF to the motion compensation of a CU (coding unit) for bi-directional affine coding after PROF (prediction refinement with optical flow) according to an embodiment is shown.
[0018] Figure 11 A method of generating a final prediction when combining PROF and BDOF according to another embodiment is shown.
[0019] Figure 12 A method of bi-prediction refinement in affine mode is shown. DETAILED DESCRIPTION
[0020] Figure 1 A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. The system 100 can be embodied as a device including the various components described below and configured to perform one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, notebook (laptop) computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of the system 100 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components in combination with the elements described below, singly or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of the system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or other electronic devices, via, for example, communication buses or by dedicated input and / or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.
[0021] The system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 can include embedded memory, input output interface, and various other circuitries known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disks, and / or optical disks. The storage device 140 can include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.
[0022] The system 100 includes an encoder / decoder module 130 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 can include its own processor and memory. The encoder / decoder module 130 represents the modules that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 130 can be implemented as a separate element of the system 100 or can be incorporated internal to the processor 110 as a combination of hardware and software as known to those skilled in the art.
[0023] Program code to be loaded onto the processor 110 or the encoder / decoder 130 to perform the various aspects described in this application can be stored in the storage device 140 and then loaded onto the memory 120 for execution by the processor 110. In accordance with various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 can store one or more of various items during the execution of the processes described in this application. Such stored items can include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and arithmetic logic processing.
[0024] In several embodiments, the memory internal to the processor 110 and / or the encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. This external memory can be the memory 120 and / or the storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, HEVC, or VVC.
[0025] As shown in block 105, input to the elements of the system 100 can be provided through various input devices. Such input devices include, but are not limited to, (i) an RF portion that receives RF signals transmitted, e.g., by a broadcaster over the air, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0026] In various embodiments, the input devices of block 105 have associated respective input processing elements known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower frequency band to select a signal frequency band which can be referred to as a channel in certain embodiments, for example, (iv) demodulating the downconverted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements for performing these functions, such as frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers, for example. The RF portion can include a tuner that performs various ones of these functions, including downconverting the received signal (s) to lower frequency (s) (e.g., an intermediate frequency or a near-baseband frequency) or to baseband, for example. In one set-top box embodiment, the RF portion and its associated input processing elements receive an RF signal transmitted through a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements between existing elements, such as amplifiers and analog-to-digital converters, for example. In various embodiments, the RF portion includes an antenna.
[0027] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting the system 100 to other electronic devices through USB and / or HDMI connections. It will be appreciated that various aspects of input processing, such as Reed-Solomon error correction, can be implemented as desired within, for example, a separate input processing IC or within the processor 110. Similarly, aspects of USB or HDMI interface processing can be implemented as desired within a separate interface IC or within the processor 110. The demodulated, error corrected, and demultiplexed streams are provided to various processing elements, including the processor 110 and the encoder / decoder 130, for example, which operate in conjunction with memory and storage elements to process the data streams as desired for presentation on output devices.
[0028] The various elements of the system 100 can be disposed within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards, for example.
[0029] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver configured to transmit and to receive data over the communication channel 190. The communication interface 150 can include, but is not limited to, a modem or network card, and the communication channel 190 can be implemented, for example, within wired and / or wireless media.
[0030] In various embodiments, data is streamed to the system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received over the communication channel 190 and the communication interface 150 appropriate for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that communicates data over an HDMI connection of the input block 105. Other embodiments provide streamed data to the system 100 using an RF connection of the input block 105.
[0031] The system 100 can provide output signals to various output devices, including a display 165, speakers 175, and other peripheral devices 185. In various examples of embodiments, the other peripheral devices 185 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to the system 100 using the communication channel 190 via the communication interface 150. In an electronic device (e.g., a television), the display 165 and speakers 175 can be integrated in a single unit with other components of the system 100. In various embodiments, the display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0032] For example, if the RF portion of the input 105 is part of a separate set-top box, the display 165 and speakers 175 can instead be separate from one or more other components. In various embodiments in which the display 165 and speakers 175 are external components, the output signals can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.
[0033] Figure 2 An example video encoder 200, such as the High Efficiency Video Codec (HEVC) encoder, is shown. Figure 2 Encoders that improve upon the HEVC standard or employ technologies similar to HEVC, such as the VVC (Various Video Codec) encoder developed by JVET (Joint Video Exploration Group), can also be shown.
[0034] In this application, the terms "reconstruction" and "decoding" are used interchangeably, as are the terms "encoding" and "encoding / decoding," the terms "pixel" and "sample," and the terms "image," "picture," and "frame." Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0035] Before being encoded, the video sequence may undergo pre-coding (201), for example, by applying color transformations to the input color picture (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or by performing remapping of the input picture components to obtain a more resilient signal distribution to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.
[0036] In encoder 200, the frame is encoded by encoder elements as described below. The frame to be encoded is segmented (202) and processed in units such as CUs. For example, each unit is encoded using an intra-frame or inter-frame mode. When a unit is encoded in intra-frame mode, it performs intra-frame prediction (260). In inter-frame mode, motion estimation (275) and compensation (270) are performed. The encoder determines (205) which of the intra-frame or inter-frame modes to use to encode the unit and indicates the intra-frame / inter-frame decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (210) the prediction block from the original image block.
[0037] The predicted residual is then transformed (225) and quantized (230). The quantized transform coefficients, along with the motion vector and other syntax elements, are entropy encoded (245) to output a bitstream. The encoder can skip the transform and apply the quantization directly to the untransformed residual signal. The encoder can bypass the transform and quantization, i.e., directly encode and decode the residual without applying the transform or quantization process.
[0038] The encoded blocks are decoded by the encoder to provide a reference for further prediction. The quantized transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined (255) with the predicted block to reconstruct the image block. In-loop filters (265) are applied to the reconstructed picture to perform, e.g., deblocking / SAO (Sample Adaptive Offset) filtering to reduce coding artifacts. The filtered image is stored at the reference picture buffer (280).
[0039] Figure 3 A block diagram of an example video decoder 300 is shown. In the decoder 300, the bitstream is decoded by the decoder elements as described below. The video decoder 300 generally performs a decoding pass opposite to the encoding pass described in Figure 2 the encoder 200. The encoder 200 generally also performs video decoding as part of encoding the video data.
[0040] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partitioning information indicates how the picture is partitioned. The decoder can therefore divide (335) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (340) and inverse transformed (350) to decode the prediction residual. The decoded prediction residual is combined (355) with the predicted block to reconstruct the image block. The predicted block can be obtained (370) from intra prediction (360) or motion-compensated prediction (i.e., inter prediction) (375). In-loop filters (365) are applied to the reconstructed image. The filtered image is stored at the reference picture buffer (380).
[0041] The decoded picture can further undergo post-decoding processing (385), such as inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0042] Affine mode
[0043] In HEVC, only translational motion models are applied to the motion-compensated prediction. To account for other types of motion, such as zooming / shrinking, rotation, perspective motion, and other irregular motion, affine motion-compensated prediction is applied in VTM-3.0. The affine motion model in VVC is either 4-parameter or 6-parameter.
[0044] The four-parameter affine motion model has the following parameters: two parameters for translational motion in horizontal and vertical directions, one parameter for scaling motion in both directions, and one parameter for rotation motion in both directions. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. The four-parameter affine motion model is coded using two motion vectors at two control point positions defined at the top-left corner (810) and the top-right corner (820) of the current CU. As shown in Figure 5 , the affine motion field of a block is described by two control point motion vectors (V0, V1). Based on the control point motion, the motion field (v x , v y ) of one affine coded block is described as
[0045]
[0046] where (v 0x , v 0y ) is the motion vector of the top-left control point (810), and (v 1x , v 1y ) is the motion vector of the top-right control point (820), and w is the width of the CU. In VTM-3.0, the motion field of an affine coded CU is derived at 4x4 block level, i.e., (v x , v y ) is derived for each of the 4x4 blocks within the current CU and applied to the corresponding 4x4 block.
[0047] The six-parameter affine motion model has the following parameters: two parameters for translational motion in horizontal and vertical directions, one parameter for scaling motion in horizontal direction and one parameter for rotation motion in horizontal direction, and one parameter for scaling motion in vertical direction and one parameter for rotation motion in vertical direction. The six-parameter affine motion model is coded with three MVs at three control points. As shown in Figure 5 , three control points of a six-parameter affine coded CU are defined at the top-left corner, the top-right corner, and the bottom-left corner (910, 920, 930) of the CU. The motion at the top-left control point (910) is related to translational motion, the motion at the top-right control point (920) is related to rotation and scaling motion in horizontal direction, and the motion at the bottom-left control point (930) is related to rotation and scaling motion in vertical direction. For the six-parameter affine motion model, the rotation and scaling motion in horizontal direction can be different from those in vertical direction. The motion vector of each sub-block (SB v x ,v y ) is derived using three MVs at the control points as follows:
[0048]
[0049] Equation 2.
[0050] where (v 2x , v 2y ) is the motion vector of the bottom-left control point (930), (x, y) is the center position of the sub-block, w and h are the width and height of the CU.
[0051] Temporal optical flow
[0052] Bi-directional optical flow (BDOF)
[0053] Conventional bi-prediction in video coding is a simple combination of two temporal prediction blocks obtained from already reconstructed reference pictures. However, due to the limitation of block-based motion compensation (MC), there can be observable residual small motion between the samples of the two prediction blocks, thus reducing the efficiency of the motion-compensated prediction. To address this issue, the bi-directional optical flow (BDOF) tool is included in VTM 3.0 (VVC Test Model 3.0, see B. Bross, J. Chen, S. Liu, “Versatile Video Coding (Draft 3),” 12th Meeting: Macao, CN, JVET-L1001, Oct. 3-12, 2018) to reduce the impact of such motion on each sample within a block. BDOF, formerly known as BIO, is used to refine the bi-predicted signal of a CU at the 4x4 sub-block level. BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth and its luminance is constant along the considered time interval. BDOF is a sample-wise motion refinement, which is performed on top of the block-wise motion compensation used for bi-prediction. The sample-level motion refinement does not use signaling. In the case of bi-prediction, the goal of BDOF is to assume a linear displacement between the two reference pictures and to refine the motion of each sample based on the Hermite interpolation of the optical flow as shown in Figure 6
[0054] In VTM-3.0, BDOF is applied only to the luma component. BDOF is applied to a CU if the following conditions are met:
[0055] • The height of the CU is not 4 and the size of the CU is not 4x8.
[0056] • The CU is not coded using affine mode or ATMVP (alternative temporal motion vector prediction) merge mode.
[0057] • The CU is coded using “true” bi-prediction mode, i.e., one of the two reference pictures precedes the current picture in display order and the other follows the current picture in display order
[0058] In particular, in the current BDOF design, the derivation of the refined motion vector for each sample in a block is based on a classical optical flow model, which describes the relationship between the motion velocity and the luminance changes in the time and spatial domains. Let be the sample value at coordinates (x, y) of the prediction block derived from the reference picture list k (k = 0, 1), and and be the horizontal and vertical gradients of the sample. Given the optical flow model, the motion refinement at (x, y) can be derived by
[0059]
[0060] In Figure 6 , it is assumed that there exists backward reference at a temporal distance to the current picture and forward reference at a temporal distance to the current picture, (MV x0 , MV y0 ) and (MV x1 , MV y1 ) indicate the block-level motion vectors used to generate two prediction blocks and in the two reference pictures. Based on the optical flow model, to predict the current block in the current picture, we have (assuming ):
[0061]
[0062]
[0063]
[0064] Eq. 4.
[0065] where is the regular bi-prediction, and the remaining offset is the BDOF adjustment based on the motion refinement and gradient calculation.
[0066] Further, the motion refinement at sample position (x, y) is calculated by minimizing the difference between the motion refinement compensated sample values (i.e., A and B in Figure 6 ):
[0067]
[0068] To ensure regularity of the derived motion refinements, it is assumed that the motion refinements are consistent for samples within a 4x4 sub-block. For each 4x4 sub-block, the motion refinements are computed by minimizing the difference between the L0 and L1 prediction samples . The bi-prediction sample values in the 4x4 sub-block are then adjusted using the motion refinements. The following steps apply to the BDOF process.
[0069] First, the horizontal and vertical gradients of the two prediction signals are computed by directly calculating the difference between two neighboring samples, and , i.e.,
[0070]
[0071] where is the sample value at the coordinates of the prediction signal in the list . Then, the auto-correlation and cross-correlation of the gradients
[0072] , , , , and are computed as
[0073]
[0074] where is a 6x6 window around the 4x4 sub-block, and
[0075]
[0076] The motion refinements are then derived using the cross-correlation and auto-correlation terms using the following equations :
[0077]
[0078] where , and are the floor functions.
[0079] Based on the motion refinements and the gradients, the following adjustments are computed for each sample in the 4x4 sub-block:
[0080]
[0081] Finally, the BDOF samples of the CU are computed by adjusting the bi-prediction samples as follows:
[0082]
[0083] Above, , and The values are 3, 6, and 12, respectively. These values are chosen so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.
[0084] To derive the gradient values, a list outside the current CU boundary needs to be generated. ( Some prediction samples in ) .like Figure 7 As shown, BDOF in VTM-3.0 uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, a bilinear filter is used to generate prediction samples in the extended region (white area), and a standard 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values are used only for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are populated (i.e., repeated) from their nearest neighbors.
[0085] Spatial optical flow
[0086] Prediction refinement using optical flow (PROF)
[0087] Affine motion model parameters can be used to derive the motion vector for each pixel in the CU. While generating pixel-based affine motion compensation predictions is very complex due to the high memory access bandwidth requirements of this sampled MC-based approach, current VVC employs a sub-block-based affine motion compensation method. In this method, the CU is divided into 4x4 sub-blocks, each assigned a motion vector (MV) derived from the affine model parameters. The MV is the MV at the center of the sub-block. All pixels within a sub-block share the same sub-block MV. Sub-block-based affine motion compensation represents a trade-off between encoding / decoding efficiency and complexity.
[0088] At JVET-N0236 (see J. Luo and Y. He, “CE2-related: Prediction refinement using optical flow for affine modes,” 14th meeting: Geneva, CH, JVET-N0236, March 19-27, 2019), optical flow-based motion refinement has been proposed to correct block-based affine motion compensation. Specifically, to achieve finer-grained motion compensation, JVET-N0236 proposed a method for refining the prediction of sub-block-based affine motion compensation using optical flow, which derives brightness refinement using a small motion difference between the sample-by-sample affine motion and the sub-block affine motion. After performing sub-block-based affine motion compensation, the brightness prediction sampling is refined by adding the difference derived from the optical flow equation. PROF is described in the following four steps.
[0089] Step 1) Perform sub-block based affine motion compensation to generate sub-block prediction .
[0090] Step 2) Calculate spatial gradient of sub-block prediction at each sample position using a 3-tap filter [-1, 0, 1] and .
[0091]
[0092]
[0093] Equation 11: Horizontal and vertical gradient of sub-block prediction signal.
[0094] The sub-block prediction is extended by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, the pixels on the extended boundary are copied from the nearest integer pixel positions in the reference picture. Thus, additional interpolation for padding area is avoided.
[0095] Step 3) Calculate luminance prediction refinement by optical flow equation.
[0096]
[0097] Equation 12: Adjust each sample in the 4x4 sub-block with PROF.
[0098] where is the difference between the calculated pixel MV and the pixel , denoted as , MV of the sub-block to which the sample position belongs (v SB ). Figure 8
[0099] Since the affine model parameters and the pixel position relative to the sub-block center do not change between sub-blocks, can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let and be the horizontal and vertical offset from the pixel position to the sub-block center, can be derived by the following equation:
[0100]
[0101] Equation 13: Derive the motion vector refinement (v ).
[0102] For 4-parameter affine model,
[0103]
[0104] For the 6-parameter affine model,
[0105]
[0106] where are the top-left, top-right and bottom-left control point motion vectors, and are the width and height of the CU.
[0107] Step 4) Finally, the luminance prediction refinement is added to the subblock prediction The final prediction is generated according to the following equation:
[0108] .
[0109] In one embodiment, the present application proposes to apply BDOF to the motion compensation of a CU coded with bi-directional affine by using the same spatial gradient derivation rules as for PROF. The present application also proposes to combine the two block-based affine motion refinements based on optical flow (PROF and BDOF) to correct the block-based affine motion compensation. Some simplifications of the prediction refinement derivation process can be achieved if the computation of the spatial gradients and is performed for both PROF and BDOF. A general bi-prediction refinement using optical flow in affine mode is also proposed.
[0110] Activation of BDOF for affine and modification of the derivation of spatial gradients
[0111] As mentioned above, for BDOF in VTM-3.0, the subblock prediction is extended by one pixel on each side for the spatial gradient computation of the prediction signal. Therefore, an extended subblock (including the 4x4 subblock and the 6x6 window of padded samples) is used as shown in Figure 7 Since only one single motion vector is applied for non-affine CUs in each reference picture list, the generation of the padded area (i.e. the extended boundaries) can be performed once for the prediction block corresponding to the whole CU. Therefore, BDOF uses one extended row / column around the CU boundaries. The out-of-boundary prediction samples are generated by a bilinear filter.
[0112] If different MVs are applied for each 4x4 subblock, the generation of the padded area should be performed separately for each subblock prediction. Several computations are needed to derive the out-of-boundary prediction and the corresponding memory storage, which is not favorable for controlling the computational complexity. Therefore, in VTM-3.0, BDOF is not applied for CUs coded with affine mode.
[0113] To reduce memory bandwidth and complexity, when PROF generates the padding region, the pixels on the extended boundary are copied from the nearest integer pixel positions in the reference picture. Thus, additional interpolation for padding the region is avoided.
[0114] In one embodiment, the present application proposes to apply BDOF to the motion compensation of an affine coded CU and to derive the spatial gradients of the prediction signal using the same padding region generation rule as PROF, which uses the nearest integer pixels as Figure 9 illustrated. When filling the block of the prediction signal with prediction samples at integer sampling positions, each sample in the padding region is generated by copying the nearest integer sample in the prediction block. When the prediction signal is at a fractional sampling position, a new block with integer sampling positions is used, where each integer sampling position in the new block is closest to the corresponding sample in the prediction block. As Figure 9 illustrated, the integer samples (960) used to generate the extended prediction samples (970) are closest to the fractional sampling position (950).
[0115] Mathematically, let and be the corresponding horizontal and vertical sampling positions in the reference picture in 1 / 16 sampling units, expressed with 1 « ( = 4) precision, the nearest integer sampling positions for the padding region and can be derived by the following equations:
[0116]
[0117]
[0118]
[0119]
[0120] Equation 14: Derivation of the nearest integer sampling positions for padding.
[0121] where the value of is equal to 4, picW and picH indicate the width and height of the reference picture, respectively.
[0122] Applying BDOF to the motion compensation of a bi-directional affine coded CU after PROF
[0123] In this embodiment, BDOF is applied to the motion compensation of a bi-directional affine coded CU after the execution of PROF. As Figure 10As shown, for bi-predictive CUs, prediction signals can be computed (1020, 1025) for both reference picture lists. After performing PROF (1030, 1035) for a CU using affine mode, BDOF (1040) can be applied to the affine coded CU after PROF to further refine the prediction signals.
[0124] In particular, for CUs using bi-predictive signals, PROF (1030, 1035) is applied to refine the prediction signals from subblock-based affine motion compensation of reference 0 (1020) and from subblock-based affine motion compensation of reference 1 (1025) . The process of PROF is the same as before without additional modification. and luma prediction refinement and are computed by the optical flow equation and their corresponding horizontal and vertical spatial gradients , , and and the MV difference between the pixel MV and the subblock MV derived by the affine model.
[0125]
[0126] Equation 15: Prediction refinement for each reference list using PROF.
[0127] The corrected luma prediction signals and using PROF are generated by adding the prediction refinement and to the prediction signals, respectively:
[0128]
[0129]
[0130] Equation 16: Adjusting prediction samples for each reference list using PROF.
[0131] After obtaining the corrected prediction signals from both reference picture lists, the BDOF process (1040) is performed on and .
[0132] In one example, the aforementioned classic BDOF is applied. Note that the input prediction signals are not the original prediction signals and but are the ones using PROF and the corrected prediction signal. Thus, the corresponding horizontal and vertical spatial gradients , , and Based on the motion refinements and spatial gradients, the prediction refinement with BDOF can be roughly expressed as:
[0133]
[0134] Equation 17: Prediction refinement with BDOF after PROF
[0135] In this case, the final prediction is generated according to the following equation:
[0136]
[0137] Equation 18: Bi-prediction sample with BDOF adjustment after PROF
[0138] In another example, the original prediction signals and are used to derive the relevant spatial gradients and motion refinements for BDOF. The advantage is that the spatial gradients , , and have already been generated by PROF and can be reused for BDOF without additional computational and memory storage complexity.
[0139] Simplification of prediction refinement when combining PROF and BDOF
[0140] In the above example, if PROF has already been performed for a CU with affine mode, the spatial gradients , , and of the prediction signals can be directly reused for the subsequent BDOF process. Thus, when PROF and BDOF are combined, there is some overlap in the computation to derive the prediction refinement and the final corrected prediction signal. In one embodiment, some further simplification of the prediction derivation process is proposed.
[0141] As proposed in one of the above examples, the spatial gradients Figure 10 , , and obtained in steps 1030 and 1035 in are directly reused for the BDOF process in step 1040. Thus, based on these spatial gradients from PROF, the prediction refinement can be generated by the following equation
[0142]
[0143] Equation 19: Prediction refinement for BDOF reusing spatial gradients from PROF
[0144] In this case, the final prediction can be simplified according to the following equation:
[0145]
[0146] Equation 20: Simplification of bi-prediction sample adjustment with BDOF and PROF combination.
[0147] where:
[0148]
[0149] Therefore, this simplification of the combination of PROF and BDOF can further change the prediction refinement derivation and final prediction correction process, as Figure 11 illustrated.
[0150] In steps 1110 and 1115, respectively, motion-compensated prediction signals and are obtained from reference 0 and reference 1, respectively. Horizontal and vertical spatial gradients , , and are then derived in steps 1120 and 1125 using
[0151] Motion refinements for PROF and BDOF are then generated in steps 1130, 1135 and 1140, based on these spatial gradients, using Equations 15 and 19. The motion refinement for the combination of PROF and BDOF is derived (1150) for the following process . In this embodiment, the steps of computing the prediction refinement with PROF and and the corresponding corrected prediction signals and are skipped. Instead, the final prediction with the refinement of the combination of PROF and BDOF is directly derived in step 1160 with Equation 20.
[0152] Advantageously, these simplifications can reduce several overlapped computations and reduce memory accesses.
[0153] Bi-prediction refinement in affine mode
[0154] In the above, BDOF based on a temporal optical flow model and PROF based on a spatial optical flow model can be combined. More generally, in one embodiment, the present application proposes a bi-prediction refinement in affine based on spatial and temporal optical flow models, which describe the relationship of motion velocity and luminance change in spatial and temporal domains, respectively. To achieve a finer granularity of affine motion compensation on a subblock basis, the luminance prediction samples are refined by adding the difference derived from the optical flow equation. The proposed bi-prediction refinement in affine is shown in Figure 12 and described below.
[0155] First, in steps 1210 and 1215, subblock-based affine motion compensation is performed in two reference picture lists to generate subblock predictions and In steps 1220 and 1225, the horizontal and vertical spatial gradients of the two subblock prediction signals are then derived, e.g., by directly computing the difference between the two neighboring samples as described earlier , , and . The subblock predictions are extended by one pixel on each side for the gradient computation, e.g., the pixels on the extended boundaries are copied from the nearest integer pixel positions in the reference pictures.
[0156] Subsequently, in step 1230, the motion refinements from the two reference picture lists are computed, (d ) and (d ). Based on the motion refinements and the gradients, the bi-prediction refinement is computed for each sample in the 4x4 subblock as shown in step 1240:
[0157]
[0158]
[0159] Equation 21: Generation of the bi-prediction refinement.
[0160] Finally, the luminance prediction refinement is added to the subblock predictions to generate (1250) the final prediction, as follows:
[0161]
[0162] Equation 22: Prediction with the proposed bi-prediction refinement in affine.
[0163] Advantageously, the proposed bi-prediction refinement uses temporal and spatial optical flow models to achieve a sample-by-sample correction of the affine prediction samples. Furthermore, considering refinements from both directions gives an even finer affine prediction.
[0164] Various methods are described herein, and each of the methods includes one or more steps or actions for achieving the described method. Unless a specific step or action is required, the order or
[0165] Various methods and other aspects described in this application can be used to modify modules, e.g., motion compensation modules (270, 375) of video encoders 200 and decoders 300 as shown in Figure 2 and Figure 3 FIGS. 1-3. Moreover, the present aspects are not limited to VVC or HEVC, and can be applied to, e.g., other standards and recommendations, as well as extensions of any such standards and recommendations. Unless otherwise noted, or technically precluded, the aspects described in this application can be used alone or in combination.
[0166] Various numerical values are used in this application. Particular values are used for example purposes and the described aspects are not limited to these particular values.
[0167] Various implementations relate to decoding. As used in this application, “decoding” can encompass all or part of the processes performed on a received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding. Based on the context of the particular description, it will be clear whether the phrase “decoding process” is intended to specifically refer to a subset of operations or to the more general decoding process, and is considered well understood by those skilled in the art.
[0168] Various implementations relate to encoding. In a similar manner as discussed above with respect to “decoding,” as used in this application, “encoding” can encompass all or part of the processes performed on an input video sequence in order to produce an encoded bitstream.
[0169] The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single implementation form (for example, discussed only as a method), implementation of the discussed features can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, an apparatus such as, for example, a processor, which is generally a processing device that includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end-users.
[0170] Reference to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", as well as other variants, means that a particular feature, structure, characteristic, and so forth being described in connection with an embodiment is included in at least one embodiment. Therefore, the appearance of the phrase "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation", as well as any other variations, throughout the application should not necessarily be construed as referring to the same embodiment. Additionally, the application can refer to "determining" various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.
[0171] Additionally, the application can refer to "accessing" various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0172] Additionally, the application can refer to "receiving" various pieces of information. Receiving is a broad term that can be understood as one or more of, for example, accessing or retrieving the information (for example, from memory). Additionally, "receiving" is often involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.
[0173] It should be understood that the use of “ / ,” “and / or,” and “at least one of’ in the examples, for example, in the contexts of “A / B,” “A and / or B,” and “at least one of A and B,” can be used to indicate that the item(s) can be either select one of the items (A) or select both items (A and B). Further, it should be understood that references to “at least one of’ a set of items can refer to individual items in the set, combinations of items in the set, and / or subsets of items in the set. Further, references to “at least one of’ a set of items can refer to individual items in the set, combinations of items in the set, and / or subsets of items in the set. As one of ordinary skill in the art will readily appreciate, these can be extended to any number of items and any combination of items.
[0174] Further, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a quantization matrix for dequantization. As such, in embodiments, the same parameters are used at both the encoder side and the decoder side. Thus, for example, an encoder can send (explicit signaling) a particular parameter to a decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter along with other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are realized in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, and the like are used to signal information to a corresponding decoder. While the foregoing involves the verb form of the word “signal,” the word “signal” is also used as a noun herein.
[0175] It will be apparent to one of ordinary skill in the art that implementations can produce signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of an described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
Claims
1. A method comprising: accessing a block to be encoded or decoded in affine mode, the block comprising a plurality of sub-blocks; for one sub-block of the plurality of sub-blocks in bi-prediction mode, obtaining first and second prediction blocks from first and second reference pictures, respectively, using sub-block based affine motion compensated prediction; generating a third prediction block for the one sub-block based on a spatial optical flow model based on the first prediction block and spatial gradients in the first prediction block; generating a fourth prediction block for the one sub-block based on the spatial optical flow model based on the second prediction block and spatial gradients in the second prediction block; obtaining a prediction adjustment for the one sub-block based on a temporal optical flow model based on the third prediction block, spatial gradients in the first prediction block, the fourth prediction block, and spatial gradients in the second prediction block; and encoding or decoding the one sub-block in response to the prediction adjustment, the third prediction block, and the fourth prediction block. using prediction refinement with optical flow motion refinement according to the spatial optical flow model.
2. The method of claim 1, wherein, using bi-directional optical flow motion refinement according to the temporal optical flow model.
3. The method of claim 1, wherein, 4. The method of claim 1, further comprising: for each sample in the one sub-block, obtaining a motion difference between a sample level motion vector and a motion vector of the one sub-block, wherein the third prediction block is obtained based further on the motion difference.
5. An apparatus comprising one or more processors, wherein the one or more processors are configured to: access a block to be encoded or decoded in affine mode, the block comprising a plurality of sub-blocks; for one sub-block of the plurality of sub-blocks in bi-prediction mode, obtain first and second prediction blocks from first and second reference pictures, respectively, using sub-block based affine motion compensated prediction; generate a third prediction block for the one sub-block based on a spatial optical flow model based on the first prediction block and spatial gradients in the first prediction block; and generate a fourth prediction block for the one sub-block based on the spatial optical flow model based on the second prediction block and spatial gradients in the second prediction block; obtain a prediction adjustment for the one sub-block based on a temporal optical flow model based on the third prediction block, spatial gradients in the first prediction block, the fourth prediction block, and spatial gradients in the second prediction block; and encode or decode the one sub-block in response to the prediction adjustment, the third prediction block, and the fourth prediction block. use prediction refinement with optical flow motion refinement according to the spatial optical flow model. use bi-directional optical flow motion refinement according to the temporal optical flow model. the one or more processors are further configured to:
6. The apparatus of claim 5, wherein, for each sample in the one sub-block, obtain a motion difference between a sample level motion vector and a motion vector of the one sub-block, 7. The apparatus of claim 5, wherein, wherein the third prediction block is obtained based further on the motion difference.
8. The apparatus of claim 5, wherein, 9. A non-transitory computer-readable storage medium having stored thereon machine executable instructions, the machine executable instructions, when executed cause a method of encoding or decoding comprising: accessing a block to be encoded or decoded in affine mode, the block comprising a plurality of sub-blocks; for one sub-block of the plurality of sub-blocks, obtaining first and second prediction blocks from first and second reference pictures, respectively, using sub-block based affine motion compensated prediction, the one sub-block being in bi-prediction mode; generating a third prediction block for the one sub-block based on a spatial optical flow model based on the first prediction block and spatial gradients in the first prediction block; and generating a fourth prediction block for the one sub-block based on the spatial optical flow model based on the second prediction block and spatial gradients in the second prediction block; obtaining a prediction adjustment for the one sub-block based on a temporal optical flow model based on the third prediction block, spatial gradients in the first prediction block, the fourth prediction block and spatial gradients in the second prediction block; and encoding or decoding the one sub-block in response to the prediction adjustment, the third prediction block and the fourth prediction block.
10. The medium of claim 9, wherein, using prediction refinement with optical flow motion refinement according to the spatial optical flow model.
11. The medium of claim 9, wherein, using bi-directional optical flow motion refinement according to the temporal optical flow model.
12. The medium of claim 9, further comprising: for each sample in the one sub-block, obtaining a motion difference between a sample level motion vector and a motion vector of the one sub-block, wherein the third prediction block is obtained based further on the motion difference.
Citation Information
Patent Citations
Method and apparatus of motion refinement based on bi-directional optical flow for video coding
CN110476424A