PROF METHOD, COMPUTING DEVICE, NON-TRANSITORY COMPUTER-READABLE STORAGE MEDIUM, BITSTREAM TRANSMISSION METHOD, AND COMPUTER PROGRAM - Patent application

By integrating prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF) tools in the VVC standard, the inefficiencies in motion compensation are addressed, enhancing coding efficiency and simplifying hardware implementation for video encoding.

JP7681169B2Active Publication Date: 2025-05-21BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024129957
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-23
Filing Date
2024-08-06
Publication Date
2025-05-21
Estimated Expiration
2040-08-24

AI Technical Summary

Technical Problem

Existing video encoding techniques, such as those in the VVC standard, face inefficiencies in motion compensation due to limitations in block-based motion estimation, particularly for affine motion models, leading to suboptimal coding efficiency and complexity in hardware implementation.

Method used

The integration of prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF) tools in the VVC standard, which harmonize bit depth representation, workflow, and gradient derivation to improve motion refinement accuracy and reduce computational complexity, enabling unified hardware pipeline designs.

Benefits of technology

Enhances coding efficiency by aligning bit depth and workflow representation between PROF and BDOF, facilitating hardware implementation and reducing computational complexity, thereby improving video encoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681169000041
    Figure 0007681169000041
  • Figure 0007681169000042
    Figure 0007681169000042
  • Figure 0007681169000043
    Figure 0007681169000043
Patent Text Reader

Abstract

To provide a method of PROF(prediction refinement with optical flow) for video coding.SOLUTION: A method includes: obtaining first and second prediction refinements based on first and second horizontal and vertical gradient values and first and second horizontal and vertical motion refinements; obtaining first and second refined samples based on a first prediction sample I(0)(i,j) and a second prediction sample I(1)(i,j), and the first and second prediction refinements; and obtaining final prediction samples of a video block based on the first and second refined samples and prediction parameters. The prediction parameter includes parameters for weighted prediction and parameters for BCW. Parameters of WP and parameters of BCW include first weight of the refined first sample and second weight of the refined second sample, respectively.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This disclosure relates to video encoding and compression, and more particularly to methods and apparatus for two inter-prediction tools considered in the versatile video coding (VVC) standard: prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF). [Background technology]

[0002] Various video encoding techniques can be used to compress video data. Video encoding is performed according to one or more video encoding standards. For example, video encoding standards include versatile video coding (VVC), joint exploration test model (JEM), high efficiency video coding (H.265 / HEVC), advanced video coding (H.264 / AVC), moving picture expert group (MPEG) coding, and the like. Video encoding generally employs prediction methods (e.g., inter-prediction, intra-prediction, and the like) that exploit redundancy present in a video image or sequence. An important goal of video encoding techniques is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]

[0003] Examples of the present disclosure provide a method and apparatus for prediction refinement with optical flow (PROF). [Means for solving the problem]

[0004] According to a first aspect of the present disclosure, a prediction refinement with optical flow (PROF) method for video coding is provided. The method includes obtaining a video block to be coded in an affine mode, and determining a first reference picture I associated with the video block. (0) and the second reference picture I (1) The method may include obtaining a first reference picture I (0) , the second reference picture I (1) Related to the first prediction sample I (0) (i,j) and the second forecast sample I (1) The method may also include obtaining a first horizontal gradient value and a second vertical gradient value based on the first reference picture I. (0) and the second reference picture I (1) The method may further include obtaining first and second horizontal and vertical motion refinements based on control point motion vectors (CPMVs) associated with the first prediction sample I. The method may include obtaining first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and the first and second horizontal and vertical motion refinements. (0) (i,j), the second prediction sample I (1)The method may further include obtaining refined first and second samples based on (i,j) and the first and second prediction refinements. The method may include obtaining a final predicted sample of the video block based on the refined first and second samples by manipulating the refined first and second samples and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for weighted prediction (WP) and parameters for biprediction with coding unit (CU)-level weight (BCW). The parameters for weighted prediction (WP) and the parameters for biprediction with coding unit (CU)-level weight (BCW) include a first weight for the refined first sample and a second weight for the refined second sample, respectively.

[0005] According to a second aspect of the present disclosure, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable memory storing instructions executable by the one or more processors. The one or more processors may obtain a video block encoded by an affine mode and generate a first reference picture I associated with the video block. (0) and the second reference picture I (1) The one or more processors may also be configured to obtain the first reference picture I (0) , the second reference picture I (1) The first predicted sample I associated with (0) (i,j) and the second forecast sample I (1) The one or more processors may be configured to obtain a first horizontal gradient value and a second vertical gradient value based on the first reference picture I(i,j). (0)and the second reference picture I (1) The one or more processors may be configured to obtain first and second horizontal and vertical motion refinements based on control point motion vectors (CPMVs) associated with the first prediction sample I. The one or more processors may be further configured to obtain first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and the first and second horizontal and vertical motion refinements. (0) (i,j), the second prediction sample I (1) (i,j) and the first and second prediction refinements to obtain refined first and second samples. The one or more processors may be further configured to obtain a final predicted sample of the video block based on the refined first and second samples by manipulating the refined first and second samples and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for WP and parameters for BCW. The parameters for WP (weighted prediction) and BCW (biprediction with coding unit (CU)-level weight) include a first weight for the refined first sample and a second weight for the refined second sample, respectively.

[0006] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided having instructions stored thereon that, when executed by one or more processors of an apparatus, include: at an encoder, obtaining a video block to be encoded by an affine mode; (0) and the second reference picture I (1)The instructions may cause the apparatus to obtain, at the encoder, a first horizontal gradient value and a second vertical gradient value based on a first predicted sample and a second predicted sample associated with a second reference picture that is the first reference picture. The instructions may cause the apparatus to obtain, at the encoder, a first horizontal gradient value and a second vertical gradient value based on a first predicted sample and a second predicted sample associated with a second reference picture that is the first reference picture. (0) and the second reference picture I (1) The instructions can further cause the apparatus to obtain, at the encoder, first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and the first and second horizontal and vertical motion refinements. (0) (i,j), the second prediction sample I (1) The instructions may further cause the apparatus to obtain, at the encoder, refined first and second samples based on (i,j) and the first and second prediction refinements. The instructions may cause the apparatus to obtain, at the decoder, a final predicted sample of the video block based on the refined first and second samples by manipulating the refined first and second samples and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for WP and parameters for BCW. The parameters for weighted prediction (WP) and biprediction with coding unit (CU)-level weight (BCW) include a first weight for the refined first sample and a second weight for the refined second sample, respectively. [Brief description of the drawings]

[0007] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Diagram 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 2 illustrates a block division in a multi-type tree structure according to an example of the present disclosure. [Figure 3B] FIG. 2 illustrates a block division in a multi-type tree structure according to an example of the present disclosure. [Figure 3C] FIG. 2 illustrates a block division in a multi-type tree structure according to an example of the present disclosure. [Figure 3D] FIG. 2 illustrates a block division in a multi-type tree structure according to an example of the present disclosure. [Figure 3E] FIG. 2 illustrates a block division in a multi-type tree structure according to an example of the present disclosure. [Figure 4] 1 is an illustration of a bi-directional optical flow (BDOF) model according to an example of the present disclosure. [Figure 5A] FIG. 2 is a diagram of an affine model according to an example of the present disclosure. [Figure 5B] FIG. 2 is a diagram of an affine model according to an example of the present disclosure. [Figure 6] FIG. 2 is a diagram of an affine model according to an example of the present disclosure. [Figure 7] FIG. 1 is an illustration of prediction refinement with optical flow (PROF) according to an example of the present disclosure. [Figure 8] 1 is a BDOF workflow according to an example of the present disclosure. [Figure 9] 1 is a workflow of PROF according to an example of the present disclosure. [Figure 10] 1 is a combined BDOF and PROF method for decoding a video signal according to an example of the present disclosure. [Figure 11] 1 is a BDOF and PROF method for decoding a video signal according to an example of the present disclosure. [Figure 12] FIG. 1 is a diagram of a workflow of PROF for bi-prediction according to an example of the present disclosure. [Figure 13]FIG. 2 is a diagram of pipeline stages of the BDOF and PROF processes in accordance with the present disclosure. [Figure 14] FIG. 1 is a diagram of a gradient derivation method for BDOF according to the present disclosure. [Figure 15] FIG. 2 is a diagram of a gradient derivation method for PROF according to the present disclosure. [Figure 16A] FIG. 2 is a diagram illustrating deriving template samples for an affine mode according to an example of the present disclosure. [Figure 16B] FIG. 2 is a diagram illustrating deriving template samples for an affine mode according to an example of the present disclosure. [Figure 17A] FIG. 13 illustrates exclusively enabling PROF and LIC for affine mode according to an example of the present disclosure. [Figure 17B] FIG. 13 illustrates a combined enablement of PROF and LIC for affine mode according to an example of the present disclosure. [Figure 18A] FIG. 1 illustrates a proposed padding method applied to a 16×16 BDOF CU, according to an example of the present disclosure. [Figure 18B] FIG. 1 illustrates a proposed padding method applied to a 16×16 BDOF CU, according to an example of the present disclosure. [Figure 18C] FIG. 1 illustrates a proposed padding method applied to a 16×16 BDOF CU, according to an example of the present disclosure. [Figure 18D] FIG. 1 illustrates a proposed padding method applied to a 16×16 BDOF CU, according to an example of the present disclosure. [Figure 19] FIG. 1 illustrates a computing environment coupled to a user interface according to an example of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not intended to be restrictive of the present disclosure.

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.

[0010] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which the same reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present disclosure as described in the appended claims.

[0011] The terms used in this disclosure are only for describing certain embodiments and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and the like are intended to include the plural forms unless the context clearly indicates otherwise. It is also to be understood that the term "and / or" as used herein is intended to mean and include any or all possible combinations of one or more of the associated listed items.

[0012] It should be understood that, although terms such as "first," "second," and "third" may be used herein to describe various information, the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, the first information may be referred to as the second information, and similarly, the second information may be referred to as the first information, without departing from the scope of this disclosure. As used herein, the term "if" may be understood to mean "when" or "in the course of" or "in response to a determination," depending on the context.

[0013] The first version of the HEVC standard was completed in October 2013, which provides approximately 50% bitrate savings or equivalent perceptual quality compared to the previous generation video coding standard H.264 / MPEG AVC. While the HEVC standard provides significant coding improvements over its predecessors, there is evidence that better coding efficiency than HEVC can be achieved with additional coding tools. Based on this, both VCEG and MPEG have started exploring new coding techniques for future video coding standardization. ITU-T VECG and ISO / IEC MPEG formed one JVET (Joint Video Exploration Team) in October 2015 to start researching advanced techniques that could enable significant improvements in coding efficiency. One reference software, called JEM (joint exploration model), was maintained by JVET by integrating several additional coding tools on top of the HEVC test model (HM).

[0014] In October 2017, ITU-T and ISO / IEC published a joint Call for Proposals (CfP) for video compression beyond HEVC. In April 2018, 23 CfP responses were obtained and evaluated at the 10th JVET meeting, which showed an improvement in compression efficiency over HEVC of approximately 40%. Based on these evaluation results, JVET launched a new project to develop a new generation video coding standard called Versatile Video Coding (VVC). In the same month, one reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0015] Like HEVC, VVC is built on a block-based hybrid video coding framework.

[0016] Figure 1 shows a general diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows an exemplary encoder 100. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra mode / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.

[0017] At encoder 100, a video frame is divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an inter-prediction or an intra-prediction approach.

[0018] A prediction residual, which represents the difference between a current video block that is part of the video input 110 and its predictor that is part of the block predictor 140, is sent from the adder 128 to the transformer 130. The transform coefficients are then sent from the transformer 130 to a quantizer 132 for entropy reduction. The quantized coefficients are then provided to an entropy encoder 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction related information 142 from the intra / inter mode decision 116, such as video block partition information, motion vectors (MVs), reference picture indexes, and intra prediction modes, are also provided via the entropy encoder 138 and stored in the compressed bitstream 144. The compressed bitstream 144 comprises the video bitstream.

[0019] The encoder 100 also requires decoder-related circuitry to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed through inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with a block predictor 140 to generate unfiltered reconstructed pixels for the current video block.

[0020] Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples of already-encoded neighboring blocks (called reference samples) in the same video frame as the current video block.

[0021] Temporal prediction (also called "inter prediction") predicts a current video block using reconstructed pixels from an already coded video picture. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more MVs that indicate the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, one reference picture index is further transmitted, and the reference picture index is used to identify which reference picture in the reference picture storage the temporal prediction signal comes from.

[0022] Motion Estimation 114 takes signals from Video Input 110 and Picture Buffer 120 and outputs a motion estimation signal to Motion Compensation 112. Motion Compensation 112 takes signals from Video Input 110, Picture Buffer 120, and a motion estimation signal from Motion Estimation 114 and outputs a motion compensation signal to Intra / Inter Mode Decision 116.

[0023] After spatial and / or temporal prediction is performed, an intra / inter mode decision 116 in the encoder 100 selects the best prediction mode, for example, based on a rate-distortion optimization method. The block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using a transform 130 and a quantization 132. The resulting quantized residual coefficients are inverse quantized by an inverse quantization 134 and inverse transformed by an inverse transform 136 to form a reconstructed residual, which is then added to the prediction block to form a reconstructed signal for the CU. Further in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied to the reconstructed CU before it is placed into reference picture storage in the picture buffer 120 and used to encode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138, where they are further compressed and packed to form the bitstream.

[0024] Figure 1 shows a block diagram of a general block-based hybrid video coding system. The input video signal is processed block by block (called coding unit (CU)). In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which splits blocks based only on quadtree, in VVC, one coding tree unit (CTU) is split into CUs based on quadtree / binary tree / tertiary tree to adapt to changing local characteristics. Furthermore, the concept of multiple split unit types in HEVC is removed, i.e., the separation of CU, prediction unit (PU) and transform unit (TU) no longer exists in VVC, instead, each CU is always used as a basic unit for both prediction and transformation without further splitting. In the multi-type tree structure, first, one CTU is split by quadtree structure. Then, each quadtree leaf node can be further split by binary tree structure and tertiary tree structure.

[0025] As shown in Figures 3A, 3B, 3C, 3D and 3E, there are five split types: 4-split, horizontal 2-split, vertical 2-split, horizontal 3-split and vertical 3-split.

[0026] FIG. 3A is a diagram illustrating block 4-partitioning in a multi-type tree structure according to the present disclosure.

[0027] FIG. 3B is a diagram illustrating vertical bisection of a block in a multi-type tree structure according to the present disclosure.

[0028] FIG. 3C is a diagram illustrating horizontal bisection of a block in a multi-type tree structure according to the present disclosure.

[0029] FIG. 3D is a diagram illustrating a vertical 3-part division of blocks in a multi-type tree structure according to the present disclosure.

[0030] FIG. 3E is a diagram illustrating a horizontal 3-part division of blocks in a multi-type tree structure according to the present disclosure.

[0031] In FIG. 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") uses pixels from samples of already coded neighboring blocks (called reference samples) in the same video picture / slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion compensated prediction") predicts the current video block using reconstructed pixels from already coded video pictures. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) that indicate the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference pictures are supported, one reference picture index is further transmitted, which is used to identify which reference picture in the reference picture store the temporal prediction signal comes from. After spatial and / or temporal prediction, a mode decision block in the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block, and the prediction residual is de-correlated and quantized using a transform. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Further in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO) and adaptive in-loop filter (ALF), can be applied to the reconstructed CU before it is put into a reference picture store and used to encode future video blocks. To form an output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to an entropy coding unit for further compression and packing to form a bitstream.

[0032] Figure 2 shows a general block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of an exemplary decoder 200. The decoder 200 has a bitstream 210, an entropy decoding 212, an inverse quantization 214, an inverse transform 216, an adder 218, an intra / inter mode selection 220, an intra prediction 222, a memory 230, an in-loop filter 228, a motion compensation 224, a picture buffer 226, prediction related information 234, and a video output 232.

[0033] The decoder 200 is similar to the reconstruction-related parts present in the encoder 100 of FIG. 1. In the decoder 200, an input video bitstream 210 is first decoded via entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed via inverse quantization 214 and inverse transform 216 to obtain a reconstructed prediction residual. A block predictor implemented in an intra / inter mode selector 220 is configured to perform either intra prediction 222 or motion compensation 224 based on the decoded prediction information. The reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block predictor are summed using an adder 218 to obtain a set of unfiltered reconstructed pixels.

[0034] The reconstructed blocks may further pass through an in-loop filter 228 before being stored in a picture buffer 226, which serves as a reference picture store. The reconstructed video in the picture buffer 226 may be used to predict future video blocks as well as transmitted to drive a display device. In situations where the in-loop filter 228 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.

[0035] FIG. 2 gives a general block diagram of a block-based video decoder. The video bitstream is first entropy decoded in an entropy decoding unit. The coding mode and prediction information are sent to either a spatial prediction unit (if intra-coded) or a temporal prediction unit (if inter-coded) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then added. The reconstructed block may further undergo in-loop filtering before being stored in a reference picture store. The reconstructed video in the reference picture store is then sent to drive a display device and used to predict future video blocks.

[0036] In general, the basic inter-prediction technique applied in VVC is kept the same as in HEVC, except that some modules are further extended and / or enhanced. In particular, for all preceding video standards, one coding block is only associated with one single MV if the coding block is uni-predicted, and associated with two MVs if the coding block is bi-predicted. Due to such limitations of traditional block-based motion compensation, small motions can still remain within the prediction samples after motion compensation, thus negatively impacting the overall efficiency of motion compensation. To improve both the granularity and refinement of MVs, two sample-wise refinement methods based on optical flow are introduced, namely, bi-directional optical flow (BDOF) for affine mode and bi-directional optical flow (BDOF) for affine mode. Two inter-coding tools, PROF (prediction refinement with optical flow) and PROF (prediction refinement with optical flow), are currently under consideration for the VVC standard. In the following, we briefly review the main technical aspects of the two inter-coding tools.

[0037] Bidirectional Optical Flow In VVC, bidirectional optical flow (BDOF) is applied to refine the predicted samples of bi-predicted coding blocks. Specifically, as shown in FIG. 4, BDOF is a sample-wise motion refinement that is executed on top of block-based motion compensation prediction when bi-prediction is used.

[0038] FIG. 4 shows a diagram of the BDOF model according to the present disclosure.

[0039] The motion refinement (v x , v y ) of each 4×4 sub-block is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample after BDOF is applied within one 6×6 window Ω around the sub-block. Specifically, the value of (v x , v y ) is derived as follows.

Equation

Equation

[0040] S 1 , S 2 , S 3 , S 5 , S 6 is calculated as follows.

Equation

number

number

number

[0041] Based on the motion refinement derived in (1), the final bi-prediction sample of a CU is calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow model as follows:

number

[0042] Affine Mode In HEVC, only the translational motion model is applied for motion compensation prediction. Meanwhile, in the real world, there are many kinds of motions, such as zoom in / zoom out, rotation, perspective motion, and other irregular motions. In VVC, affine motion compensation prediction is applied by signaling one flag for each inter-coded block, indicating whether the translational motion model or the affine motion model is applied for inter prediction. In the current VVC design, two affine modes are supported for one affine-coded block, including a 4-parameter affine mode and a 6-parameter affine mode.

[0043] The four-parameter affine model has the following parameters: two parameters for horizontal and vertical translational motions, respectively, one parameter for zoom motion, and one parameter for rotational motions in both directions. The horizontal zoom parameter is the same as the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To achieve better adaptation of the motion vectors and affine parameters, in VVC, these affine parameters are converted into two MVs (also called control point motion vectors (CPMVs)) located at the top-left and top-right corners of the current block. As shown in Figures 5A and 5B, the affine motion field of a block is calculated by dividing the two control point MVs (V 0 , V 1 )

[0044] FIG. 5A shows a diagram of a four-parameter affine model in accordance with the present disclosure.

[0045] FIG. 5B shows a diagram of a four-parameter affine model according to the present disclosure.

[0046] Based on the control point motion, the motion field of one affine coding block (v x ,v y ) can be written as follows:

number

[0047] The 6-parameter affine mode has two parameters for horizontal and vertical translational motion, one parameter for zoom motion and one parameter for horizontal rotational motion, one parameter for zoom motion and one parameter for vertical rotational motion. The 6-parameter affine motion model is coded in 3 MVs with 3 CPMVs.

[0048] FIG. 6 shows a diagram of a six-parameter affine model according to the present disclosure.

[0049] As shown in Figure 6, the three control points of one six-parameter affine block are located at the top-left, top-right, and bottom-left corners of the block. The motion at the top-left control point is associated with translation motion, the motion at the top-right control point is associated with horizontal rotation and zoom motion, and the motion at the bottom-left control point is associated with vertical rotation and zoom motion. Compared with the four-parameter affine motion model, the horizontal rotation and zoom motion of the six-parameter model may not be the same as the vertical motion. (V 0 , V 1 , V 2 ) are the MVs of the top left, top right, and bottom left corners of the current block in FIG. 6. Then, the motion vectors (v x ,v y ) is derived using three MVs at the control points as follows:

number

[0050] Prediction refinement by optical flow in affine mode (PROF) To improve the accuracy of affine motion compensation, PROF is currently considered in the current VVC to refine the sub-block-based affine motion compensation based on the optical flow model. Specifically, after performing the sub-block-based affine motion compensation, the luminance prediction samples of one affine block are modified by one sample refinement value derived based on the optical flow equation. In detail, the operation of PROF can be summarized as the following four steps:

[0051] Step 1: Sub-block based affine motion compensation is performed to generate sub-block predictions I(i,j) using the sub-block MVs as derived in (6) for the 4-parameter affine model and in (7) for the 6-parameter affine model.

[0052] Step 2: The spatial gradient g x (i,j) and g y The values ​​of (i,j) and each predicted sample are calculated as follows: JPEG0007681169000010.jpg15154

[0053] To compute the gradient, one additional row / column of prediction samples needs to be generated on each side of one sub-block. To reduce memory bandwidth and complexity, samples on the extension boundary are copied from the nearest integer pixel location in the reference picture to avoid an additional interpolation process.

[0054] Step 3: The luminance prediction refinement value is calculated by:

number

number

[0055] FIG. 7 illustrates the PROF process for affine modes in accordance with the present disclosure.

[0056] Since the affine model parameters and pixel position relative to the subblock center do not change for each subblock, Δv(i,j) can be calculated for the first subblock and reused for other subblocks in the same CU. If Δx and Δy are the horizontal and vertical offsets from a sample position (i,j) to the center of the subblock to which the sample belongs, then Δv(i,j) can be derived as follows:

number

[0057] Based on the affine sub-block MV derivation equations (6) and (7), the MV difference Δv(i,j) can be derived. Specifically, for the four-parameter affine model,

number

[0058] For the six-parameter affine model:

number

[0059] Local Lighting Compensation Local illumination compensation (LIC) is a coding tool used to address the problem of local illumination changes that exist between temporally adjacent pictures. A pair of weights and offset parameters are applied to a reference sample to obtain a predicted sample of one current block. The general mathematical model is given as follows:

number

number

[0060] I represents the number of samples in the template. P c [x i ] is the i-th sample of the template for the current block, and P r [x i ] is the reference sample for the i-th template sample based on the motion vector v.

[0061] In addition to being applied to normal inter blocks that contain at most one motion vector for each prediction direction (L0 or L1), LIC is also applied to affine mode coded blocks, where one coded block is further divided into multiple smaller sub-blocks, and each sub-block may be associated with different motion information. To derive reference samples for LIC of affine mode coded blocks, the reference samples of the top template of one affine coded block are fetched using the motion vectors of each sub-block in the top sub-block row, while the reference samples of the left template are fetched using the motion vectors of the sub-blocks in the left sub-block column, as shown in Figures 16A and 16B (described below). Then, a LLMSE derivation method similar to (12) is applied to derive LIC parameters based on the composite template.

[0062] 16A shows a diagram for deriving a template sample for an affine mode according to the present disclosure. The diagram includes a Cur Frame 1620 and a Cur CU 1622. The Cur Frame 1620 is the current frame. The Cur CU 1622 is the current coding unit.

[0063] FIG. 16B shows a diagram for deriving template samples for affine modes. The diagram is shown in Ref Frame 1640, Col CU 1642, A Ref 1643, B Ref1644, C Ref1645, D Ref1646, E Ref1647, F 1643, B Ref1644, C Ref1645, D Ref1646, E Ref1647, F Ref1648, and G Ref1649. Ref Frame1640 is a reference frame. Col CU1642 is a collocated coding unit. A Ref1643, B Ref1644, C Ref1645, D Ref1646, E Ref1647, F Ref1648, G Ref1649 are reference samples.

[0064] Inefficiency of Prediction Refinement by Optical Flow for Affine Modes Although PROF can improve the coding efficiency of affine mode, its design can still be further improved. In particular, considering that both PROF and BDOF are built on the optical flow concept, it is highly desirable to harmonize the designs of PROF and BDOF as much as possible so that PROF can make the most of the existing logic of BDOF to facilitate hardware implementation. Based on such considerations, the following inefficiencies regarding the interaction between the current PROF design and BDOF design are identified in this disclosure.

[0065] As explained in 1. “Prediction refinement by affine mode optical flow”, in equation (8), the accuracy of the gradient is determined based on the internal bit depth. On the other hand, the MV difference, i.e., Δv x and △v yis always derived with an accuracy of 1 / 32-pel. Correspondingly, based on equation (9), the accuracy of the derived PROF refinement depends on the internal bit depth. However, similar to BDOF, PROF is applied on top of predicted sample values ​​with a medium level high bit depth (i.e., 16 bits) to maintain a higher PROF derivation accuracy. Therefore, regardless of the internal coding bit depth, the accuracy of the prediction refinement derived by PROF must match the accuracy of the intermediate predicted sample, i.e., 16 bits. In other words, the representation bit depth of the MV difference and gradient in the existing PROF design is not perfectly aligned to derive accurate prediction refinement compared to the predicted sample accuracy (i.e., 16 bits). Meanwhile, based on the comparison of equations (1), (4) and (8), the existing PROF and BDOF use different accuracy to represent sample gradient and MV difference. As pointed out earlier, such a non-uniform design is not desirable for hardware because the existing BDOF logic cannot be reused.

[0066] 2. As discussed in the section "Prediction refinement by optical flow in affine mode", when one current affine block is bi-predicted, PROF is applied to the prediction samples of lists L0 and L1 separately, and then the enhanced L0 and L1 prediction signals are averaged to generate the final bi-prediction signal. In contrast, instead of deriving PROF refinement for each prediction direction separately, BDOF derives a prediction refinement once, which is then applied to enhance the joint L0 and L1 prediction signals. Figures 8 and 9 (described later) compare the current BDOF and PROF workflows for bi-prediction. In practical codec hardware pipeline designs, different main encoding / decoding modules are usually assigned to each pipeline stage so that more coding blocks can be processed in parallel. However, the differences between the BDOF workflow and the PROF workflow may make it difficult to have one and the same pipeline design that can be shared by BDOF and PROF, which is disadvantageous for practical codec implementation.

[0067] 8 illustrates a workflow of BDOF according to this disclosure. The workflow 800 includes L0 motion compensation 810, L1 motion compensation 820, and BDOF 830. The L0 motion compensation 810 can be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture previous to the current picture in the video block. For example, the L1 motion compensation 820 can be a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. The BDOF 830 takes the motion compensation samples from the L1 motion compensation 810 and the L1 motion compensation 820 and outputs prediction samples, as described above with respect to FIG. 4.

[0068] FIG. 9 illustrates a workflow of an existing PROF according to the present disclosure. The workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0 PROF 930, L1 PROF 940, and averaging 960. The L0 motion compensation 910 can be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture previous to the current picture in the video block. For example, the L1 motion compensation 920 can be a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. The L0 PROF 930 takes the L0 motion compensation samples from the L0 motion compensation 910 as described with respect to FIG. 7 above, and outputs a motion refinement value. The L1 PROF 940 takes the L1 motion compensation samples from the L1 motion compensation 920 as described with respect to FIG. 7 above, and outputs a motion refinement value. Averaging 960 averages the motion refinement values ​​output of L0 PROF 930 and L1 PROF 940.

[0069] 3.For both BDOF and PROF, the gradient needs to be calculated for each sample in the current coding block, which requires generating one additional row / column of predicted samples on each side of the block. To avoid the additional computational complexity of sample interpolation, the predicted samples in the extension region around the block are directly copied from the reference samples at integer positions (i.e., without interpolation). However, in existing designs, integer samples at different positions are selected to generate gradient values ​​for BDOF and PROF. Specifically, integer reference samples located to the left of the predicted sample (for horizontal gradients) and above the predicted sample (for vertical gradients) are used in BDOF, while the integer reference sample closest to the predicted sample is used for gradient calculation in PROF. Similar to the bit depth representation problem, such a non-uniform gradient calculation method is also undesirable for hardware codec implementation.

[0070] 4. As pointed out earlier, the motivation of PROF is to compensate for small MV differences between the MV of each sample and the sub-block MV derived at the center of the sub-block to which the sample belongs. According to the current PROF design, PROF is always invoked when one coding block is predicted by affine mode. However, as shown in equations (6) and (7), the sub-block MV of one affine block is derived from the control point MV. Therefore, when the difference between the control point MVs is relatively small, the MV at each sample position should be consistent. In such cases, the benefit of applying PROF may be very limited, so considering the performance / complexity tradeoff, it may not be worth doing PROF.

[0071] Improving Prediction Refinement by Optical Flow for Affine Modes In this disclosure, a method is provided to improve and simplify the existing PROF design to facilitate hardware codec implementation. Particular attention is paid to harmonizing the BDOF and PROF designs to maximize sharing of existing BDOF logic with PROF. In general, the main aspects of the technology proposed in this disclosure are summarized below.

[0072] 1. In order to improve the coding efficiency of PROF while achieving one or more unified designs, we propose one way to unify the representation bit depths of sample gradients and MV differences used by BDOF and PROF.

[0073] 2. To facilitate hardware pipeline design, we propose to harmonize the workflow of PROF with that of BDOF for bi-prediction. Specifically, unlike existing PROF, which derives prediction refinements separately for L0 and L1, our proposed method derives a single prediction refinement that is applied to the combined prediction signal of L0 and L1.

[0074] 3. We propose two methods to reconcile the derivation of integer reference samples to compute the gradient values ​​used by BDOF and PROF.

[0075] 4. To reduce the computational complexity, we propose an early stopping method that adaptively disables the PROF process for affine coding blocks when certain conditions are met.

[0076] Improved bit depth representation design of PROF gradient and MV difference As analyzed in the "Problem Description" section, the representation bit depth of MV difference and sample gradients in the current PROF are not aligned to derive accurate prediction refinement. Moreover, the representation bit depth of sample gradients and MV difference are inconsistent between BDOF and PROF, which is disadvantageous for hardware. In this section, we propose an improved bit depth representation method by extending the bit depth representation method of BDOF to PROF. Specifically, the proposed method calculates the horizontal and vertical gradients at each sample position as follows:

number

[0077] In addition, assuming horizontal and vertical offsets Δx and Δy, expressed with 1 / 4-pel accuracy, from one sample location to the center of the sub-block to which the sample belongs, the corresponding PROF MV difference Δv(x,y) at the sample location is derived as follows:

number

number

[0078] For the six-parameter affine model:

number

[0079] In the above description, a pair of fixed right shifts are applied to calculate the gradient and MV difference values, as shown in equations (13) and (14). In practice, different bitwise right shifts can be applied to (13) and (14) to achieve various representation precisions of gradient and MV difference due to different tradeoffs between intermediate calculation precision and bit width of the internal PROF derivation process. For example, if the input video contains a lot of noise, the derived gradient may not be reliable to represent the true local horizontal / vertical gradient value at each sample. In such a case, it makes sense to use more bits to represent the MV difference than the gradient. On the other hand, if the input video exhibits stationary motion, the MV difference derived by the affine model should be very small. If so, no additional benefit can be gained from using high-precision MV difference to increase the precision of the derived PROF refinement. In other words, in such a case, it is more beneficial to use more bits to represent the gradient value. Based on the above considerations, one embodiment of the present disclosure proposes one general method to calculate the gradient and MV difference of PROF, specifically, assuming the horizontal gradient and vertical gradient at each sample position is n-th the difference of the adjacent predicted samples. a It is calculated by applying a right shift.

number

number

number

[0080] In another embodiment of the present disclosure, another PROF bit depth control method is proposed as follows: In this method, the horizontal gradient and vertical gradient at each sample position are right-shifted to the difference value of the adjacent predicted samples. a By applying the bits, it is still calculated as in (18). The corresponding PROF MV difference Δv(x,y) at the sample position should be calculated as follows:

number

[0081] Furthermore, to keep the entire PROF derivation at the proper internal bit depth, clipping is applied to the MV difference derived as follows:

number

number

[0082] Integrated PROF and BDOF Workflow for Bi-Forecasting As mentioned above, when one affine-coded block is bi-predicted, the current PROF is applied in one direction. More specifically, PROF sample refinements are derived separately and applied to the predicted samples in lists L0 and L1. Then, the refined prediction signals from lists L0 and L1 are averaged to generate the final bi-predicted signal for the block. This is in contrast to the BDOF design, where sample refinements are derived and applied to the bi-predicted signal. Such a difference between the bi-prediction workflows of BDOF and PROF may be disadvantageous for practical codec pipeline design.

[0083] To facilitate the hardware pipeline design, one simplification method according to the present disclosure is to modify the bi-prediction process of PROF so that the workflow of the two prediction refinement methods is harmonized. Specifically, instead of applying the refinement for each prediction direction separately, the proposed PROF method derives prediction refinement once based on the control point MV of lists L0 and L1. Then, to enhance the quality, the derived prediction refinement is applied to the combined L0 and L1 prediction signals. Specifically, based on the MV difference derived in equation (14), the final bi-prediction sample of one affine coding block is calculated by the proposed method as follows:

number

[0084] FIG. 12 shows a diagram of a PROF process when the proposed bi-predictive PROF method according to the present disclosure is applied. The PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bi-predictive PROF 1230. For example, the L0 motion compensation 1210 can be a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture previous to the current picture in the video block. The L1 motion compensation 1220 can be a list of motion compensation samples from a next reference picture, for example. The next reference picture is a reference picture after the current picture in the video block. The bi-predictive PROF 1230 takes in motion compensation samples from the L1 motion compensation 1210 and the L1 motion compensation 1220 as described above, and outputs bi-predictive samples.

[0085] FIG. 12 shows the corresponding PROF process when the proposed bi-prediction PROF method is applied. PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bi-prediction PROF 1230. For example, L0 motion compensation 1210 can be a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture previous to the current picture in the video block. L1 motion compensation 1220 can be a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. Bi-prediction PROF 1230 takes the motion compensation samples from L1 motion compensation 1210 and L1 motion compensation 1220 as described above, and outputs bi-prediction samples.

[0086] To demonstrate the potential benefits of the proposed method for hardware pipeline design, Fig. 13 shows an example to illustrate the pipeline stages when both BDOF and the proposed PROF are applied. In Fig. 13, the decoding process of one inter block mainly includes three steps.

[0087] 1. Parse / decode the MV of the coding block and fetch the reference samples. 2. Generate L0 and / or L1 prediction signals for the coding block. 3. If the coding block is predicted by a non-affine mode, perform a sample-by-sample refinement of the generated bi-predicted samples based on BDOF, and if the coding block is predicted by an affine mode, perform it based on PROF.

[0088] FIG. 13 shows an example of pipeline stages when both BDOF and proposed PROF are applied according to the present disclosure. FIG. 13 shows the potential advantages of the proposed method for hardware pipeline design. Pipeline stages 1300 include parsing / decoding MV and fetching reference samples 1310, motion compensation 1320, BDOF / PROF 1330. Pipeline stages 1300 encode video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block starts parsing / decoding MV, fetching reference samples 1310, moving to motion compensation 1320, then motion compensation 1320, BDOF / PROF 1330 sequentially. This means that BLK0 does not start processing in pipeline stages 1300 until BLK0 moves to motion compensation 1320. It is the same for all stages and video blocks as time goes from T0 to T1, T2, T3, and T4.

[0089] In FIG. 13, the decoding process of one inter block mainly includes three steps. First, parse / decode the MV of the coding block and fetch the reference samples. Second, generate the L0 and / or L1 prediction signals for the coding block. Third, perform sample-wise refinement of the generated bi-predicted samples based on BDOF if the coding block is predicted by one non-affine mode, and based on PROF if the coding block is predicted by an affine mode.

[0090] As shown in Fig. 13, after the proposed harmonization method is applied, both BDOF and PROF are applied directly to the bi-predicted samples. Assuming that BDOF and PROF are applied to different types of coding blocks (i.e., BDOF is applied to non-affine blocks and PROF is applied to affine blocks), the two coding tools cannot be invoked simultaneously. Therefore, their corresponding decoding processes can be performed by sharing the same pipeline stage. This is more efficient than the existing PROF design, which has difficulty in allocating the same pipeline stage to both BDOF and PROF, due to their different bi-prediction workflows.

[0091] In the above, the proposed method only considers the harmony of the workflow of BDOF and PROF, but in the existing design, the basic operation units of the two encoding tools are also different in size. For example, in the case of BDOF, one encoding block has a size of W s ×H s where W s =min(W,16) and H s=min(H,16), where W and H are the width and height of the coding block. BODF operations such as gradient calculation and sample refinement derivation are performed independently for each subblock. Meanwhile, as mentioned before, the affine coding block is divided into 4×4 subblocks, and each subblock is assigned one individual MV derived based on either the 4-parameter or 6-parameter affine model. Since PROF is only applied to affine blocks, its basic operation unit is 4×4 subblock. Similar to the bi-prediction workflow problem, using a different basic operation unit size for PROF than BDOF is also inconvenient for hardware implementation, making it difficult for BDOF and PROF to share the same pipeline stage of the entire decoding process. To solve such problems, in one embodiment, it is proposed to make the subblock size of affine mode the same as that of BDOF.

[0092] According to the proposed method, when a coding block is coded by the affine mode, it is coded by W s ×H s where W s =min(W,16) and H s =min(H,16), where W and H are the width and height of the coding block. Each sub-block is assigned one separate MV and is considered as one independent PROF computation unit. It is worth mentioning that the independent PROF computation unit ensures that the PROF operation on it is performed without referring to information from neighboring PROF computation units. Specifically, the PROF MV difference at one sample position is calculated as the difference between the MV at the sample position and the MV at the center of the PROF computation unit where the sample is located, and the gradient used by the PROF derivation is calculated by padding samples along each PROF computation unit.

[0093] The claimed advantages of the proposed method mainly include the following aspects: 1) a simplified pipeline architecture with a combined basic arithmetic unit size for both motion compensation and BDOF / PROF refinement, 2) reduced memory bandwidth usage due to an enlarged sub-block size for affine motion compensation, and 3) reduced per-sample computational complexity of fractional sample interpolation.

[0094] It should be mentioned that due to the computational complexity reduction by the proposed method (i.e., item 3), the existing 6-tap interpolation filter constraint for affine-coded blocks can be removed. Instead, the default 8-tap interpolation for non-affine-coded blocks is also used for affine-coded blocks. The overall computational complexity in this case still compares favorably with the existing PROF design (based on 4x4 sub-blocks with 6-tap interpolation filters).

[0095] Harmonization of BDOF and PROF gradient derivations As mentioned above, both BDOF and PROF calculate the gradients of each sample in the current coding block, which accesses one additional row / column of predicted samples on each side of the block. To avoid additional interpolation complexity, the required predicted samples in the extension region around the block boundary are directly copied from the integer reference samples. However, as pointed out in the "Problem Description" section, integer samples at different positions are used to calculate the gradient values ​​of BDOF and PROF.

[0096] In order to achieve uniform design, two methods are proposed to integrate the gradient derivation methods used by BDOF and PROF. The first method proposes to align the gradient derivation method of PROF to be the same as that of BDOF. Specifically, by the first method, the integer positions used to generate the predicted samples in the extended region are determined by flooring down the fractional sample positions, i.e., the selected integer sample positions are located to the left of the fractional sample positions (for horizontal gradients) and above the fractional sample positions (for vertical gradients). The second method proposes to make the gradient derivation method of BDOF the same as that of PROF, more specifically, when the second method is applied, the integer reference sample closest to the predicted sample is used for gradient calculation.

[0097] Figure 14 shows an example of using the BDOF gradient derivation method according to the present disclosure. In Figure 14, the white circles represent reference samples at integer positions, the triangles represent fractional predicted samples of the current block, and the grey circles represent integer reference samples used to fill the extension region of the current block.

[0098] Figure 15 shows an example of using the gradient derivation method of PROF according to the present disclosure. In Figure 15, the white circles represent reference samples at integer positions, the triangles represent fractional predicted samples of the current block, and the grey circles represent integer reference samples used to fill the extension region of the current block.

[0099] Figures 14 and 15 show the corresponding integer sample positions used to derive the gradients of BDOF and PROF when the first method (Figure 12) and the second method (Figure 13) are applied, respectively. In Figures 14 and 15, the open circles represent the reference samples at integer positions, the triangles represent the fractional predicted samples of the current block, and the patterned circles represent the integer reference samples used to fill the extension region of the current block for gradient derivation.

[0100] Moreover, according to the existing BDOF and PROF designs, the predicted sample padding is performed at different coding levels. Specifically for BDOF, padding is applied along the boundary of each sbWidth×sbHeight subblock, where sbWidth=min(CUWidth, 16) and sbHeight=min(CUHeight, 16). CUWidth and CUHeight are the width and height of the CU. Meanwhile, the padding of PROF is always applied at the 4×4 subblock level. Although the padding subblock size is still different in the above description, only the padding method is unified between BDOF and PROF. This is also not user-friendly for practical hardware implementation, considering that different modules need to be implemented for the padding process of BDOF and PROF. In order to achieve another unified design, it is proposed to unify the subblock padding size of BDOF and PROF. In one embodiment of the present disclosure, it is proposed to apply the predicted sample padding of BDOF at the 4×4 level. Specifically, with this method, a CU is first divided into multiple 4x4 sub-blocks, and after motion compensation of each 4x4 sub-block, the extension samples along the top / bottom and left / right boundaries are padded by copying the corresponding integer sample positions. Figures 18A, 18B, 18C, and 18D show an example in which the proposed padding method is applied to one 16x16BDOF CU, where the dashed lines represent the 4x4 sub-block boundaries and the gray bands represent the padded samples of each 4x4 sub-block.

[0101] FIG. 18A illustrates the proposed padding method applied to a 16×16 BDOF CU, where the dashed line represents the top-left 4×4 sub-block boundary 1820, in accordance with the present disclosure.

[0102] FIG. 18B illustrates the proposed padding method applied to a 16×16 BDOF CU, where the dashed line represents the top-right 4×4 sub-block boundary 1840 in accordance with this disclosure.

[0103] FIG. 18C illustrates the proposed padding method applied to a 16×16 BDOF CU, where the dashed line indicates the bottom-left 4×4 sub-block boundary 1860, in accordance with the present disclosure.

[0104] FIG. 18D illustrates the proposed padding method applied to a 16×16 BDOF CU, where the dashed line indicates the bottom-right 4×4 sub-block boundary 1880, in accordance with the present disclosure.

[0105] High-level signaling syntax for enabling / disabling BDOF, PROF, DMVR In the existing BDOF and PROF design, two different flags are signaled in the sequence parameter set (SPS) to separately control the enabling / disabling of the two coding tools. However, due to the similarity between BDOF and PROF, it is more desirable to enable and / or disable BDOF and PROF from a high level by one and the same control flag. Based on this consideration, one new flag called sps_bdof_prof_enabled_flag is introduced in SPS, as shown in Table 1.

[0106] As shown in Table 1, enabling and disabling of BDOF depends only on sps_bdof_prof_enabled_flag. When the flag is equal to 1, BDOF is enabled for coding the video content of the sequence. Otherwise, when sps_bdof_prof_enabled_flag is equal to 0, BDOF is not applied. Meanwhile, in addition to sps_bdof_prof_enabled_flag, an SPS-level affine control flag, namely sps_affine_enabled_flag, is also used to conditionally enable and disable PROF. When both sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are equal to 1, PROF is enabled for all coding blocks that are coded in affine mode. When the flag sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0, PROF is disabled. [Table 1] Table 1. SPS syntax table modifications with proposed BDOF / PROF enable / disable flags

[0107] sps_bdof_prof_enabled_flag specifies whether optical flow prediction refinement and bidirectional optical flow are enabled. If sps_bdof_prof_enabled_flag is equal to 0, both optical flow prediction refinement and bidirectional optical flow are disabled. If sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 1, both optical flow prediction refinement and bidirectional optical flow are enabled. Otherwise (sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0), bidirectional optical flow is enabled and optical flow prediction refinement is disabled.

[0108] sps_bdof_prof_dmvr_slice_preset_flag specifies when slice_disable_bdof_prof_dmvr_flag is signaled at the slice level. If the flag is equal to 1, the syntax slice_disable_bdof_prof_dmvr_flag is signaled for each slice that references the current sequence parameter set. Otherwise (if sps_bdof_prof_dmvr_slice_present_flag is 0), the syntax slice_disabled_bdof_prof_dmvr_flag is not signaled at the slice level. If the flag is not signaled, it is inferred to be 0.

[0109] In addition to the above SPS BDOF / PROF syntax, it is proposed to introduce another control flag at slice level, i.e. slice_disable_bdof_prof_dmvr_flag is introduced to disable BDOF, PROF and DMVR. SPS flag sps_bdof_prof_dmvr_slice_present_flag is signaled in SPS if either DMVR or BDOF / PROF sps level control flags are true and is used to indicate the presence of slice_disable_bdof_prof_dmvr_flag. If present, slice_disable_bdof_dmvr_flag is signaled. Table 2 shows the modified slice header syntax table after the proposed syntax is applied. [Table 2] Table 2. SPS syntax table modifications with proposed BDOF / PROF enable / disable flags

[0110] Early termination of PROF based on control point MV difference According to the current PROF design, PROF is always invoked when one coding block is predicted by affine mode. However, as shown in equations (6) and (7), the sub-block MVs of one affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is relatively small, the MVs at each sample position should be consistent. In such a case, the benefit of applying PROF may be very limited. Therefore, to further reduce the average computational complexity of PROF, we proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between sample-wise MVs and sub-block-wise MVs within one 4 × 4 sub-block. Since the values ​​of PROF MV differences of samples inside one 4 × 4 sub-block are symmetric with respect to the sub-block center, the maximum horizontal and vertical PROF MV differences can be calculated based on the following equation (10).

number

[0111] According to this disclosure, different metrics can be used in determining whether the MV difference is small enough to skip the PROF process. In one example, based on equation (19), if the sum of the absolute maximum horizontal MV difference and the absolute maximum vertical MV difference is less than one predefined threshold, the PROF process can be skipped, that is,

number

[0112] In another example, |△v x max ||△v y max If the maximum value of | is less than or equal to the threshold, the PROF process can be skipped.

number

[0113] MAX(a, b) is a function that returns the larger value between the input values ​​a and b.

[0114] In addition to the above two examples, the ideas of the present disclosure are also applicable when other metrics are used in determining whether the MV difference is small enough to skip the PROF process.

[0115] In the above method, PROF is skipped based on the magnitude of MV difference. Meanwhile, in addition to MV difference, PROF sample refinement is also calculated based on local gradient information at each sample position in one motion compensation block. For prediction blocks with less high frequency details (such as flat regions), the gradient value tends to be small, so that the derived sample refinement value is small. In view of this, according to another embodiment of the present disclosure, it is proposed to apply PROF only to prediction samples of blocks that contain sufficient high frequency information.

[0116] Different metrics can be used in determining whether a block contains enough high frequency information to make it worthwhile to invoke the PROF process on the block. In one example, the decision is made based on the average magnitude (i.e., absolute value) of the gradient of the samples in the predicted block. If the average magnitude is less than a threshold, the predicted block is classified as a flat region and PROF should not be applied, otherwise the predicted block is considered to contain enough high frequency detail that PROF is still applicable. In another example, the maximum magnitude of the gradient of the samples in the predicted block can be used. If the maximum magnitude is less than a threshold, PROF should be skipped for the block. In yet another example, the difference I between the maximum and minimum sample values ​​of the predicted block is used. max -I min A difference value of 0.01 can be used to determine whether PROF applies to a block. If such difference value is less than a threshold, PROF is skipped for the block. Note that the ideas of this disclosure are also applicable when some other metric is used in determining whether a given block contains sufficient high frequency information.

[0117] Dealing with the interaction between LIC and PROF in affine mode Since the neighboring reconstructed samples (i.e., templates) of the current block are used by LIC to derive the linear model parameters, the decoding of one LIC-coded block depends on the perfect reconstruction of its neighboring samples. Due to such interdependence, in a practical hardware implementation, LIC needs to be performed in the reconstruction stage where the neighboring reconstructed samples become available for LIC parameter derivation. Since the block reconstructions must be performed serially (i.e., one after the other), the throughput (i.e., the amount of work that can be performed in parallel per unit time) is one important issue to consider when applying other coding methods in conjunction with a LIC-coded block. In this section, we proposed two methods to handle the interaction when both PROF and LIC are valid for the affine mode.

[0118] In the first embodiment of the present disclosure, it is proposed to apply the PROF mode and the LIC mode exclusively to one affine coding block. As mentioned above, in the existing design, PROF is implicitly applied to all affine blocks without signaling, while one LIC flag is signaled or inherited at the coding block level to indicate whether the LIC mode is applied to one affine block. According to the method of the present invention, it is proposed to conditionally apply PROF based on the value of the LIC flag of one affine block. If the flag is equal to 1, only LIC is applied by adjusting the prediction samples of the entire coding block based on the LIC weight and offset. Otherwise (i.e., if the LIC flag is equal to 0), PROF is applied to the affine coding block to refine the prediction samples of each sub-block based on the optical flow model.

[0119] FIG. 17A shows one exemplary flowchart of the decoding process based on the proposed method, in which PROF and LIC are prohibited from being applied simultaneously.

[0120] FIG. 17A shows a diagram of the decryption process based on the proposed method where PROF and LIC are not allowed according to the present disclosure. The decryption process 1720 includes steps LIC flag on? 1722, LIC 1724, and PROF 1726. LIC flag on? 1722 is a step of determining whether the LIC flag is set or not, and taking the next step according to the determination. LIC 1724 is the application of the LIC where the LIC flag is set. If the LIC flag is not set, PROF 1726 is the application of the PROF.

[0121] In the second embodiment of the present disclosure, it is proposed to apply LIC after PROF to generate a predicted sample of one affine block. Specifically, after the sub-block-based affine motion compensation is performed, the predicted sample is refined based on PROF sample refinement, and then LIC is performed by applying a pair of weights and offsets (as derived from a template and its reference sample) to the PROF-adjusted predicted sample to obtain the final predicted sample of the block as follows:

number

[0122] FIG. 17B shows a diagram of a decoding process in which PROF and LIC are applied according to the present disclosure. The decoding process 1760 includes affine motion compensation 1762, LIC parameter derivation 1764, PROF 1766, and LIC sample adjustment 1768. Affine motion compensation 1762 applies affine motion and is input to LIC parameter derivation 1764 and PROF 1766. LIC parameter derivation 1764 is applied to derive LIC parameters. PROF 1766 is applied to PROF. LIC sample adjustment 1768 is a LIC weight and offset parameter that is combined with PROF.

[0123] Figure 17B is a diagram showing an example of a decoding workflow when the second method is applied. As shown in Figure 17B, since LIC uses a template (i.e., adjacent reconstructed samples) to calculate a LIC linear model, the LIC parameters can be derived as soon as the adjacent reconstructed samples are available. This means that PROF refinement and LIC parameter derivation can be performed simultaneously.

[0124] The LIC weights and offsets (i.e., α and β) and the PROF refinements (i.e., △i[x]) are generally floating-point numbers. For hardware-friendly implementation, these floating-point operations are usually implemented as a multiplication of an integer value followed by a right shift operation and a number of bits. In the current LIC and PROF designs, the two tools are designed separately, so two different right shifts are implemented as N LIC Bits and N PROF Depending on the bit, it is applied in two stages.

[0125] According to a third embodiment of the present disclosure, in order to improve the coding gain when PROF and LIC are jointly applied to an affine coded block, it is proposed to apply LIC-based and PROF-based sample adjustments with high accuracy by combining those two right-shift operations into one and applying it at the end to derive the final predicted sample of the current block (as shown in (12)).

[0126] Addressing the multiplication overflow problem when combining PROF with weighted prediction and bi-prediction with CU-level weights According to the current PROF design in the VVC working draft, PROF can be applied in conjunction with weighted prediction (WP).

[0127] 10 illustrates a method of prediction refinement with optical flow (PROF) for decoding a video signal according to the present disclosure, which can be applied, for example, to a decoder.

[0128] In step 1010, the decoder determines a first reference picture I associated with a video block encoded in an affine mode of the video signal. (0) and the second reference picture I (1) can be obtained.

[0129] In step 1012, the decoder reads the first reference picture I (0) , the second reference picture I (1) The first predicted sample I associated with (0) (i,j) and the second forecast sample I (1) Based on (i,j), first and second horizontal gradient values ​​and vertical gradient values ​​can be obtained.

[0130] In step 1014, the decoder (0) and the second reference picture I (1)CPMV (control point motion vector) associated with Based on the vectors, a first and a second horizontal and vertical motion refinement can be obtained.

[0131] In step 1016, the decoder may derive first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and the first and second horizontal and vertical motion refinements.

[0132] In step 1018, the decoder calculates the first predicted sample I (0) (i,j), the second prediction sample I (1) Based on (i,j) and the first and second prediction refinements, refined first and second samples may be obtained.

[0133] In step 1020, the decoder may obtain a final predicted sample of the video block based on the refined first and second samples by manipulating the refined first and second samples and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for weighted prediction (WP) and parameters for bi-prediction with coding unit (CU) level weights (BCW).

[0134] Specifically, when synthesizing a prediction signal of one affine CU, the signal may be generated in the following procedure.

[0135] 1. For each sample at location (x,y), refine the L0 prediction △I based on PROF. 0 Calculate (x,y) and refine it to the original L0 predicted sample I 0 Add to (x,y), i.e.

number

[0136] 2. For each sample at location (x,y), perform L1 prediction refinement △I based on PROF. 1 Calculate (x,y) and refine it to the original L1 prediction sample I 1 Add to (x,y), i.e.

number

[0137] 3. Combine the refined L0 and L1 prediction samples, i.e.

number

[0138] As can be seen from the above formula, the refinement per sample, i.e., △I 0 (x,y), △I 1 By (x,y), we define the predicted sample after PROF (i.e., I 0 '(x,y) and I 1 'The dynamic range of (x,y) is the original predicted sample (i.e., I 0 (x,y) and I 1 (x,y)) which is one larger than the dynamic range of the signal I(x,y). Assuming that the refined prediction samples are multiplied by the WP and BCW weighting coefficients, it increases the length of the multipliers required. For example, based on the current design, if the intra coding bit depth is in the range of 8 to 12 bits, the predicted signal I 0 (x,y) and I 1 The dynamic range of (x,y) is 16 bits. However, after PROF, the predicted signal I 0 '(x,y) and I 1 'The dynamic range of (x,y) is 17 bits. Therefore, when PROF is applied, it may cause a 16-bit multiplication overflow problem.

[0139] 11 illustrates obtaining a final predicted sample of a video block according to the present disclosure. This method can be applied to a decoder, for example.

[0140] At step 1110, the decoder may obtain a final predicted sample for the video block based on the refined first and second samples and the prediction parameters.

[0141] In step 1112, the decoder may adjust the refined first and second samples by shifting them to the right by a first shift value.

[0142] In step 1114, the decoder may obtain a combined prediction sample by combining the refined first and second samples.

[0143] In step 1116, the decoder may obtain a final predicted sample for the video block by left shifting the combined predicted sample by the first shift value.

[0144] In order to solve such an overflow problem, several methods are proposed below.

[0145] 1. In the first method, it is proposed to disable WP and BCW when PROF is applied to one affine CU.

[0146] 2. In the second method, we apply one clipping operation to the derived sample refinement before adding it to the original predicted sample to obtain the refined predicted sample I 0 '(x,y) and I 1 'The dynamic range of (x,y) is the original prediction sample I 0 (x,y) and I 1 It is proposed to have the same dynamic bit depth as the dynamic bit depth of (x,y). Specifically, by such a method, the sample refinement ΔI 0 (x,y) and △I 1 (x,y) is modified by introducing a clipping operation as shown below.

number

[0147] 3. In the third method, instead of clipping the sample refinement, we propose to clip the refined prediction samples directly so that the refined samples have the same dynamic range as the original prediction samples. Specifically, by the third method, the refined L0 and L1 samples are as follows:

number

[0148] 4. The fourth method proposes to apply a certain right shift to the refined L0 and L1 predicted samples before WP and BCW, and then the final predicted samples are adjusted to the original accuracy by an additional left shift. Specifically, the final predicted samples are derived as follows:

number

[0149] The above methods may be implemented using an apparatus including one or more circuits, including application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus may use the circuits in combination with other hardware or software components to perform the above-mentioned methods. Each module, sub-module, unit, or sub-unit disclosed above may be at least partially implemented using one or more circuits.

[0150] 19 illustrates a computing environment 1910 coupled to a user interface 1960. The computing environment 1910 may be part of a data processing server. The computing environment 1910 includes a processor 1920, a memory 1940, and an I / O interface 1950.

[0151] The processor 1920 typically controls the overall operation of the computing environment 1910, such as operations related to display, data acquisition, data communication, and image processing. The processor 1920 may include one or more processors to execute instructions to perform all or a portion of the steps in the methods described above. Additionally, the processor 1920 may include one or more modules that facilitate interaction between the processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single chip machine, a GPU, or the like.

[0152] The memory 1940 is configured to store various types of data to support the operation of the computing environment 1910. The memory 1940 may include predefined software 1942. Examples of such data include instructions for any application or method operating on the computing environment 1910, video data sets, image data, etc. The memory 1940 may be implemented by using any type of volatile or non-volatile memory device, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk, or a combination thereof.

[0153] The I / O interface 1950 provides an interface between the processor 1920 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1950 may be coupled to an encoder and a decoder.

[0154] In some embodiments, a non-transitory computer readable storage medium is also provided that includes a number of programs, such as those contained in memory 1940, executable by processor 1920 in computing environment 1910 to perform the methods described above. For example, the non-transitory computer readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage, etc.

[0155] The non-transitory computer-readable storage medium stores a plurality of programs that are executed by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, cause the computing device to perform the motion prediction method described above.

[0156] In some embodiments, the computing environment 1910 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0157] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations and alternative embodiments will become apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0158] The examples have been chosen and described to explain the principles of the disclosure and to enable others skilled in the art to understand the disclosure for various implementations and to best utilize the underlying principles and various implementations with various modifications suited to the particular use intended. Therefore, it should be understood that the scope of the disclosure should not be limited to the particular examples of implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the disclosure.

[0159] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to Provisional Application No. 62 / 891,273, filed August 23, 2019, the entire contents of which are incorporated herein by reference in their entirety for all purposes.

Claims

1. A method of prediction refinement with optical flow (PROF) for video coding, comprising: The encoder is Obtaining a video block to be coded by an affine mode; Obtaining a first reference picture and a second reference picture associated with the video block; obtaining first and second horizontal and vertical gradient values ​​based on a first predicted sample I(0)(i,j) and a second predicted sample I(1)(i,j) associated with the first reference picture and the second reference picture; obtaining first and second horizontal and vertical motion refinements based on control point motion vectors (CPMVs) associated with the first and second reference pictures; obtaining first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and the first and second horizontal and vertical motion refinements; obtaining refined first and second samples based on the first prediction sample I(0)(i,j), the second prediction sample I(1)(i,j), and the first and second prediction refinements; obtaining a final predicted sample for the video block based on the refined first and second samples and prediction parameters; The prediction parameters include parameters for weighted prediction (WP) or parameters for bi-prediction with coding unit (CU)-level weight (BCW); the parameters of WP and the parameters of BCW each include a first weight of the refined first sample and a second weight of the refined second sample, A method of prediction refinement with optical flow (PROF).

2. The parameters of the WP further include a first offset and a second offset; The parameters of the BCW further include a third offset.

2. The method of prediction refinement with optical flow (PROF) according to claim 1.

3. Obtaining the first and second prediction refinements includes: obtaining the first and second prediction refinements based on the first and second horizontal and vertical gradient values ​​and first and second horizontal and vertical motion refinements; clipping the first and second prediction refinements based on a prediction refinement threshold.

2. The method of prediction refinement with optical flow (PROF) according to claim 1.

4. the prediction refinement threshold is equal to the coding bit depth plus the maximum of either 1 or 13; The method according to claim 3.

5. Obtaining the refined first and second samples includes: obtaining refined first and second samples based on the first prediction sample I(0)(i,j), the second prediction sample I(1)(i,j), and the first and second prediction refinements; clipping the refined first and second samples based on a refined sample threshold; 2. The method of prediction refinement with optical flow (PROF) of claim 1, comprising:

6. the refined sample threshold is equal to the coding bit depth plus the maximum of either 4 or 16; The prediction refinement with PROF (prediction refinement with (optical flow) method.

7. Obtaining the final predicted sample for the video block based on the refined first and second samples and the prediction parameters includes: Apply only the WP or only the BCW; 2. The method of prediction refinement with optical flow (PROF) of claim 1, comprising:

8. one or more processors; a non-transitory computer-readable storage medium storing instructions executable by the one or more processors; Including, The one or more processors are configured to execute a method according to any one of claims 1 to 7. Computing device.

9. A non-transitory computer-readable storage medium storing a plurality of programs for execution by a computing device including one or more processors, comprising: The plurality of programs, when executed by the one or more processors, A non-transitory computer readable storage medium causing the computing device to perform the method of any one of claims 1 to 7.

10. The bitstream is generated by a method according to any one of claims 1 to 7. Bitstream transmission method.

11. 1. A computer program comprising instructions for storing a bitstream, the computer program comprising: The bitstream comprises encoded video data generated by a method according to any one of claims 1 to 7. Computer program.

Citation Information

Patent Citations

  • Motion-compensation prediction based on BI-directional optical flow

    WO2019010156A1