Method and apparatus for predictive refinement using optical flow

By integrating BDOF and PROF for motion vector prediction in video coding, the method addresses inefficiencies in VVC, enhancing coding efficiency and video quality through refined motion compensation.

JP7813753B2Active Publication Date: 2026-02-13BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023144689
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-07-12
Filing Date
2023-09-06
Publication Date
2026-02-13
Estimated Expiration
2040-07-10

AI Technical Summary

Technical Problem

Existing video coding standards like VVC face inefficiencies in motion vector prediction due to limitations in block-based motion compensation, particularly in handling small motions and complex motions, which affect coding efficiency.

Method used

Integration of bidirectional optical flow (BDOF) and prediction refinement by optical flow (PROF) for decoding video signals, where video blocks are divided into sub-blocks with multiple motion vectors, and motion refinement is applied based on BDOF and PROF, depending on coding mode, to enhance prediction accuracy.

Benefits of technology

Improves coding efficiency by refining motion vectors, reducing residual errors, and enhancing the accuracy of motion compensation, leading to better video quality at lower bitrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007813753000038
    Figure 0007813753000038
  • Figure 0007813753000039
    Figure 0007813753000039
  • Figure 0007813753000040
    Figure 0007813753000040
Patent Text Reader

Abstract

To provide a method, apparatus, and storage medium for motion vector prediction in video coding.SOLUTION: A method for motion vector prediction in video coding includes: dividing a video block to multiple non-overlapped video subblocks; obtaining a first reference picture I(0) and a second reference picture I(1); obtaining first prediction samples I(0)(i,j)'s; obtaining second prediction samples I(1)(i,j)'s; obtaining horizontal and vertical gradient values of the first prediction samples I(0)(i,j)'s and second prediction samples I(1)(i,j)'s; obtaining motion refinements for samples in the video subblock based on a BDOF when the video block is not coded in affine mode; obtaining motion refinements for samples in the video subblock based on a PROF when the video block is coded in affine mode; and obtaining prediction samples of the video block based on the motion refinements.SELECTED DRAWING: Figure 10
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is based on and claims priority to Provisional Application No. 62 / 872,700, filed July 10, 2019, and Provisional Application No. 62 / 873,837, filed July 12, 2019, both of which are incorporated herein by reference in their entirety for all purposes.

[0002] This disclosure relates to video coding and compression, and more particularly to methods and apparatus for two inter-prediction tools being investigated in the Versatile Video Coding (VVC) standard: prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF). [Background technology]

[0003] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Search and Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, etc. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancy present in video images or sequences. An important goal of video coding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]

[0004] Examples of this disclosure provide methods and apparatus for motion vector prediction in video coding. [Means for solving the problem]

[0005] According to a first aspect of the present disclosure, a method for integrating bidirectional optical flow (BDOF) and prediction refinement by optical flow (PROF) for decoding a video signal is provided. The decoder may divide a video block into a plurality of non-overlapping video sub-blocks, and at least one of the plurality of non-overlapping video sub-blocks may be associated with two motion vectors. The decoder may generate a first reference picture I associated with the two motion vectors of the at least one of the plurality of non-overlapping video sub-blocks. (0) and the second reference picture I (1) In display order, the first reference picture I (0) may precede the current picture, and the second reference picture I (1) The decoder uses the first reference picture I (0) The first predicted sample of the video sub-block from the reference block in (0) (i,j)'s, where i and j may represent the coordinates of one sample with the current picture. The decoder may then obtain the second reference picture I (1) The second predicted sample of the video sub-block from the reference block in (1) (i,j)'s. The decoder can obtain the first predicted sample I (0) (i,j)'s and the second prediction sample I (1) (i,j)'s horizontal and vertical gradient values. The decoder may obtain motion refinement for samples within the video sub-block based on BDOF when the video block is not coded in affine mode. The decoder may obtain motion refinement for samples within the video sub-block based on PROF when the video block is coded in affine mode. The decoder may then obtain predicted samples of the video block based on the motion refinement.

[0006] According to a second aspect of the present disclosure, there is provided a BDOF and PROF method for decoding a video signal, the method including: at a decoder, determining a first reference picture I associated with a video block; (0) and the second reference picture I (1) In display order, the first reference picture I (0) may precede the current picture, and the second reference picture I (1) may be after the current picture. In the decoder, the first reference picture I (0) The first predicted sample of a video block from a reference block in (0) (i,j), where i and j represent the coordinates of one sample with the current picture. The method may further include, at the decoder, obtaining a second reference picture I (1) The second predicted sample of video block I from the reference block in (1) (i,j). The method may further include receiving, by the decoder, at least one flag. The at least one flag may be signaled by the encoder in a sequence parameter set (SPS) to signal whether BDOF and PROF are enabled for the current video block. The method, at the decoder, when the at least one flag is enabled, may include obtaining the first predicted sample I when the video block is not coded in affine mode. (0) (i,j) and the second predicted sample I (1) (i,j) to derive a motion refinement for the video block based on the first prediction sample I when the video block is coded in affine mode. (0) (i,j) and the second predicted sample I (1) (i, j). The method may additionally include applying PROF to derive a motion refinement for the video block based on (i, j). The method may also include, at the decoder, obtaining a predictive sample for the video block based on the motion refinement.

[0007] According to a third aspect of the present disclosure, a computing device for decoding a video signal is provided. The computing device may include one or more processors and a non-transitory computer-readable memory storing instructions executable by the one or more processors. The one or more processors may be configured to divide a video block into a plurality of non-overlapping video sub-blocks. At least one of the plurality of non-overlapping video sub-blocks may be associated with two motion vectors. The one or more processors may generate a first reference picture I associated with the two motion vectors of at least one of the plurality of non-overlapping video sub-blocks. (0) and the second reference picture I (1) The first reference picture I in display order may be further configured to be obtained. (0) may precede the current picture, and the second reference picture I (1) may be after the current picture. The one or more processors may (0) The first predicted sample of the video sub-block from the reference block in (0) (i,j)'s, where i and j represent the coordinates of one sample with the current picture. The one or more processors may be further configured to obtain the second reference picture I (1) The second predicted sample of the video sub-block from the reference block in (1) The one or more processors may be further configured to obtain the first predicted sample I(i,j). (0) (i,j)'s and the second prediction sample I (1)(i,j)'s horizontal and vertical gradient values. The one or more processors may be further configured to obtain motion refinement for samples within the video sub-block based on BDOF when the video block is not coded in affine mode. The one or more processors may be further configured to obtain motion refinement for samples within the video sub-block based on PROF when the video block is coded in affine mode. The one or more processors may be further configured to obtain prediction samples of the video block based on the motion refinement.

[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon instructions that, when executed by one or more processors of the apparatus, cause a decoder to: (0) and the second reference picture I (1) In display order, the first reference picture I (0) may precede the current picture, and the second reference picture I (1) The instruction is to, in the decoder, (0) The first predicted sample of a video block from a reference block in (0) (i,j), where i and j represent the coordinates of one sample with the current picture. (1) The second predicted sample of video block I from the reference block in (1)(i,j). The instructions may further cause the apparatus to receive, by the decoder, at least one flag. The at least one flag may be signaled by the encoder in the SPS and signal whether BDOF and PROF are enabled for the current video block. The instructions may further cause, at the decoder, when the at least one flag is enabled, to obtain the first predicted sample I when the video block is not coded in affine mode. (0) (i,j) and the second predicted sample I (1) The instructions may further cause the apparatus to apply BDOF to derive a motion refinement for the video block based on (i, j). (0) (i,j) and the second predicted sample I (1) The instructions may further cause the device to apply PROF to derive a motion refinement for the video block based on (i, j). The instructions may further cause the device, at a decoder, to obtain a predictive sample for the video block based on the motion refinement.

[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]

[0010] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 10 illustrates an example of block partitioning in a multi-type tree structure according to an example of the present disclosure. [Figure 3B] FIG. 10 illustrates an example of block partitioning in a multi-type tree structure according to an example of the present disclosure. [Figure 3C]FIG. 10 illustrates an example of block partitioning in a multi-type tree structure according to an example of the present disclosure. [Figure 3D] FIG. 10 illustrates an example of block partitioning in a multi-type tree structure according to an example of the present disclosure. [Figure 3E] FIG. 10 illustrates an example of block partitioning in a multi-type tree structure according to an example of the present disclosure. [Figure 4] 1 is a diagrammatic illustration of a BDOF model according to an example of the present disclosure; [Figure 5A] 1 is an illustration of an affine model according to an example of the present disclosure. [Figure 5B] 1 is an illustration of an affine model according to an example of the present disclosure. [Figure 6] 1 is an illustration of an affine model according to an example of the present disclosure. [Figure 7] 1 is an illustration of a PROF according to an example of the present disclosure. [Figure 8] 1 is a BDOF workflow according to an example of the present disclosure. [Figure 9] 1 is a workflow of PROF according to an example of the present disclosure. [Figure 10] 1 is a diagram illustrating a method for integrating BDOF and PROF for decoding a video signal according to an example of the present disclosure. [Figure 11] 1 is a BDOF and PROF method for decoding a video signal according to an example of the present disclosure. [Figure 12] 10 is an illustration of a workflow of PROF for bi-prediction according to an example of the present disclosure. [Figure 13] 1 is an illustration of pipeline stages of the BDOF and PROF processes according to the present disclosure. [Figure 14] 1 is an example of a method for deriving a gradient of BDOF according to the present disclosure. [Figure 15] 1 is an example of a method for deriving the gradient of PROF according to the present disclosure. [Figure 16A] 10 is an illustration of deriving template samples for an affine mode according to an example of the present disclosure. [Figure 16B]10 is an illustration of deriving template samples for an affine mode according to an example of the present disclosure. [Figure 17A] 10 is an illustration of exclusively enabling PROF and LIC for affine mode according to an example of the present disclosure. [Figure 17B] 10 is an illustration of jointly enabling PROF and LIC for affine mode according to an example of the present disclosure. [Figure 18A] FIG. 10 illustrates a proposed padding method applied to a 16×16 BDOF CU according to an example of the present disclosure. [Figure 18B] FIG. 10 illustrates a proposed padding method applied to a 16×16 BDOF CU according to an example of the present disclosure. [Figure 18C] FIG. 10 illustrates a proposed padding method applied to a 16×16 BDOF CU according to an example of the present disclosure. [Figure 18D] FIG. 10 illustrates a proposed padding method applied to a 16×16 BDOF CU according to an example of the present disclosure. [Figure 19] FIG. 1 illustrates an exemplary computing environment coupled with a user interface according to an example of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Reference will now be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which like numbers in different drawings, unless otherwise indicated, represent the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure. Instead, these implementations are merely examples of apparatus and methods consistent with aspects related to the present disclosure as recited in the appended claims.

[0012] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or," as used herein, means and includes any and all possible combinations of one or more of the associated listed items.

[0013] While terms such as "first," "second," and "third" may be used herein to describe various pieces of information, it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another category of information. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" may be understood to mean "when" or "upon" or "in response to a judgment," depending on the context.

[0014] The first version of the HEVC standard, which offers approximately 50% bitrate savings or equivalent perceptual quality compared to the previous generation video coding standard H.264 / MPEG AVC, was finalized in October 2013. While the HEVC standard offers significant coding improvements over its predecessor, there is evidence that better coding efficiency than HEVC can be achieved with additional coding tools. Based on this, both VCEG and MPEG have begun work on exploring new coding techniques for future video coding standards. The Joint Video Exploration Team (JVET) was formed in October 2015 by ITU-T VECG and ISO / IEC MPEG to initiate significant research into advanced technologies that could enable significant improvements in coding efficiency. One reference software, called the Joint Exploration Model (JEM), was maintained by JVET by integrating several additional coding tools onto the HEVC Test Model (HM).

[0015] In October 2017, a joint call for proposals (CfP) for video compression with capabilities exceeding HEVC was issued by ITU-T and ISO / IEC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating a compression efficiency gain of approximately 40% over HEVC. Based on these evaluation results, JVET launched a new project to develop a new generation of video coding standard named Versatile Video Coding (VVC). In the same month, a reference software code base called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0016] Like HEVC, VVC is built on a block-based hybrid video coding framework.

[0017] Figure 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows a typical encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.

[0018] In encoder 100, a video frame is partitioned into video blocks for processing. For a given video block, a prediction is formed based on either an inter-prediction technique or an intra-prediction technique.

[0019] A prediction residual, which represents the difference between a current video block, which is part of video input 110, and its predictor, which is part of block predictor 140, is sent from summer 128 to transform 130. The transform coefficients are then sent from transform 130 to quantization 132 for entropy reduction. The quantized coefficients are then provided to entropy coding 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 142 from intra / inter mode decision 116, such as video block partition information, motion vectors (MVs), reference picture indexes, and intra prediction modes, is also provided through entropy coding 138 and stored in compressed bitstream 144. Compressed bitstream 144 comprises the video bitstream.

[0020] Decoder-related circuitry is also required in encoder 100 to reconstruct pixels for prediction purposes. First, a prediction residual is reconstructed through inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with block predictor 140 to generate unfiltered reconstructed pixels for the current video block.

[0021] Spatial prediction (or "intra prediction") uses pixels from already coded samples of neighboring blocks (called reference samples) in the same video frame as the current video block to predict the current video block.

[0022] Temporal prediction (also called "inter-prediction") uses reconstructed pixels from an already coded video picture to predict a current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more MVs that indicate the amount and direction of motion between the current CU and its temporal reference. Furthermore, if multiple reference pictures are supported, one reference picture index is additionally transmitted that is used to identify which reference picture in the reference picture storage the temporal prediction signal comes from.

[0023] Motion estimation 114 takes in signals from video input 110 and picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes in video input 110, signals from picture buffer 120 and a motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.

[0024] After spatial prediction and / or temporal prediction are performed, intra / inter mode decision 116 within encoder 100 chooses the best prediction mode, for example, based on a rate-distortion optimization method. Block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantized residual coefficients are inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Additionally, in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied on the reconstructed CU before it is placed in reference picture storage in picture buffer 120 and used to encode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138 for further compression and packing to form the bitstream.

[0025] Figure 1 shows a block diagram of a typical block-based hybrid video coding system. The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128 x 128 pixels in size. However, unlike HEVC, which partitions blocks solely based on a quadtree, VVC splits a single coding tree unit (CTU) into CUs based on a quadtree, binary tree, or ternary tree to accommodate varying local characteristics. In addition, the concept of multiple partitioning unit types in HEVC is eliminated; that is, the distinction between CUs, prediction units (PUs), and transform units (TUs) no longer exists in VVC. Instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. In the multi-type tree structure, a CTU is first partitioned using a quadtree structure. Then, each quadtree leaf node can be further partitioned using binary and ternary tree structures.

[0026] As shown in Figures 3A, 3B, 3C, 3D, and 3E, there are five split types: quadrant, horizontal bisection, vertical bisection, horizontal third, and vertical third.

[0027] FIG. 3A shows a diagram illustrating block quadrants in a multi-type tree structure according to the present disclosure.

[0028] FIG. 3B shows a diagram illustrating block vertical bisection in a multi-type tree structure according to the present disclosure.

[0029] FIG. 3C shows a diagram illustrating block horizontal bisection in a multi-type tree structure according to the present disclosure.

[0030] FIG. 3D shows a diagram illustrating a vertical third division of blocks in a multi-type tree structure according to the present disclosure.

[0031] FIG. 3E shows a diagram illustrating a horizontal third division of blocks in a multi-type tree structure according to the present disclosure.

[0032] In FIG. 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") uses pixels from samples of already-encoded neighboring blocks (called reference samples) in the same video picture / slice to predict the current video block. Spatial prediction reduces spatial redundancy inherent in video signals. Temporal prediction (also called "inter prediction" or "motion-compensated prediction") uses reconstructed pixels from already-encoded video pictures to predict the current video block. Temporal prediction reduces temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. Also, if multiple reference pictures are supported, one reference picture index is additionally transmitted, which is used to identify which reference picture in the reference picture storage the temporal prediction signal comes from. After spatial prediction and / or temporal prediction, a mode decision block in the encoder chooses the best prediction mode, for example, based on a rate-distortion optimization method. The prediction block is then subtracted from the current video block, and the prediction residual is decorrelated using transform and quantization. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Furthermore, in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and adaptive in-loop filter (ALF), may be applied on the reconstructed CU before it is placed in the reference picture store and used to encode future video blocks. To form an output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit for further compression and packing to form the bitstream.

[0033] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of an exemplary decoder 200. The decoder 200 includes a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.

[0034] The decoder 200 is similar to the reconstruction-related section in the encoder 100 of Figure 1. In the decoder 200, an incoming video bitstream 210 is first decoded through entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 214 and inverse transform 216 to obtain a reconstructed prediction residual. A block predictor mechanism implemented in an intra / inter mode selector 220 is configured to perform either intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block predictor mechanism using a summer 218.

[0035] The reconstructed blocks may further pass through an in-loop filter 228 before being stored in a picture buffer 226, which serves as a reference picture store. The reconstructed video in the picture buffer 226 may be transmitted to drive a display device, as well as used to predict future video blocks. In situations where the in-loop filter 228 is turned on, a filtering operation is performed on these reconstructed pixels to derive a final reconstructed video output 232.

[0036] Figure 2 provides a schematic block diagram of a block-based video decoder. The video bitstream is first entropy decoded in an entropy decoding unit. The coding mode and prediction information are sent to either a spatial prediction unit (if intra-coded) or a temporal prediction unit (if inter-coded) to form a prediction block. The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then summed. The reconstructed block may further pass in-loop filtering before being stored in a reference picture store. The reconstructed video in the reference picture store is then sent to drive a display device and used to predict future video blocks.

[0037] In general, the basic inter-prediction techniques applied in VVC remain the same as those of HEVC, except that some modules are further extended and / or enhanced. In particular, in all previous video standards, one coding block can be associated with only one MV when the coding block is uni-predicted, or only two MVs when the coding block is bi-predicted. Due to such limitations of traditional block-based motion compensation, small motion may still remain in the predicted samples after motion compensation, thus adversely affecting the overall efficiency of motion compensation. To improve both the granularity and accuracy of MVs, two sample-by-sample refinement methods based on optical flow, namely, bidirectional optical flow (BDOF) for affine mode and prediction refinement by optical flow (PROF), are currently being researched for the VVC standard. Below, the main technical aspects of the two inter-coding tools are briefly reviewed.

[0038] Bidirectional Optical Flow In VVC, BDOF is applied to refine the predicted samples of bi-predicted coded blocks. Specifically, as shown in Figure 4, BDOF is a sample-by-sample motion refinement performed on top of block-based motion compensation prediction when bi-prediction is used.

[0039] FIG. 4 illustrates an example of a BDOF model according to the present disclosure.

[0040] Motion refinement for each 4x4 sub-block (v x ,v y ) is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample after BDOF is applied within a 6 × 6 window Ω around the sub-block. Specifically, (v x ,v y ) value is

number

[0041] The values ​​of S1, S2, S3, S5 and S6 are

number

number

number

[0042] Based on the motion refinement derived in (1), the final bi-predictive samples of the CU are

number

[0043] Affine Mode In HEVC, only a translational motion model is applied for motion compensation prediction. Meanwhile, in the real world, there are many types of motion, such as zoom in / out, rotation, perspective motion, and other irregular motion. In VVC, affine motion compensation prediction is applied by signaling one flag per inter-coded block to indicate whether a translational motion model or an affine motion model is applied for inter prediction. In the current VVC design, two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode, are supported for one affine-coded block.

[0044] The four-parameter affine model has the following parameters: two parameters for translational motion in the horizontal and vertical directions, respectively, one parameter for zoom motion in both directions, and one parameter for rotational motion. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To achieve better adaptation of the motion vectors and affine parameters, in VVC, these affine parameters are converted into two MVs (also called control point motion vectors (CPMVs)) located at the upper left and upper right corners of the current block. As shown in Figures 5A and 5B, the affine motion field of a block is described by two control point MVs (V0, V1).

[0045] FIG. 5A shows an example of a four-parameter affine model according to the present disclosure.

[0046] FIG. 5B shows an example of a four-parameter affine model according to the present disclosure.

[0047] Based on the control point motion, the motion field (v x ,v y )teeth,

number

[0048] The 6-parameter affine mode has the following parameters: two parameters for translational motion in the horizontal and vertical directions, one parameter for zoom motion in the horizontal direction and one parameter for rotational motion, and one parameter for zoom motion in the vertical direction and one parameter for rotational motion. The 6-parameter affine motion model is coded using three MVs in three CPMVs.

[0049] FIG. 6 illustrates an example of a six-parameter affine model according to the present disclosure.

[0050] As shown in Figure 6, the three control points of one six-parameter affine block are located at the top-left corner, top-right corner, and bottom-left corner of the block. The motion at the top-left control point is for translational motion, the motion at the top-right control point is for rotational and zooming motion in the horizontal direction, and the motion at the bottom-left control point is for rotational and zooming motion in the vertical direction. Compared with the four-parameter affine motion model, the rotational and zooming motions in the horizontal direction of the six-parameter affine block may not be the same as those in the vertical direction. Assuming (V0, V1, V2) are the MVs of the top-left corner, top-right corner, and bottom-left corner of the current block in Figure 6, the motion vectors (v x ,v y ) is calculated using three MVs at the control points.

number

[0051] Optical flow-based prediction refinement for affine modes To improve the accuracy of affine motion compensation, PROF, which refines sub-block-based affine motion compensation based on an optical flow model, is currently being studied in VVC. Specifically, after implementing sub-block-based affine motion compensation, the luma prediction sample of one affine block is modified by one sample refinement value derived based on the optical flow equation. In detail, the operation of PROF can be summarized as the following four points:

[0052] Step 1: Subblock-based affine motion compensation is performed to generate subblock predictions I(i,j) using the subblock MVs derived in (6) for the 4-parameter affine model and (7) for the 6-parameter affine model.

[0053] 2. The spatial gradient of each predicted sample g x (i,j) and g y (i,j) is

number

[0054] To calculate the gradient, one additional row / column of prediction samples needs to be generated on each side of one sub-block. To reduce memory bandwidth and complexity, samples on the extended boundary are copied from the nearest integer-pixel position in the reference picture to avoid an additional interpolation process.

[0055] Thing 3: The luma prediction refinement value is

number

[0056] FIG. 7 illustrates a PROF process for affine modes according to the present disclosure.

[0057] Since the affine model parameters and pixel locations relative to the sub-block center do not change by sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks within the same CU. Let Δx and Δy be the horizontal and vertical offsets from sample location (i,j) to the center of the sub-block to which the sample belongs, then Δv(i,j) is given by

number

[0058] Based on the affine sub-block MV derivation equations (6) and (7), the MV difference Δv(i,j) can be derived. Specifically, for a four-parameter affine model, The file is TIFF0007813753000015.tif23170.

[0059] For a six-parameter affine model, TIFF0007813753000016.tif44170, where (v 0x ,v 0y ), (v1x ,v 1y ), (v 2x ,v 2y ) are the top-left, top-right, and bottom-left control points MV of the current coding block, and w and h are the width and height of the block. In the existing PROF design, the MV difference Δv x and Δv y is always derived to 1 / 32 pel accuracy.

[0060] Local Lighting Compensation Local illumination compensation (LIC) is a coding tool used to address the problem of local illumination changes that exist between temporally adjacent pictures. A pair of weight and offset parameters is applied to a reference sample to obtain a predicted sample for one current block. The general mathematical model is:

number

number

[0061] In addition to being applied to regular inter-blocks, which contain at most one motion vector per prediction direction (L0 or L1), LIC is also applied to affine-mode coded blocks, where one coded block is further split into multiple smaller sub-blocks, and each sub-block may be associated with different motion information. To derive reference samples for LIC of an affine-mode coded block, as shown in Figures 16A and 16B, the reference samples in the top template of one affine-coded block are fetched using the motion vectors of each sub-block in the top sub-block row, while the reference samples in the left template are fetched using the motion vectors of the sub-blocks in the left sub-block column. Then, the same LLMSE derivation method as shown in (12) is applied to derive LIC parameters based on the composite template.

[0062] 16A shows an example for deriving template samples for affine mode according to the present disclosure. This example includes a Cur Frame 1620 and a Cur CU 1622. The Cur Frame 1620 is the current frame. The Cur CU 1622 is the current coding unit.

[0063] Figure 16B shows an example for deriving template samples for the affine mode. This example includes Ref Frame 1640, Col CU 1642, A Ref 1643, B Ref 1644, C Ref 1645, D Ref 164,6 E Ref 1647, F Ref 1648, and G Ref 1649. Ref Frame 1640 is a reference frame. Col CU 1642 is a co-located coding unit. A Ref 1643, B Ref 1644, C Ref 1645, D Ref 164,6 E Ref 1647, F Ref 1648, and G Ref 1649 are reference samples.

[0064] The terms used in this disclosure are for the purpose of describing illustrative examples only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It should also be understood that the terms "or" and "and / or," as used herein, are intended to mean and include any and all possible combinations of one or more of the associated listed items, unless the context clearly dictates otherwise.

[0065] While terms such as "first," "second," and "third" may be used herein to describe various pieces of information, it should be understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another category of information. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" may be understood to mean "when" or "upon" or "in response to a judgment," depending on the context.

[0066] References throughout this specification to "one example," "an example," "exemplary example," etc. in the singular or plural form mean that one or more particular features, structures, or characteristics described in connection with one example are included in at least one example of the present disclosure. Thus, the appearances of phrases such as "in one example," "in an example," "in an exemplary example," etc. in the singular or plural form in various places throughout this specification are not necessarily all referring to the same example. Furthermore, particular features, structures, or characteristics in one or more examples may be combined in any suitable manner.

[0067] Current BDOF, PROF, and LIC designs Although PROF can improve the coding efficiency of affine modes, its design can still be further improved. In particular, considering the fact that both PROF and BDOF are built on the concept of optical flow, it is highly desirable to harmonize the designs of PROF and BDOF as much as possible so that PROF can make the most of the existing logic of BDOF to facilitate hardware implementation. Based on such considerations, the following inefficiencies regarding the interaction between current PROF and BDOF designs are identified in this disclosure.

[0068] As explained in the section "Optical Flow-Based Prediction Refinement for Affine Modes," in equation (8), the accuracy of the gradient is determined based on the internal bit depth. Meanwhile, the MV difference, i.e., Δv x and Δv yis always derived with 1 / 32-pel precision. Correspondingly, based on equation (9), the precision of the derived PROF refinement depends on the internal bit depth. However, similar to BDOF, to maintain higher PROF derivation precision, PROF is applied on predicted sample values ​​with an intermediate high bit depth (i.e., 16 bits). Therefore, regardless of the internal coding bit depth, the precision of the prediction refinement derived by PROF should match the precision of the intermediate predicted sample, i.e., 16 bits. In other words, the representation bit depth of the MV difference and gradient in the existing PROF design is not perfectly suited to derive accurate prediction refinement compared to the prediction sample precision (i.e., 16 bits). Meanwhile, based on a comparison of equations (1), (4), and (8), the existing PROF and BDOF use different precisions to represent sample gradients and MV differences. As previously pointed out, such a non-integrated design is undesirable for hardware because the existing BDOF logic cannot be reused.

[0069] As discussed in the "Optical Flow-Based Prediction Refinement for Affine Mode" section, when a current affine block is bi-predicted, PROF is applied to the prediction samples in lists L0 and L1 separately, and then the extended L0 and L1 prediction signals are averaged to generate the final bi-predicted signal. Instead of deriving PROF refinement separately for each prediction direction, BDOF derives prediction refinement once, and then the prediction refinement is applied to extend the combined L0 and L1 prediction signals. Figures 8 and 9 (described below) compare the current BDOF and PROF workflows for bi-prediction. In practical codec hardware pipeline designs, different major encoding / decoding modules are usually assigned to each pipeline stage so that more coding blocks can be processed in parallel. However, due to the differences between the BDOF and PROF workflows, this can lead to difficulties in having one and the same pipeline design that can be shared by BDOF and PROF, which is undesirable for practical codec implementations.

[0070] Figure 8 shows a BDOF workflow according to this disclosure. Workflow 800 includes L0 motion compensation 810, L1 motion compensation 820, and BDOF 830. L0 motion compensation 810 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 820 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. BDOF 830 takes motion compensation samples from L1 motion compensation 810 and L1 motion compensation 820 and outputs prediction samples, as described above with respect to Figure 4.

[0071] FIG. 9 shows a workflow of an existing PROF according to this disclosure. Workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0 PROF 930, L1 PROF 940, and average 960. L0 motion compensation 910 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 920 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. L0 PROF 930 takes the L0 motion compensation samples from L0 motion compensation 910 and outputs a motion refinement value, as described above with respect to FIG. 7. L1 PROF 940 takes the L1 motion compensation samples from L1 motion compensation 920 and outputs a motion refinement value, as described above with respect to FIG. 7. Average 960 averages the motion refinement value output of L0 PROF 930 and L1 PROF 940.

[0072] 1) For both BDOF and PROF, gradients need to be calculated for each sample within the current coding block, which requires generating one additional row / column of predicted samples on each side of the block. To avoid the additional computational complexity of sample interpolation, predicted samples in the extended region around the block are directly copied from reference samples at integer positions (i.e., without interpolation). However, according to existing designs, integer samples at different locations are selected to generate gradient values ​​for BDOF and PROF. Specifically, for BDOF, integer reference samples to the left of the predicted sample (for horizontal gradients) and above the predicted sample (for vertical gradients) are used for gradient calculation, while for PROF, integer reference samples closest to the predicted sample are used. Similar to the bit depth representation issue, such a non-integrated gradient calculation method is also undesirable for hardware codec implementation.

[0073] 2) As pointed out earlier, the motivation for PROF is to compensate for small MV differences between the MV of each sample and the sub-block MV derived at the center of the sub-block to which the sample belongs. According to the current PROF design, PROF is always invoked when a coding block is predicted by the affine mode. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is relatively small, the MVs at each sample position should be consistent. In such cases, the benefit of applying PROF may be very limited, and considering the performance / complexity tradeoff, it may not be worthwhile to implement PROF.

[0074] Improved BDOF, PROF, and LIC In this disclosure, a method is provided for improving and simplifying the existing PROF design to facilitate hardware codec implementation. In particular, special attention is paid to harmonizing the BDOF and PROF designs to maximize sharing of existing BDOF logic with PROF. In general, the main aspects of the proposed technology in this disclosure are summarized as follows:

[0075] 1) In order to improve the coding efficiency of PROF while achieving a more integrated design, a method is proposed to integrate the representation bit depths of the sample gradients and MV differences used by BDOF and PROF.

[0076] 2) To facilitate hardware pipeline design, we propose to harmonize the workflow of PROF with that of BDOF for bi-prediction. Specifically, unlike existing PROF, which derives prediction refinement separately for L0 and L1, the proposed method derives prediction refinement applied to the combined L0 and L1 prediction signals at once.

[0077] 3) Two methods are proposed to harmonize the derivation of integer reference samples to calculate the gradient values ​​used by BDOF and PROF.

[0078] 4) To reduce the computational complexity, an early termination method is proposed to adaptively disable the PROF process for affine coded blocks when some conditions are met.

[0079] Improved bit-depth representation design of PROF gradient and MV difference As analyzed in the "Problem Statement" section, the representation bit depths of MV difference and sample gradient in the current PROF are not aligned to derive accurate prediction refinement. Furthermore, the representation bit depths of sample gradient and MV difference are inconsistent between BDOF and PROF, which is unfavorable for hardware. In this section, an improved bit depth representation method is proposed by extending the bit depth representation method of BDOF to PROF. Specifically, in the proposed method, the horizontal gradient and vertical gradient at each sample position are

number

[0080] Additionally, let Δx and Δy be the horizontal and vertical offsets from one sample location to the center of the sub-block to which the sample belongs, expressed with quarter-pel precision. The corresponding PROF MV difference Δv(x,y) at the sample location is

number

[0081] For a six-parameter affine model, TIFF0007813753000022.tif44170, where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the top-left, top-right, and bottom-left control points MV of the current coding block, expressed in 1 / 16 pel precision, and w and h are the width and height of the block.

[0082] In the above discussion, a pair of fixed right shifts is applied to calculate the gradient and MV difference values, as shown in equations (13) and (14). In practice, different bitwise right shifts can be applied to (13) and (14) to achieve various representation precisions of the gradient and MV difference for different tradeoffs between the intermediate calculation precision and the bit width of the internal PROF derivation process. For example, if the input video contains a lot of noise, the derived gradient may not be reliable to represent the true local horizontal gradient value / vertical gradient value at each sample. In such a case, it makes sense to use more bits to represent the MV difference than the gradient. On the other hand, when the input video exhibits stable motion, the MV difference derived by the affine model should be very small. In such a case, using a high-precision MV difference cannot provide the additional benefit of increasing the accuracy of the derived PROF refinement. In other words, in such a case, it is more beneficial to use more bits to represent the gradient value. Based on the above considerations, in one embodiment of the present disclosure, a general method for calculating the gradient and MV difference for PROF is proposed below. Specifically, the horizontal and vertical gradients at each sample location are n a is calculated by applying right shifts to the difference of adjacent predicted samples, i.e.

number

number

number

[0083] In another embodiment of the present disclosure, another PROF bit depth control method is proposed as follows: In this method, the horizontal gradient and vertical gradient at each sample position are still n times the right shift. a The corresponding PROF MV difference Δv(x,y) at a sample position is calculated as follows: It should be calculated as TIFF0007813753000026.tif16170.

[0084] Additionally, to keep the overall PROF derivation at the appropriate internal bit depth, clipping is applied to the derived MV difference as follows: TIFF0007813753000027.tif17170 where limit is where n is the threshold value equal to TIFF0007813753000028.tif11170, and clip3(min,max,x) is a function that clips a given value x inside the range [min,max]. b The value of is 2 max(5,bit-depth-7) Finally, the PROF refinement of the sample is set as follows: Calculated as TIFF0007813753000029.tif10170.

[0085] Harmonized Workflow of BDOF and PROF for Biprediction As discussed previously, when an affine-coded block is bi-predicted, the current PROF is applied unilaterally. More specifically, PROF sample refinements are derived separately and applied to the predicted samples in lists L0 and L1. The refined prediction signals from lists L0 and L1, respectively, are then averaged to generate the final bi-predictive signal for the block. This is in contrast to BDOF designs, where sample refinements are derived and applied to the bi-predictive signal. Such differences between BDOF and PROF bi-predictive workflows may be undesirable for practical codec pipeline designs.

[0086] To facilitate hardware pipeline design, one simplification method according to the present disclosure is to modify the bi-prediction process of PROF so that the workflows of the two prediction refinement methods are harmonized. Specifically, instead of applying refinement separately for each prediction direction, the proposed PROF method derives prediction refinement based on the control point MVs of lists L0 and L1 at once, and then applies the derived prediction refinement to the combined L0 and L1 prediction signals to improve quality. Specifically, based on the MV difference derived in equation (14), the final bi-prediction sample of one affine coding block is calculated by the proposed method as follows:

number

[0087] FIG. 12 shows an example of a PROF process when the proposed bi-predictive PROF method according to this disclosure is applied. PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bi-predictive PROF 1230. L0 motion compensation 1210 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 1220 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. Bi-predictive PROF 1230 takes motion compensation samples from L1 motion compensation 1210 and L1 motion compensation 1220, as described above, and outputs bi-predictive samples.

[0088] Figure 12 illustrates the corresponding PROF process when the proposed bi-predictive PROF method is applied. PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bi-predictive PROF 1230. L0 motion compensation 1210 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 1220 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. Bi-predictive PROF 1230 takes motion compensation samples from L1 motion compensation 1210 and L1 motion compensation 1220, as described above, and outputs bi-predictive samples.

[0089] To demonstrate the potential benefits of the proposed method for hardware pipeline design, Figure 13 shows an example to illustrate the pipeline stages when both BDOF and the proposed PROF are applied. In Figure 13, the decoding process of one inter-block mainly includes three things:

[0090] 1) Parse / decode the MV of the coded block and fetch the reference samples.

[0091] 2) Generating an L0 prediction signal and / or an L1 prediction signal for the coding block.

[0092] 3) Perform sample-wise refinement of the generated bi-predictive samples based on BDOF when the coding block is predicted by one non-affine mode or PROF when the coding block is predicted by an affine mode.

[0093] FIG. 13 illustrates exemplary pipeline stages when both BDOF and the proposed PROF are applied according to the present disclosure. FIG. 13 demonstrates the potential benefits of the proposed method for hardware pipeline design. Pipeline stages 1300 include MV parsing / decoding and reference sample fetching 1310, motion compensation 1320, and BDOF / PROF 1330. Pipeline stages 1300 encode video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block starts at MV parsing / decoding and reference sample fetching 1310, then moves sequentially to motion compensation 1320, then motion compensation 1320, and BDOF / PROF 1330. This means that BLK0 does not begin processing in pipeline stages 1300 until it moves to motion compensation 1320. As time passes from T0 to T1, T2, T3, and T4, it is the same for all stages and video blocks.

[0094] In FIG. 13, the decoding process of one inter block mainly includes three things.

[0095] First, parse / decode the MV of the coded block and fetch the reference samples.

[0096] Second, generate an L0 prediction signal and / or an L1 prediction signal for the coding block.

[0097] Third, perform sample-wise refinement of the generated bi-predictive samples based on BDOF when the coding block is predicted by one non-affine mode or PROF when the coding block is predicted by an affine mode.

[0098] As shown in Figure 13, after the proposed harmonization method is applied, both BDOF and PROF are directly applied to bi-predictive samples. Given that BDOF and PROF are applied to different types of coding blocks (i.e., BDOF is applied to non-affine blocks, and PROF is applied to affine blocks), the two coding tools cannot be invoked simultaneously. Therefore, their corresponding decoding processes can be implemented by sharing the same pipeline stages. This is more efficient than existing PROF designs, which make it difficult to allocate the same pipeline stages to both BDOF and PROF due to the different workflows of bi-prediction.

[0099] In the above discussion, the proposed method only considers the harmonization of BDOF and PROF workflows. However, according to the existing design, the basic operation units of the two coding tools are implemented in different sizes. Specifically, in the case of BDOF, one coding block is W s ×H s where W s =min(W,16) and H s= min(H,16), where W and H are the width and height of the coding block. BODF operations, such as gradient calculation and sample refinement derivation, are performed independently for each subblock. Meanwhile, as previously described, an affine coding block is divided into 4x4 subblocks, and each subblock is assigned one individual MV derived based on either a 4-parameter affine model or a 6-parameter affine model. Since PROF is applied only to affine blocks, its basic operation unit is a 4x4 subblock. Similar to the bi-prediction workflow issue, using a different basic operation unit size for PROF than for BDOF is also unfavorable for hardware implementation, making it difficult for BDOF and PROF to share the same pipeline stage in the overall decoding process. To solve this problem, one embodiment proposes adjusting the subblock size of the affine mode to be the same as the subblock size of the BDOF. Specifically, according to the proposed method, when one coding block is coded in the affine mode, one coding block is coded using W s ×H s where W s =min(W,16) and H s= min(H,16), where W and H are the width and height of the coding block. Each subblock is assigned one individual MV and considered as one independent PROF operation unit. It is worth mentioning that an independent PROF operation unit ensures that the PROF operation on it is performed without referring to information from neighboring PROF operation units. Specifically, the PROF MV difference at one sample position is calculated as the difference between the MV at the sample position and the MV at the center of the PROF operation unit where the sample is located, and the gradient used by the PROF derivation is calculated by padding samples along each PROF operation unit. The claimed benefits of the proposed method mainly include the following aspects: 1) a simplified pipeline architecture with a unified basic operation unit size for both motion compensation and BDOF / PROF refinement; 2) reduced memory bandwidth usage due to the enlarged subblock size for affine motion compensation; and 3) reduced per-sample computation complexity of fractional sample interpolation.

[0100] It should also be mentioned that due to the reduced computational complexity of the proposed method (i.e., item 3), the existing 6-tap interpolation filter constraint for affine-coded blocks can be removed. Instead, the default 8-tap interpolation for non-affine-coded blocks is also used for affine-coded blocks. The overall computational complexity in this case is still comparable to the existing PROF design (based on 4x4 sub-blocks with 6-tap interpolation filters).

[0101] Reconciliation of gradient derivations for BDOF and PROF As explained earlier, both BDOF and PROF calculate the gradient of each sample inside the current coding block, which accesses one additional row / column of predicted samples on each side of the block. To avoid additional interpolation complexity, the required predicted samples in the extension region around the block boundary are directly copied from the integer reference samples. However, as pointed out in the "Problem Statement" section, integer samples at different locations are used to calculate the gradient values ​​of BDOF and PROF.

[0102] To achieve a more unified design, two methods are proposed below to integrate the gradient derivation methods used by BDOF and PROF. In the first method, it is proposed to adjust the gradient derivation method of PROF to be the same as that of BDOF. Specifically, in the first method, the integer positions used to generate predicted samples in the extension region are determined by flooring down the fractional sample positions, i.e., the selected integer sample positions are to the left of the fractional sample positions (for horizontal gradients) and above the fractional sample positions (for vertical gradients).

[0103] In the second method, we propose to adjust the gradient derivation method of BDOF to be the same as that of PROF. More specifically, when the second method is applied, the integer reference sample closest to the predicted sample is used for the gradient calculation.

[0104] Figure 14 shows an example of using the BDOF gradient derivation method according to the present disclosure. In Figure 14, the empty circles represent reference samples at integer positions, the triangles represent fractional predicted samples of the current block, and the grey circles represent integer reference samples used to fill the extension region of the current block.

[0105] Figure 15 shows an example of using the PROF gradient derivation method according to the present disclosure. In Figure 15, the open circles represent reference samples at integer positions, the triangles represent fractional predicted samples of the current block, and the grey circles represent integer reference samples used to fill the extension region of the current block.

[0106] Figures 14 and 15 illustrate the corresponding integer sample locations used to derive gradients for BDOF and PROF when the first method (Figure 12) and the second method (Figure 13) are applied, respectively. In Figures 14 and 15, the empty circles represent reference samples at integer positions, the triangles represent fractional predicted samples of the current block, and the patterned circles represent integer reference samples used to fill the extension region of the current block for gradient derivation.

[0107] In addition, according to the existing BDOF and PROF designs, predicted sample padding is implemented at different coding levels. Specifically, for BDOF, padding is applied along the boundaries of sbWidth×sbHeight sub-blocks, where sbWidth=min(CUWidth, 16) and sbHeight=min(CUHeight, 16). CUWidth and CUHeight are the width and height of one CU. Meanwhile, padding for PROF is always applied at the 4×4 sub-block level. In the above discussion, only the padding method is integrated between BDOF and PROF, but the padding sub-block size is still different. Given that different modules are required to be implemented for the padding processes of BDOF and PROF, this is also undesirable for practical hardware implementation. To achieve a more integrated design, it is proposed to integrate the sub-block padding size of BDOF and PROF. In one embodiment of the present disclosure, it is proposed to apply predicted sample padding for BDOF at the 4×4 level. Specifically, by this method, a CU is first divided into multiple 4x4 sub-blocks, and after motion compensation of each 4x4 sub-block, the extension samples along the top / bottom and left / right boundaries are padded by copying the corresponding integer sample positions.

[0108] Figures 18A, 18B, 18C, and 18D illustrate an example in which the proposed padding method is applied to one 16x16 BDOF CU, where the dashed lines represent 4x4 sub-block boundaries and the gray bands represent padded samples of each 4x4 sub-block.

[0109] FIG. 18A illustrates the proposed padding method applied to a 16x16 BDOF CU according to the present disclosure, where the dashed line represents the top-left 4x4 sub-block boundary 1820.

[0110] FIG. 18B illustrates the proposed padding method applied to a 16x16 BDOF CU according to the present disclosure, where the dashed line represents the top right 4x4 sub-block boundary 1840.

[0111] FIG. 18C illustrates the proposed padding method applied to a 16x16 BDOF CU according to the present disclosure, where the dashed line represents the bottom-left 4x4 sub-block boundary 1860.

[0112] FIG. 18D illustrates the proposed padding method applied to a 16x16 BDOF CU, where the dashed line represents the bottom right 4x4 sub-block boundary 1880, according to the present disclosure.

[0113] 10 illustrates a method for integrating BDOF and PROF for decoding a video signal according to the present disclosure. This method can be applied to, for example, a decoder.

[0114] At 1010, the decoder may divide the video block into a plurality of non-overlapping video sub-blocks. At least one of the plurality of non-overlapping video sub-blocks may be associated with two motion vectors.

[0115] At step 1012, the decoder generates a first reference picture I associated with two motion vectors of at least one of the plurality of non-overlapping video sub-blocks. (0) and the second reference picture I (1) In display order, the first reference picture I (0) is before the current picture, and the second reference picture I (1) is after the current picture.

[0116] At step 1014, the decoder (0) The first predicted sample of the video sub-block from the reference block in (0) (i,j)'s can be obtained, where i and j can represent the coordinates of one sample with the current picture.

[0117] At step 1016, the decoder (1) The second predicted sample of the video sub-block from the reference block in (1)(i,j)'s can be obtained.

[0118] At step 1018, the decoder calculates the first predicted sample I (0) (i,j)'s and the second prediction sample I (1) (i,j)'s horizontal and vertical gradient values ​​can be obtained.

[0119] At 1020, the decoder may obtain motion refinement for samples in the video sub-block based on the BDOF when the video block is not coded in affine mode.

[0120] At 1022, the decoder may obtain motion refinement for samples in the video sub-block based on the PROF when the video block is coded in affine mode.

[0121] At 1024, the decoder may obtain a prediction sample for the video block based on the motion refinement.

[0122] High-level signaling syntax for enabling / disabling BDOF, PROF, and DMVR In the existing BDOF and PROF designs, two different flags are signaled in the SPS to separately control the enabling / disabling of the two coding tools. However, due to the similarity between BDOF and PROF, it is more desirable to enable and / or disable BDOF and PROF from a high level with one and the same control flag. Based on this consideration, one new flag, sps_bdof_prof_enabled_flag, is introduced in the SPS, as shown in Table 1. As shown in Table 1, the enabling and disabling of BDOF depends only on sps_bdof_prof_enabled_flag. When the flag is equal to 1, BDOF is enabled for encoding the video content in the sequence. Otherwise, when sps_bdof_prof_enabled_flag is equal to 0, BDOF is not applied. Meanwhile, in addition to sps_bdof_prof_enabled_flag, an SPS-level affine control flag, namely, sps_affine_enabled_flag, is also used to conditionally enable and disable PROF. PROF is enabled for all coding blocks coded in affine mode when the flags sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are both equal to 1. PROF is disabled when the flags sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are equal to 1 and 0, respectively.

[0123] 11 illustrates a BDOF and PROF method for decoding a video signal according to the present disclosure, which may be applied to, for example, a decoder.

[0124] At step 1110, the decoder extracts the first reference picture I associated with the video block. (0) and the second reference picture I (1) In display order, the first reference picture I (0) is before the current picture, and the second reference picture I (1)is after the current picture.

[0125] At step 1112, the decoder (0) The first predicted sample of a video block from a reference block in (0) (i,j), where i and j may represent the coordinates of one sample with the current picture.

[0126] At step 1114, the decoder (1) The second predicted sample of video block I from the reference block in (1) (i,j) can be obtained.

[0127] At 1116, the decoder may receive at least one flag signaled by the encoder in the SPS that signals whether BDOF and PROF are enabled for the current video block.

[0128] At step 1118, the decoder determines whether the first predicted sample I is a first predicted sample I when the video block is coded in affine mode when at least one flag is enabled. (0) (i,j) and the second predicted sample I (1) PROF may be applied to derive motion refinement for the video block based on (i,j).

[0129] At 1120, the decoder may obtain motion refinement for samples within the video block based on the BDOF applied to the video block.

[0130] At step 1122, the decoder may obtain a prediction sample for the video block based on the motion refinement. [Table 1]

[0131] sps_bdof_prof_enabled_flag specifies whether bidirectional optical flow and optical flow prediction refinement are enabled. When sps_bdof_prof_enabled_flag is equal to 0, both bidirectional optical flow and optical flow prediction refinement are disabled. When sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 1, both bidirectional optical flow and optical flow prediction refinement are enabled. Otherwise (sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0), bidirectional optical flow is enabled and optical flow prediction refinement is disabled.

[0132] sps_bdof_prof_dmvr_slice_preset_flag specifies when the flag slice_disable_bdof_prof_dmvr_flag is signaled at the slice level. When the flag is equal to 1, the syntax slice_disable_bdof_prof_dmvr_flag is signaled for each slice that references the current sequence parameter set. Otherwise (when sps_bdof_prof_dmvr_slice_present_flag is equal to 0), the syntax slice_disabled_bdof_prof_dmvr_flag is not signaled at the slice level. When this flag is not signaled, it is inferred to be 0.

[0133] In addition to the above SPS BDOF / PROF syntax, it is proposed to introduce another control flag at the slice level, namely, slice_disable_bdof_prof_dmvr_flag to disable BDOF, PROF, and DMVR. The SPS flag sps_bdof_prof_dmvr_slice_present_flag, which is signaled in the SPS when either the sps level control flags of DMVR or BDOF / PROF are true, is used to indicate the presence of slice_disable_bdof_prof_dmvr_flag. If present, slice_disable_bdof_dmvr_flag is signaled. Table 2 illustrates a modified slice header syntax table after the proposed syntax is applied. [Table 2]

[0134] Early termination of PROF based on control point MV difference According to the current PROF design, PROF is always invoked when a coding block is predicted by the affine mode. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is relatively small, the MVs at each sample position should be consistent. In such cases, the benefit of applying PROF may be very limited. Therefore, to further reduce the average computation complexity of PROF, it is proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between the sample-wise MVs and the sub-block-wise MVs within a 4x4 sub-block. Since the PROF MV difference values ​​of samples within a 4x4 sub-block are symmetric with respect to the sub-block center, the maximum horizontal and vertical PROF MV differences are calculated based on equation (10):

number

[0135] According to this disclosure, different metrics may be used in determining when the MV difference is small enough to skip the PROF process.

[0136] In one example, based on equation (19), the sum of the absolute maximum horizontal MV difference and the absolute maximum vertical MV difference is less than or equal to one predefined threshold, i.e.,

number

[0137] In another example, When the maximum value of TIFF0007813753000035.tif27170 is less than or equal to the threshold, the PROF process may be skipped.

number

[0138] In addition to the above two examples, the principles of this disclosure are also applicable when other metrics are used in determining whether the MV difference is small enough to skip the PROF process.

[0139] In the above method, PROF is skipped based on the magnitude of the MV difference. Meanwhile, in addition to the MV difference, PROF sample refinement is also calculated based on local gradient information at each sample location within one motion compensation block. For prediction blocks that contain less frequent details (e.g., flat areas), the gradient value tends to be small, and as a result, the derived sample refinement value should be small. Taking this into consideration, according to another embodiment of the present disclosure, it is proposed to apply PROF only to prediction samples of blocks that contain sufficient frequent information.

[0140] Different metrics may be used in determining whether a block contains enough high-frequency information to merit invoking the PROF process for that block. In one example, the decision is based on the average magnitude (i.e., absolute value) of the gradient of the samples in the predicted block. When the average magnitude is less than a threshold, the predicted block is classified as a flat area and PROF should not be applied. Otherwise, the predicted block is deemed to contain enough high-frequency detail that PROF is still applicable. In another example, the maximum magnitude of the gradient of the samples in the predicted block may be used. If the maximum magnitude is less than a threshold, PROF should be skipped for the block. In yet another example, the difference I between the maximum and minimum sample values ​​of the predicted block max -I min may be used to determine whether PROF should be applied to a block. If such difference value is less than a threshold, PROF should be skipped for the block. It is worth noting that the principles of this disclosure are also applicable when some other metric is used in determining whether a given block contains sufficient frequent information.

[0141] Handles the interaction between PROF and LIC for affine modes Because neighboring reconstructed samples (i.e., templates) of the current block are used by LIC to derive linear model parameters, decoding of one LIC-encoded block depends on the complete reconstruction of its neighboring samples. Due to this interdependence, in practical hardware implementations, LIC must be performed in a reconstruction stage where neighboring reconstructed samples are available for LIC parameter derivation. Because block reconstruction must be performed sequentially (i.e., one by one), throughput (i.e., the amount of work per unit time that can be done in parallel) is an important issue to consider when applying other encoding methods to a LIC-encoded block. In this section, two methods are proposed to handle the interaction when both PROF and LIC are enabled for affine modes.

[0142] In a first embodiment of the present disclosure, it is proposed to exclusively apply the PROF mode and the LIC mode to one affine coding block. As previously discussed, in existing designs, a single LIC flag is signaled or inherited at the coding block level to indicate whether the LIC mode is applied to an affine block, while the PROF mode is implicitly applied to all affine blocks without signaling. According to the method of the present disclosure, it is proposed to conditionally apply the PROF mode based on the value of the LIC flag of one affine block. When the flag is equal to 1, only the LIC mode is applied by adjusting the prediction samples of the entire coding block based on the LIC weight and offset. Otherwise (i.e., the LIC flag is equal to 0), the PROF mode is applied to the affine coding block to refine the prediction samples of each sub-block based on an optical flow model.

[0143] FIG. 17A illustrates one exemplary flowchart of the decoding process based on the proposed method in which PROF and LIC are not allowed to be applied simultaneously.

[0144] 17A shows an example of a decoding process based on the proposed method in which PROF and LIC are not allowed according to the present disclosure. The decoding process 1720 includes LIC flag on? 1722, LIC 1724, and PROF 1726. LIC flag on? 1722 determines whether the LIC flag is set, and the following actions are taken according to the determination: LIC 1724 applies LIC if the LIC flag is set; and PROF 1726 applies PROF if the LIC flag is not set.

[0145] In the second embodiment of the present disclosure, it is proposed to apply LIC after PROF to generate prediction samples for one affine block. Specifically, after sub-block-based affine motion compensation is performed, the prediction samples are refined based on PROF sample refinement, and then LIC is applied as follows:

number

[0146] FIG. 17B shows an example of a decoding process in which PROF and LIC are applied according to the present disclosure. The decoding process 1760 includes affine motion compensation 1762, LIC parameter derivation 1764, PROF 1766, and LIC sample adjustment 1768. Affine motion compensation 1762 applies affine motion and is input to LIC parameter derivation 1764 and PROF 1766. LIC parameter derivation 1764 is applied to derive LIC parameters. PROF 1766 is the PROF being applied. LIC sample adjustment 1768 is the LIC weight and offset parameters combined with PROF.

[0147] Figure 17B illustrates an exemplary decoding workflow when the second method is applied. As shown in Figure 17B, since LIC uses a template (i.e., adjacent reconstructed samples) to calculate a LIC linear model, LIC parameters can be derived as soon as adjacent reconstructed samples are available. This means that PROF refinement and LIC parameter derivation can be performed simultaneously.

[0148] The LIC weights and offsets (i.e., α and β) and the PROF refinement (i.e., ΔI[x]) are generally floating-point. In preferred hardware implementations, these floating-point operations are usually implemented as a multiplication with an integer value followed by a right-shift operation by several bits. In existing LIC and PROF designs, the two tools are designed separately, so each has N LIC Bit and N PROF Two different right shifts by bits are applied in two stages.

[0149] According to a third embodiment of the present disclosure, to improve the coding gain when PROF and LIC are jointly applied to an affine-coded block, it is proposed to apply the LIC-based sample adjustment and the PROF-based sample adjustment with high precision by combining the two right-shift operations into one and applying it last to derive the final predicted samples (as shown in (12)) of the current block.

[0150] The above methods may be implemented using an apparatus including one or more circuit configurations, including an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components. The apparatus may use the circuit configurations in combination with other hardware or software components to implement the above-described methods. Each module, sub-module, unit, or sub-unit disclosed above may be implemented at least in part using one or more circuit configurations.

[0151] 19 shows a computing environment 1910 coupled with a user interface 1960. The computing environment 1910 may be part of a data processing server. The computing environment 1910 includes a processor 1920, a memory 1940, and an I / O interface 1950.

[0152] The processor 1920 typically controls the overall operation of the computing environment 1910, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1920 may include one or more processors for executing instructions to perform all or part of the methods described above. Additionally, the processor 1920 may include one or more modules that facilitate interconnection between the processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.

[0153] Memory 1940 is configured to store various types of data to support the operation of computing environment 1910. Memory 1940 may include predetermined software 1942. Examples of such data include instructions for any applications or methods operating on computing environment 1910, video data sets, image data, etc. Memory 1940 may be implemented using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.

[0154] The I / O interface 1950 provides an interface between the processor 1920 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1950 may be coupled to an encoder and a decoder.

[0155] In some embodiments, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs, such as contained in memory 1940, executable by processor 1920 in computing environment 1910 to implement the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.

[0156] A non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the above-described method for motion prediction.

[0157] Here, the computing environment 1910 may be implemented with one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described methods.

[0158] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present disclosure. Many modifications, variations, and alternative implementations will be apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.

[0159] Examples have been chosen and described to explain the principles of the present disclosure and to enable those skilled in the art to understand the disclosure in its various implementations and to best utilize the underlying principles and various implementations with various modifications suited to the particular applications contemplated. It is therefore to be understood that the scope of the present disclosure is not limited to the specific examples of implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the present disclosure.

Claims

1. 1. A method for encoding a video signal, comprising: dividing a video block of a current picture into a plurality of video sub-blocks, at least one of the plurality of video sub-blocks being associated with two motion vectors; determining a first reference picture and a second reference picture associated with the two motion vectors of the at least one of the plurality of video sub-blocks, the first reference picture being before a current picture and the second reference picture being after the current picture in display order; determining a first predicted sample of the video sub-block from the first reference picture; determining a second predicted sample for the video sub-block from the second reference picture; and When a bidirectional optical flow (BDOF) is applied, determining horizontal and vertical gradient values ​​of the first and second predicted samples of the BDOF; determining horizontal and vertical gradient values ​​of the first and second prediction samples of optical flow prediction refinement (PROF) when the PROF is applied; determining motion refinement for samples within the video sub-block according to the horizontal and vertical gradient values ​​of the BDOF when the BDOF is applied, or according to the horizontal and vertical gradient values ​​of the PROF and motion vector differentials when the PROF is applied, wherein the horizontal and vertical gradient values ​​of the BDOF are equal to the horizontal and vertical gradient values ​​of the PROF; determining a prediction sample for the video block based on the motion refinement; and Including, The method comprises: deriving extended prediction samples along left and right boundaries of first and second prediction blocks of the video sub-block by copying from integer reference samples that are nearest to their respective fractional sample positions in the horizontal direction; deriving extended prediction samples along top and bottom boundaries of first and second prediction blocks of the video sub-block by copying from integer reference samples that are nearest to their respective fractional sample positions in a vertical direction.

2. The method described in claim 1, wherein the bit depth of the video data is greater than 12.

3. The method described in claim 1, wherein the motion vector differential is clipped to a symmetric range, the symmetric range including the range [-31, 31].

4. The method of claim 1, wherein the bit depth of the video data is equal to 12.

5. The method of claim 1, wherein a right-shift operation is performed while determining the horizontal and vertical gradient values ​​of the BDOF. a right shift operation is performed while determining the horizontal and vertical gradient values ​​of the PROF; The method of claim 1 , wherein a right shift operation is performed while determining the motion vector differential.

6. 1. A computing device, comprising: one or more processors; a non-transitory computer-readable storage medium storing instructions executable by the one or more processors; A computing device, wherein the one or more processors are configured to perform the method of any one of claims 1 to 5.

7. A computer-readable storage medium storing instructions that, when executed by at least one processor of a computing device, cause the computing device to perform the method of any one of claims 1 to 5.

8. A computer program comprising a plurality of instructions which, when executed by one or more processors, cause said one or more processors to perform the method of any one of claims 1 to 5.

9. A method for transmitting a bitstream, comprising: Executing a method for encoding a video signal according to any one of claims 1 to 5 to generate a bitstream; transmitting said bitstream; A method comprising: