Method and device for refining optical flow prediction
The optical flow prediction refinement method solves the problem of insufficient inter-frame prediction accuracy in the VVC standard. The gradient calculation and motion refinement technology improve the coding efficiency and hardware compatibility, achieving more efficient video coding.
Patent Information
- Application Number
- CN202211693052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-23
- Filing Date
- 2020-08-24
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2040-08-24
AI Technical Summary
Existing video coding and decoding technologies have the problem of insufficient motion compensation accuracy in inter-frame prediction. Especially in the VVC standard, the traditional block-based motion compensation method cannot effectively improve coding efficiency. The bidirectional optical flow and optical flow prediction refinement methods have inconsistencies and incompatibilities in design, which affect hardware implementation.
The optical flow prediction refinement (PROF) method is adopted to obtain the reference picture and prediction sample of the video block, calculate the gradient value and motion refinement, and use weighted prediction and bidirectional prediction technology to prevent multiplication overflow and improve prediction accuracy. It is coordinated with the bidirectional optical flow (BDOF) design to achieve higher coding efficiency.
It improves the accuracy and efficiency of video coding, optimizes hardware implementation, and enhances the accuracy of inter-frame prediction and coding performance.
Smart Images

Figure CN116320473B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 202080059798.0 (application date is August 24, 2020, and invention name is "Method and device for refinement of optical flow prediction"). Technical Field
[0002] The present application relates to video coding and compression. More specifically, the present application relates to methods and apparatus for two inter-frame prediction tools studied in the Versatile Video Coding (VVC) standard, namely, Optical Flow Prediction Refinement (PROF) and Bidirectional Optical Flow (BDOF). Background Art
[0003] Various video codec technologies can be used to compress video data. Video coding and decoding are performed according to one or more video codec standards. For example, video codec standards include Versatile Video Codec (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video coding and decoding generally adopts a prediction method (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilizes the redundancy present in video images or sequences. An important goal of video coding and decoding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the Invention
[0004] Embodiments of the present application provide a method and apparatus for refining optical flow (PROF) prediction in video encoding and decoding.
[0005] According to a first aspect of the present application, there is provided an optical flow prediction refinement (PROF), which is implemented at an encoder. The method may include obtaining a video block encoded by an affine mode; obtaining a first reference picture I associated with the video block; (0) and the second reference picture I (1) The method may further include: (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) to obtain a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value and a second vertical gradient value. The method may also include: (0) and the second reference picture I (1)The method may further include obtaining a first prediction refinement and a second prediction refinement based on the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, the second vertical gradient value and the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement and the second vertical motion refinement. The method may further include obtaining a first prediction refinement and a second prediction refinement based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), the first prediction refinement, and the second prediction refinement to obtain a first refined sample and a second refined sample. The method may further include obtaining a final prediction sample of the video block based on the first refined sample and the second refined sample by manipulating the first refined sample, the second refined sample, and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for weighted prediction (WP) or parameters for weighted bidirectional prediction (BCW) at the coding unit (CU) level.
[0006] According to a second aspect of the present application, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable storage medium storing instructions executed by the one or more processors. The one or more processors may be configured to obtain a video block encoded using an affine mode; obtain a first reference picture associated with the video block; (0) and the second reference picture I (1) The one or more processors may also be configured to generate a first reference image based on the first reference image I (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) obtains a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value, and a second vertical gradient value. The one or more processors may also be configured to obtain a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value, and a second vertical gradient value based on the first reference picture I (0) and the second reference picture I (1) The one or more processors may be configured to obtain a first horizontal motion refinement, a first vertical motion refinement, a second horizontal motion refinement, and a second vertical motion refinement based on the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, the second vertical gradient value, the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement, and the second vertical motion refinement. The one or more processors may be configured to obtain a first prediction refinement and a second prediction refinement based on the first prediction sample I (0) (i, j), the second prediction sample I (1)(i, j), the first prediction refinement, and the second prediction refinement to obtain a first refined sample and a second refined sample. The one or more processors may be further configured to obtain a final prediction sample of the video block based on the first refinement sample and the second refinement sample by manipulating the first refinement sample, the second refinement sample, and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for WP or parameters for BCW.
[0007] According to a third aspect of the present application, a non-transitory computer-readable storage medium is provided, which stores instructions. When executed by one or more processors of a device, the instructions may cause the device to obtain a video block encoded using an affine mode; obtain a first reference picture I associated with the video block; (0) and the second reference picture I (1) The instruction may cause the device to generate a signal based on the first reference image I (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) obtains a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value, and a second vertical gradient value. These instructions can enable the device to obtain a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value, and a second vertical gradient value based on the first reference picture I (0) and the second reference picture I (1) The apparatus may obtain a first horizontal motion refinement, a first vertical motion refinement, a second horizontal motion refinement, and a second vertical motion refinement based on the associated control point motion vector (CPMV). The instructions may cause the apparatus to obtain a first prediction refinement and a second prediction refinement based on the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, the second vertical gradient value, the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement, and the second vertical motion refinement. The instructions may cause the apparatus to obtain a first prediction refinement and a second prediction refinement based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), the first prediction refinement, and the second prediction refinement to obtain a first refined sample and a second refined sample. The instructions may cause the apparatus to obtain a final prediction sample of the video block based on the first refinement sample and the second refinement sample by manipulating the first refinement sample, the second refinement sample, and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for WP or parameters for BCW.
[0008] It is to be understood that both the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples according to the application and, together with the description, serve to explain the principles of the application.
[0010] Figure 1 is a block diagram illustrating an encoder according to an example of the present application.
[0011] Figure 2 is a block diagram illustrating a decoder according to an example of the present application.
[0012] Figure 3A is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0013] Figure 3B is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0014] Figure 3C is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0015] Figure 3D is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0016] Figure 3E is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0017] Figure 4 is a diagram of a bidirectional optical flow (BDOF) model according to an example of the present application.
[0018] Figure 5A is a diagram of an affine model according to an example of the present application.
[0019] Figure 5B is a diagram of an affine model according to an example of the present application.
[0020] Figure 6 is a diagram of an affine model according to an example of the present application.
[0021] Figure 7 is a diagram of optical flow prediction refinement (PROF) according to an example of the present application.
[0022] Figure 8 This is the BDOF workflow according to an example of this application.
[0023] Figure 9 This is the workflow of PROF according to the example of this application.
[0024] Figure 10 is a unified method for BDOF and PROF for decoding video signals according to examples of the present application.
[0025] Figure 11is a method for decoding BDOF and PROF of a video signal according to an example of the present application.
[0026] Figure 12 is an example diagram of the workflow of PROF for bidirectional prediction according to an example of the present application.
[0027] Figure 13 is an example diagram of pipeline stages of the BDOF and PROF processes according to the present application.
[0028] Figure 14 is an example diagram of the gradient derivation method of BDOF according to the present application.
[0029] Figure 15 This is an example diagram of the gradient derivation method of PROF according to the present application.
[0030] Figure 16A is an example diagram of a derivation template sample for an affine mode according to an example of the present application.
[0031] Figure 16B is an example diagram of a derivation template sample for an affine mode according to an example of the present application.
[0032] Figure 17A is an example diagram of exclusively enabling PROF and LIC for affine mode according to an example of the present application.
[0033] Figure 17B is an example diagram of jointly enabling PROF and LIC for affine mode according to an example of the present application.
[0034] Figure 18A is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0035] Figure 18B is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0036] Figure 18C is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0037] Figure 18D is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0038] Figure 19 is a diagram illustrating a computing environment coupled with a user interface according to examples of the present application. DETAILED DESCRIPTION
[0039] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different figures represent the same or similar elements, unless otherwise specified. The embodiments set forth in the following description of exemplary embodiments are not intended to represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with aspects related to the present application as described in the appended claims.
[0040] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0041] It should be understood that although the terms "first," "second," "third," etc. may be used herein to describe various types of information, such information should not be limited by these terms. These terms are merely used to distinguish one type of information from another. For example, first information could be referred to as second information without departing from the scope of the present invention. Similarly, second information could be referred to as first information. As used herein, the term "if" can be understood to mean "when," "upon," or "in response to a determination," depending on the context.
[0042] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceptual quality compared to the previous-generation video codec standard, H.264 / MPEG AVC. While the HEVC standard offers significant codec improvements over its predecessor, there is evidence that even higher coding efficiency can be achieved using additional codec tools compared to HEVC. Building on this momentum, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG established a Joint Video Exploration Team (JVET) to initiate significant research into advanced technologies that could significantly improve coding efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional coding tools on top of the HEVC Test Model (HM).
[0043] In October 2017, ITU-T and ISO / IEC issued a joint call for proposals (CfP) for video compression with capabilities exceeding HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating compression efficiency improvements of approximately 40% over HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard, called the Versatile Video Codec (VVC). That same month, a reference software codebase, called the VVC Test Model (VTM), was established to demonstrate a reference implementation of the VVC standard.
[0044] Like HEVC, VVC is built on a block-based hybrid video codec framework.
[0045] Figure 1 An overall diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1 A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, a block predictor 140, an adder 128, a transform 130, quantization 132, prediction related information 142, intra prediction 118, a picture buffer 120, inverse quantization 134, an inverse transform 136, an adder 126, a memory 124, an in-loop filter 122, entropy coding 138, and a bitstream 144.
[0046] In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.
[0047] The prediction residual representing the difference between the current video block (part of the video input 110) and its predictor (part of the block predictor 140) is sent from the adder 128 to the transform 130. The transform coefficients are then sent from the transform 130 to the quantization 132 for entropy reduction. The quantized coefficients are then fed to the entropy coding 138 to generate the compressed video bitstream. Figure 1 As shown, prediction related information 142 from the intra / inter mode decision 116, such as video block segmentation information, motion vector (MV), reference picture index, and intra prediction mode, is also fed through entropy coding 138 and saved into a compressed bitstream 144. The compressed bitstream 144 includes a video bitstream.
[0048] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through inverse quantization 134 and inverse transformation 136. This reconstructed prediction residual is combined with a block predictor 140 to generate unfiltered reconstructed pixels for the current video block.
[0049] Spatial prediction (or "intra-prediction") uses pixels from samples of already-encoded neighboring blocks in the same video frame as the current video block (called reference samples) to predict the current video block.
[0050] Temporal prediction (also known as "inter prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, a reference picture index is additionally sent, which identifies which reference picture in the reference picture pool the temporal prediction signal comes from.
[0051] Motion estimation 114 receives the video input 110 and a signal from the picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 receives the video input 110, a signal from the picture buffer 120, and a motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0052] After performing spatial and / or temporal prediction, the intra / inter mode decision 116 in the encoder 100 selects the optimal prediction mode based on, for example, a rate-distortion optimization method. The block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using a transform 130 and a quantizer 132. The resulting quantized residual coefficients are inversely quantized by an inverse quantizer 134 and inversely transformed by an inverse transform 136 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal for the CU. Furthermore, in-loop filtering 122 (e.g., a deblocking filter, sample adaptive offset (SAO), and / or an adaptive in-loop filter (ALF)) may be applied to the reconstructed CU before the reconstructed CU is placed in the reference picture memory of the picture buffer 120 and used to encode and decode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to an entropy coder 138 for further compression and packaging to form the bitstream.
[0053] Figure 1A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called a CU). In VTM-1.0, a CU can be up to 128×128 pixels. However, unlike HEVC, which partitions blocks based solely on a quadtree, in VVC, a coding tree unit (CTU) is split into multiple CUs to accommodate different local characteristics based on quadtree / binarytree / ternarytree. In addition, the concept of multiple partition unit types in HEVC is removed, that is, the separation of CU, prediction unit (PU), and transform unit (TU) no longer exists in VVC; instead, each CU is always used as the basic unit for prediction and transform without further partitioning. In the multi-type tree structure, a CTU is first partitioned by a quadtree structure. Then, the leaf nodes of each quadtree can be further partitioned using binary and ternary tree structures.
[0054] like Figure 3A 、 3B As shown in , 3C, 3D, and 3E, there are five types of segmentation, namely, four-way segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal three-way segmentation, and vertical three-way segmentation.
[0055] Figure 3A is a diagram illustrating block quad partitioning in a multi-type tree structure according to the present application.
[0056] Figure 3B is a diagram illustrating vertical binary partitioning of blocks in a multi-type tree structure according to the present application.
[0057] Figure 3C is a diagram illustrating block-level binary partitioning in a multi-type tree structure according to the present application.
[0058] Figure 3D is a diagram illustrating vertical ternary partitioning of blocks in a multi-type tree structure according to the present application.
[0059] Figure 3E is a diagram illustrating block-level ternary partitioning in a multi-type tree structure according to the present application.
[0060] exist Figure 1In the video codec, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples of already coded neighboring blocks in the same video picture / slice (called reference samples) to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion compensated prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture memory the temporal prediction signal comes from. After spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode based on, for example, rate-distortion optimization methods. The prediction block is then subtracted from the current video block, and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are inversely quantized and inversely transformed to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering (such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF)) can be applied to the reconstructed CU before the reconstructed CU is placed in the reference picture memory and used for encoding and decoding future video blocks. In order to form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit to be further compressed and packaged to form the bitstream.
[0061] Figure 2 The overall block diagram of the video decoder for VVC is shown. Specifically, Figure 2 A block diagram of a typical decoder 200 is shown. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction related information 234, and video output 232.
[0062] Decoder 200 is similar to the one residing in Figure 1The decoder 200 first decodes the input video bitstream 210 by entropy decoding 212 to derive quantized coefficient levels and prediction-related information. These quantized coefficient levels are then processed by inverse quantization 214 and inverse transform 216 to obtain reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter mode selector 220 is configured to perform intra prediction 222 or motion compensation 224 based on the decoded prediction information. The unfiltered reconstructed pixel set is obtained by adding the reconstructed prediction residual from the inverse transform 216 to the prediction output generated by the block predictor mechanism using an adder 218.
[0063] Before the reconstructed block is stored in a picture buffer 226, which serves as a reference picture memory, the reconstructed block may also be passed through an in-loop filter 228. The reconstructed video in the picture buffer 226 may be sent to drive a display device and may be used to predict future video blocks. With the in-loop filter 228 turned on, filtering operations are performed on the reconstructed pixels to derive the final reconstructed video output 232.
[0064] Figure 2 A general block diagram of a block-based video decoder is provided. The video bitstream is first entropy decoded in the entropy decoding unit. The codec mode and prediction information are sent to the spatial prediction unit (if intra-frame coding) or the temporal prediction unit (if inter-frame coding) to form the prediction block. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then added together. The reconstructed block may also be in-loop filtered before being stored in the reference picture memory. The reconstructed video in the reference picture memory is then sent to drive the display device and used to predict future video blocks.
[0065] In general, the basic inter prediction techniques applied in VVC remain the same as in HEVC, except that several modules are further extended and / or enhanced. Specifically, for all previous video standards, a coding block can only be associated with a single MV when it is unidirectionally predicted, or with two MVs when it is bidirectionally predicted. Due to this limitation of traditional block-based motion compensation, small motions are still left in the predicted samples after motion compensation, which has a negative impact on the overall efficiency of motion compensation. In order to improve the granularity and accuracy of these MVs, two optical flow-based sample-by-sample refinement methods are currently being studied for the VVC standard, namely Bidirectional Optical Flow (BDOF) and Optical Flow Prediction Refinement for Affine Mode (PROF). The following briefly reviews the main technical aspects of these two inter codec tools.
[0066] Bidirectional optical flow
[0067] In VVC, Bidirectional Optical Flow (BDOF) is applied to refine the prediction samples of the coded blocks for bi - directional prediction. Specifically, as Figure 4 shown, when bi - directional prediction is used, BDOF is a per - sample motion refinement performed on the basis of block - based motion - compensated prediction.
[0068] Figure 4 An example diagram of the BDOF model according to the present application is shown.
[0069] The motion refinement (v x , v y ) of each 4×4 sub - block is calculated by minimizing the difference between the L0 and L1 prediction samples after applying the BDOF within a 6×6 window Ω around the sub - block (v x , v y ). Specifically, the value of (v x , v y ) is derived as:
[0070]
[0071] where, is the floor function; clip3(min, max, x) is a function that clips the given value x to the range [min, max]; the symbol >> represents a bit - by - bit right - shift operation; the symbol << represents a bit - by - bit left - shift operation; th BDOF is a motion - refinement threshold for preventing propagation errors due to irregular local motion, which is equal to 1<<max(5, bit - depth - 7), where bit - depth is the internal bit - depth. In (1),
[0072] The values of S1, S2, S3, S5, and S6 are calculated as:
[0073]
[0074]
[0075]
[0076] <
[0081] θ(i,j)=(I (1) (i,j)>>max(4,bit-depth-8))-(I (0) (i,j)>>max(4,bit-depth-8))(3);
[0082] Among them, I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k (k = 0, 1), which is generated with medium-high precision (i.e., 16 bits); and The horizontal and vertical gradients of a sample are obtained by directly calculating the difference between two adjacent samples of the sample, that is,
[0083]
[0084]
[0085] Based on the motion refinement derived in (1), the final bidirectional prediction samples of the CU are calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow model, as shown in the following formula:
[0086]
[0087] Among them, shift and o offset are the right shift and offset values for merging the L0 and L1 prediction signals for bidirectional prediction, equal to 15-bit depth and 1<<(14-bit depth)+2·(1<<13), respectively. Based on the above bit depth control method, the maximum bit depth of the intermediate parameters in the entire BDOF process is ensured to not exceed 32 bits, and the maximum input of the multiplication is within 15 bits. That is, a 15-bit multiplier is sufficient for BDOF implementation.
[0088] Affine mode
[0089] In HEVC, only the translational motion model is applied to motion-compensated prediction. However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, affine motion-compensated prediction is applied by signaling a flag for each inter-frame coding block to indicate whether the translational motion or affine motion model is applied to inter-frame prediction. In the current VVC design, one affine coding block supports two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode.
[0090] The 4-parameter affine model has the following parameters: two parameters for translation in the horizontal and vertical directions, one parameter for scaling, and one parameter for rotation in both directions. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To better adapt the MV and affine parameters, in VVC, these affine parameters are converted into two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. Figure 5A and Figure 5B As shown, the affine motion field of the block is described by two control points MV(V0, V1).
[0091] Figure 5A A diagram showing a 4-parameter affine model according to the present application is shown.
[0092] Figure 5B A diagram showing a 4-parameter affine model according to the present application is shown.
[0093] Based on the control point motion, an affine encoding block motion field (v x ,v y ) is described as:
[0094]
[0095] The 6-parameter affine mode has the following parameters: two parameters for horizontal and vertical translation, one parameter for horizontal scaling, and one parameter for rotation, and one parameter for vertical scaling and one parameter for rotation. The 6-parameter affine motion model is encoded and decoded using three MVs at three CPMVs.
[0096] Figure 6 A diagram showing a 6-parameter affine model according to the present application is shown.
[0097] like Figure 6 As shown, the three control points of a 6-parameter affine block are located at the top left, top right, and bottom left corners of the block. The motion at the top left control point is associated with translation, the motion at the top right control point is associated with horizontal rotation and scaling, and the motion at the bottom left control point is associated with vertical rotation and scaling. Compared to the 4-parameter affine motion model, the 6-parameter rotation and scaling motion in the horizontal direction may be different from those in the vertical direction. Assume that (V0, V1, V2) is Figure 6 The MVs of the upper left corner, upper right corner and lower left corner of the current block in the MV, then the three MVs at the control point are used to convert the MV of each sub-block (v x ,v y ) is derived as:
[0098]
[0099]
[0100] Refinement of optical flow prediction for affine mode
[0101] To improve the accuracy of affine motion compensation, PROF is currently being studied in VVC, which refines sub-block-based affine motion compensation based on the optical flow model. Specifically, after performing sub-block-based affine motion compensation, the brightness prediction sample of an affine block is modified by a sample refinement value derived based on the optical flow equation. Specifically, the operation of PROF can be summarized in the following four steps:
[0102] Step 1: Perform sub-block based affine motion compensation to generate sub-block predictions I(i,j) using the sub-block MVs derived in (6) for the 4-parameter affine model and in (7) for the 6-parameter affine model.
[0103] Step 2: Spatial gradient g of each prediction sample x (i,j) and g y (i,j) is calculated as follows:
[0104] g x (i,j)=(I(i+1,j)-I(i-1,j))>>(max(2,14-bit-depth)-4)
[0105] g y (i,j)=(I(i,j+1)-I(i,j-1))>>(max (2,14-bit-depth) - 4) (8).
[0106] To compute these gradients, an additional row / column of prediction samples is generated on each side of a subblock. To reduce memory bandwidth and complexity, samples on the extended boundary are copied from the nearest integer pixel position in the reference picture to avoid additional interpolation.
[0107] Step 3: The brightness prediction refinement value is calculated by the following formula:
[0108] ΔI(i,j)= g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (9);
[0109] Where Δv(i,j) is the difference between the pixel MV calculated for the sample position (i,j) (denoted by v(i,j)) and the sub-block MV of the sub-block where the pixel (i,j) is located. In addition, in the current PROF design, after the prediction refinement is added to the original prediction sample, a clipping operation is performed to clip the value of the refined prediction sample to within 15 bits.
[0110] I r (i,j)=I(i,j)+ΔI(i,j)
[0111] I r (i,j)=clip3(-2 14 ,2 14 -1,I r (i,j));
[0112] Among them, I(i,j) and I r (i, j) are the original prediction sample and the refined prediction sample at position (i, j), respectively.
[0113] Figure 7 A diagram showing the PROF process for affine mode according to the present application is shown.
[0114] Since these affine model parameters and the pixel positions relative to the sub-block center do not change from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let Δx and Δy be the horizontal and vertical offsets from the sample position (i,j) to the center of the sub-block to which the sample belongs, then Δv(i,j) can be derived as:
[0115]
[0116] Based on the affine sub-block MV derivation equations (6) and (7), the MV difference Δv(i,j) can be derived. Specifically, for the 4-parameter affine model,
[0117]
[0118] For the 6-parameter affine model,
[0119]
[0120] Among them, (v 0x ,v 0y )、(v 1x ,v 1y )、(v 2x ,v 2y) are the upper left, upper right and lower left control points MV of the current coding block, w and h are the width and height of the block. In the existing PROF design, the MV difference Δv is always derived with an accuracy of 1 / 32 pixel. x and Δv y .
[0121] Local brightness compensation
[0122] Local illumination compensation (LIC) is a codec tool used to address local illumination variations between temporally adjacent images. A pair of weight parameters and offset parameters are applied to reference samples to obtain a predicted sample for the current block. The overall mathematical model is as follows:
[0123] P[x]=α*P r [x+v]+β (11);
[0124] Among them, P r [x+v] is the reference block indicated by the motion vector v, [α , β] is the corresponding weight parameter and offset parameter pair for the reference block, and P[x] is the final predicted block. This pair of weight parameters and offset parameters is estimated based on the template of the current block (i.e., the adjacent reconstructed samples) and the reference block of the template (the reference block is derived using the motion vector of the current block) using the minimum linear mean square error (LLMSE) algorithm. By minimizing the mean square error between these template samples and the reference samples of the template, the mathematical representation of α and β can be derived as follows:
[0125]
[0126]
[0127] Among them, I represents the number of samples in the template, P c [x i ] is the i-th sample of the template of the current block, P r [x i ] is the reference sample of the i-th template sample based on the motion vector v.
[0128] In addition to being applied to regular inter blocks that contain at most one motion vector per prediction direction (L0 or L1), LIC is also applied to affine mode coded blocks, where a coded block is further split into multiple smaller sub-blocks, each of which can be associated with different motion information. In order to derive the reference samples for LIC for affine mode coded blocks, as described below, Figure 16A and 16BAs shown in (12), the reference samples in the top template of an affine coded block are obtained by using the motion vectors of each sub-block in the top sub-block row, while the reference samples in the left template are obtained by using the sub-blocks in the left sub-block column. Afterwards, the same LLMSE derivation method as shown in (12) is applied to derive the LIC parameters based on the composite template.
[0129] Figure 16A 16 shows a diagram for deriving template samples for affine mode according to the present application. The diagram includes CurFrame 1620 and CurCU 1622. CurFrame 1620 is the current frame, and CurCU 1622 is the current coding unit.
[0130] Figure 16B A diagram for deriving template samples for affine mode is shown. The diagram contains Ref Frame 1640, Col CU 1642, A Ref 1643, B Ref 1644, C Ref 1645, D Ref 1646, E Ref 1647, F Ref 1648, and G Ref 1649. Ref Frame 1640 is the reference frame, Col CU 1642 is the co-located coding unit, and A Ref 1643, B Ref 1644, C Ref 1645, D Ref 1646, E Ref 1647, F Ref 1648, and G Ref 1649 are reference samples.
[0131] Shortcomings of optical flow prediction refinement for affine models
[0132] While PROF can improve the encoding and decoding efficiency of affine modes, its design still needs further improvement. In particular, given that both PROF and BDOF are built on the concept of optical flow, it is highly desirable to coordinate the designs of PROF and BDOF as much as possible so that PROF can maximize the use of BDOF's existing logic to facilitate hardware implementation. Based on this consideration, this application identifies the following deficiencies in the interaction between the current PROF and BDOF designs.
[0133] 1. As described in the “Optical flow prediction refinement for affine mode” section, in Equation (8), the accuracy of the gradient is determined based on the internal bit depth. On the other hand, the MV difference is always derived with an accuracy of 1 / 32 pixel, i.e., Δv x and Δv y. Accordingly, based on equation (9), the accuracy of the derived PROF refinement depends on the internal bit depth. However, similar to BDOF, PROF is applied on the prediction sample values of medium and high bit depths (i.e., 16 bits) to maintain higher PROF derivation accuracy. Therefore, regardless of the internal coding bit depth, the prediction refinement accuracy derived by PROF should match the accuracy of the intermediate prediction samples, i.e., 16 bits. In other words, the representation bit depth of MV differences and gradients in the existing PROF design is not fully matched, and accurate prediction refinement relative to the prediction sample accuracy (i.e., 16 bits) cannot be derived. At the same time, based on the comparison of equations (1), (4) and (8), the existing PROF and BDOF use different precisions to represent sample gradients and MV differences. As pointed out earlier, this non-uniform design is not desirable for hardware because the existing BDOF logic cannot be reused.
[0134] 2. As discussed in the "Optical Flow Prediction Refinement for Affine Mode" section, when bidirectionally predicting a current affine block, PROF is applied to the prediction samples in lists L0 and L1 separately; then, the enhanced L0 and L1 prediction signals are averaged to generate the final bidirectional prediction signal. In contrast, BDOF does not derive PROF refinement separately for each prediction direction, but derives prediction refinement once and then applies it to enhance the merged L0 and L1 prediction signals. (As described below) Figure 8 and Figure 9 The current BDOF and PROF workflows for bidirectional prediction are compared. In actual codec hardware pipeline designs, different main encoding / decoding modules are usually assigned to each pipeline stage to enable parallel processing of more coding blocks. However, due to the differences between the BDOF and PROF workflows, this can make it difficult to have a single pipeline design that can be shared by both BDOF and PROF, which is not friendly to actual codec implementations.
[0135] Figure 8 The workflow of BDOF according to the present application is shown. The workflow 800 includes L0 motion compensation 810, L1 motion compensation 820 and BDOF 830. L0 motion compensation 810 can be, for example, a list of motion compensated samples from a previous reference picture. The previous reference picture is a reference picture that was previously from the current picture in the video block. L1 motion compensation 820 can be, for example, a list of motion compensated samples from a next reference picture. The next reference picture is a reference picture that follows the current picture in the video block. BDOF 830 obtains motion compensated samples from L0 motion compensation 810 and L1 motion compensation 820 and outputs predicted samples, as previously described. Figure 4 As described in .
[0136] Figure 9The workflow of the existing PROF according to the present application is shown. The workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0 PROF 930, L1 PROF 940 and averaging 960. L0 motion compensation 910 can be, for example, a list of motion compensated samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 920 can be, for example, a list of motion compensated samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. L0 PROF 930 obtains the L0 motion compensated samples from L0 motion compensation 910 and outputs a motion refinement value, as previously described. Figure 7 L1 PROF 940 takes L1 motion compensation samples from L1 motion compensation 920 and outputs motion refinement values, as previously described in Figure 7 Average 960 averages the motion refinement values output by L0 PROF 930 and L1 PROF 940.
[0137] 3. For BDOF and PROF, gradients need to be calculated for each sample within the current coding block, which requires generating an additional row / column of prediction samples on each side of the block. In order to avoid the additional computational complexity of sample interpolation, the prediction samples in the extended area around the block are copied directly from the reference samples at integer positions (i.e., no interpolation). However, according to the existing design, integer samples at different positions are selected to generate the gradient values for BDOF and PROF. Specifically, for BDOF, integer reference samples located to the left of the prediction sample (horizontal gradient) and above the prediction sample (vertical gradient) are used; for PROF, the integer reference sample closest to the prediction sample is used for gradient calculation. Similar to the bit depth representation problem, this non-uniform gradient calculation method is also undesirable for hardware codec implementation.
[0138] 4. As pointed out earlier, the motivation of PROF is to compensate for small MV differences between the MV of each sample and the sub-block MV derived at the center of the sub-block to which the sample belongs. According to the current PROF design, PROF is always called when predicting a coding block in affine mode. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is small, the MV of each sample position should be consistent. In this case, since the benefit of applying PROF may be very limited, it may not be worthwhile to perform PROF when considering the performance / complexity trade-off.
[0139] Improved refinement of optical flow prediction for affine models
[0140] This application presents a method for improving and simplifying existing PROF designs to facilitate hardware codec implementations. Specifically, special attention is paid to coordinating the design of BDOF and PROF to maximize sharing of existing BDOF logic with PROF. In general, the main aspects of the techniques presented in this application are summarized below.
[0141] 1. In order to improve the coding efficiency of PROF and achieve a more unified design, a method is proposed to unify the representation bit depth of sample gradients and MV differences used by BDOF and PROF.
[0142] 2. To facilitate hardware pipeline design, we propose to coordinate the PROF workflow with the BDOF workflow to achieve bidirectional prediction. Specifically, unlike the existing PROF method, which derives prediction refinements for L0 and L1 separately, the proposed method derives a single prediction refinement, which is applied to the combined L0 and L1 prediction signals.
[0143] 3. Two methods are proposed to coordinate the derivation of integer reference samples to calculate the gradient values used by BDOF and PROF.
[0144] 4. To reduce the computational complexity, an early termination method is proposed to adaptively disable the PROF process for affine coded blocks when certain conditions are met.
[0145] Improved bit depth representation design for PROF gradient and MV difference
[0146] As analyzed in the "Problem Statement" section, the representation bit depths of MV differences and sample gradients in the current PROF are not aligned to obtain accurate prediction refinement. In addition, the representation bit depths of sample gradients and MV differences between BDOF and PROF are inconsistent, which is not hardware-friendly. In this section, an improved bit depth representation method is proposed by extending the bit depth representation method of BDOF to PROF. Specifically, in the proposed method, the horizontal and vertical gradients at each sample position are calculated as:
[0147] g x (i,j)=(I(i+1,j)-I(i-1,j))>>max(6,bit-depth-6)
[0148] g y (i,j)=(I(i,j+1)-I(i,j-1))>>max (6,bit-depth-6) (13).
[0149] Furthermore, assuming that Δx and Δy are the horizontal and vertical offsets from a sample position to the center of the sub-block to which the sample belongs, expressed in 1 / 4 pixel precision, the corresponding PROFMV difference Δv(x,y) at the sample position is derived as:
[0150] Δv x (i,j)=(c*Δx+d*Δy)>>(13-dMvBits)
[0151] Δv y (i,j)=(e*Δx+f*Δy)>>(13-dMvBits) (14);
[0152] Where dMvBits is the bit depth of the gradient value used by the BDOF process, that is, dMvBits = max(5, (bit-depth-7)) + 1. In equations (13) and (14), c, d, e, and f are affine parameters derived based on the affine control point MV. Specifically, for a 4-parameter affine model,
[0153]
[0154] For the 6-parameter affine model,
[0155]
[0156] Among them, ((v 0x ,v 0y )、(v 1x ,v 1y ) and (v 2x ,v 2y ) are the upper left, upper right, and lower left control point MVs of the current coding block, expressed with 1 / 16 pixel precision, and w and h are the width and height of the block.
[0157] In the previous discussion, as shown in equations (13) and (14), a pair of fixed right shifts are applied to calculate these gradient values and MV differences. In practice, for different trade-offs between the intermediate calculation accuracy and the bit width of the internal PROF derivation process, different bit-by-bit right shifts can be applied to (13) and (14) to achieve different representation accuracy of these gradients and MV differences. For example, when the input video contains a lot of noise, the derived gradients may not reliably represent the true local horizontal / vertical gradient values at each sample. In this case, it makes more sense to use more bits to represent MV differences than gradients. On the other hand, when the input video shows stable motion, the MV differences derived from the affine model should be very small. If so, using high-precision MV differences does not provide additional benefits to improve the accuracy of the derived PROF refinement. In other words, in this case, it is more advantageous to use more bits to represent gradient values. Based on the above considerations, in one or more embodiments of the present application, a general method for calculating gradients and MV differences for PROF is proposed below. Specifically, it is assumed that the horizontal and vertical gradients at each sample position are obtained by n a Differences between adjacent prediction samples are calculated by n a It is calculated by right shift, that is,
[0158] g x (i,j)=(I(i+1,j)-I(i-1,j))>>n a
[0159] g y (i,j)=(I(i,j+1)-I(i,j-1))>>n a (15);
[0160] The corresponding PROF MV difference Δv(x,y) at this sample location should be calculated as:
[0161] Δv x (i,j)=(c*Δx+d*Δy)>>(13-n a )
[0162] Δv y (i,j)=(e*Δx+f*Δy)>>(13-n a ) (16);
[0163] Where Δx and Δy are the horizontal and vertical offsets from a sample position to the center of the sub-block to which the sample belongs, expressed in 1 / 4 pixel precision, and c, d, e, and f are parameters derived based on the 1 / 16 pixel affine control point MV. Finally, the final PROF refinement of the sample is calculated as:
[0164] ΔI(i,j)= (g x (i,j)*Δvx (i,j)+g y (i,j)*Δv y (i,j)+1)>>1 (17).
[0165] In another embodiment of the present application, another PROF bit depth control method is proposed as follows. In this method, n a Perform n on the differences between adjacent prediction samples a The horizontal and vertical gradients at each sample position are calculated by right shifting the bits, as in (18). The corresponding PROFMV difference Δv(x,y) at this sample position should be calculated as:
[0166] Δv x (i,j)=(c*Δx+d*Δy)>>(14-n a ),
[0167] Δv y (i,j)=(e*Δx+f*Δy)>>(14-n a ).
[0168] Furthermore, to keep the entire PROF derivation within the appropriate internal bit depth, the derived MV differences are cropped as follows:
[0169] Δv x (i,j)=Clip3(-limit,limit,Δv x (i,j)),
[0170] Δv y (i,j)=Clip3(-limit,limit,Δv y (i,j));
[0171] Among them, limit is equal to clip3(min,max,x) is a function that clips the given value x to the range [min,max]. In an example, n b The value is set to 2 max(5,bit-depth-7) Finally, the PROF refinement of this sample is calculated as:
[0172] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j).
[0173] BDOF and PROF coordinated workflow for bidirectional prediction
[0174] As discussed previously, when an affine coded block is bidirectionally predicted, the current PROF is applied unilaterally. More specifically, these PROF sample refinements are derived separately and applied to the prediction samples in lists L0 and L1. Afterwards, the refined prediction signals from lists L0 and L1 are averaged to generate the final bidirectional prediction signal for the block. This is in contrast to the BDOF design, where these sample refinements are derived and applied to the bidirectional prediction signal. This difference between the bidirectional prediction workflows of BDOF and PROF may not be friendly to practical codec pipeline design.
[0175] To facilitate hardware pipeline design, according to the present application, a simplified approach is to modify the bidirectional prediction process of PROF so that the workflows of the two prediction refinement methods are consistent. Specifically, instead of applying refinement to each prediction direction separately, the proposed PROF method derives a prediction refinement based on the control point MVs of lists L0 and L1; these derived prediction refinements are then applied to the merged L0 and L1 prediction signals to improve quality. Specifically, based on the MV differences derived in equation (14), the final bidirectional prediction samples of an affine coded block are calculated by the proposed method as:
[0176] pred PROF (i,j)=(I (0) (i,j)+I (1) (i,j)+ΔI(i,j)+o offset )>>shift,
[0177] ΔI(i,j)=(g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)+1)>>1
[0178] I r (i,j)=I(i,j)+ΔI(i,j) (18);
[0179] Among them, shift and o offset are the right shift value and offset value used to merge the L0 and L1 prediction signals for bidirectional prediction, which are equal to (15-bit-depth) and 1<<(14-bit-depth)+(2<<13), respectively. In addition, as shown in (18), the cropping operation in the existing PROF design (as shown in (9)) is removed in the proposed method.
[0180] Figure 12The PROF process according to the present application is shown when applying the proposed bidirectional prediction PROF method. PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bidirectional prediction PROF 1230. L0 motion compensation 1210 can be, for example, a list of motion compensated samples from a previous reference picture. The previous reference picture is a reference picture before the current picture in the video block. L1 motion compensation 1220 can be, for example, a list of motion compensated samples from a next reference picture. The next reference picture is a reference picture after the current picture in the video block. As described above, bidirectional prediction PROF 1230 receives motion compensated samples from L0 motion compensation 1210 and L1 motion compensation 1220 and outputs bidirectional prediction samples.
[0181] To demonstrate the potential benefits of the proposed approach for hardware pipeline design, Figure 13 An example is shown to illustrate the pipeline stages when BDOF and the proposed PROF are applied simultaneously. Figure 13 In , the decoding process of an inter-frame block mainly includes three steps:
[0182] 1. Parse / decode the MV of the coded block and obtain the reference sample.
[0183] 2. Generate L0 and / or L1 prediction signals for the coding block.
[0184] 3. Perform sample-by-sample refinement on the generated bidirectional prediction samples based on the BDOF when the coding block is predicted by a non-affine mode or the PROF when the coding block is predicted by an affine mode.
[0185] Figure 13 A diagram showing exemplary pipeline stages when applying BDOF and the proposed PROF according to the present application is shown. Figure 13 This paper demonstrates the potential benefits of the proposed approach for hardware pipeline design. Pipeline stage 1300 includes parsing / decoding MV and acquiring reference samples 1310, motion compensation 1320, and BDOF / PROF 1330. Pipeline stage 1300 will encode video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block will begin by parsing / decoding MV and acquiring reference samples 1310 and move sequentially to motion compensation 1320, motion compensation 1320, and BDOF / PROF 1330. This means that BLK0 does not begin processing in pipeline stage 1300 until it moves to motion compensation 1320. This is true for all stages and video blocks as time passes from T0 to T1, T2, T3, and T4.
[0186] Figure 13 In , the decoding process of an inter-frame block mainly includes three steps:
[0187] First, parse / decode the MV of the coded block and obtain the reference sample.
[0188] Secondly, the L0 and / or L1 prediction signals of the coding block are generated.
[0189] Thirdly, based on the BDOF when the coding block is predicted by a non-affine mode or the PROF when the coding block is predicted by an affine mode, sample-by-sample refinement is performed on the generated bidirectional prediction samples.
[0190] like Figure 13 As shown in Figure 3, after applying the proposed coordination method, both BDOF and PROF are directly applied to bidirectionally predicted samples. Given that BDOF and PROF are applied to different types of coding blocks (i.e., BDOF is applied to non-affine blocks and PROF is applied to affine blocks), the two encoding tools cannot be invoked simultaneously. Therefore, their corresponding decoding processes are performed by sharing the same pipeline stage. This is more efficient than the existing PROF design, in which it is difficult to assign the same pipeline stage to BDOF and PROF due to their different bidirectional prediction workflows.
[0191] In the above discussion, the proposed method only considers the coordination of BDOF and PROF workflows. However, according to the existing design, the basic operation units for these two coding tools are also performed with different sizes. For example, for BDOF, a coding block is split into multiple blocks of size W. s ×H s sub-blocks, where W s =min(W,16),H s =min(H,16), where W and H are the width and height of the coding block, respectively. BODF operations such as gradient calculation and sample refinement derivation are performed independently for each sub-block. On the other hand, as mentioned above, the affine coding block is divided into 4×4 sub-blocks, and each sub-block is assigned a separate MV derived based on a 4-parameter or 6-parameter affine model. Since PROF is only applied to the affine block, its basic operation unit is a 4×4 sub-block. Similar to the bidirectional prediction workflow problem, using different basic operation unit sizes from BDOF to PROF is not friendly to hardware implementation, and makes it difficult for BDOF and PROF to share the same pipeline stage of the entire decoding process. To solve such a problem, in one embodiment, it is proposed to align the sub-block size of the affine mode to be the same as the sub-block size of BDOF.
[0192] Here, according to the proposed method, if a coding block is affine-coded, it will be split into blocks of size W s ×H s sub-blocks, where W s=min(W,16),H s = min(H, 16), where W and H are the width and height of the coding block. Each sub-block is assigned a separate MV and is treated as an independent PROF operation unit. It is worth mentioning that the independent PROF operation unit ensures that the PROF operation performed on it does not need to refer to information from adjacent PROF operation units. Specifically, the PROFMV difference at a sample position is calculated as the difference between the MV at the sample position and the MV at the center of the PROF operation unit where the sample is located; the gradient used for PROF derivation is calculated by filling in samples along each PROF operation unit.
[0193] The benefits of the proposed method mainly include the following aspects: 1) simplified pipeline architecture with a unified basic operation unit size for motion compensation and BDOF / PROF refinement; 2) reduced memory bandwidth usage due to the increased sub-block size for affine motion compensation; 3) reduced per-sample computation complexity of fractional sample interpolation.
[0194] It should also be mentioned that due to the reduced computational complexity using the proposed method (i.e., item 3), the existing 6-tap interpolation filter constraint for affine coded blocks can be removed. Instead, the default 8-tap interpolation used for non-affine coded blocks is also used for affine coded blocks. In this case, the overall computational complexity is still advantageous over the existing PROF design (i.e., based on 4×4 sub-blocks with 6-tap interpolation filters).
[0195] Harmonization of gradient derivation for BDOF and PROF
[0196] As mentioned earlier, both BDOF and PROF compute the gradient for each sample within the current coding block, accessing one additional row / column of prediction samples on each side of the block. To avoid additional interpolation complexity, the required prediction samples in the extended area around the block boundary are copied directly from the integer reference samples. However, as pointed out in the "Problem Statement" section, integer samples at different locations are used to calculate the gradient values for BDOF and PROF.
[0197] In order to achieve a more unified design, two methods are proposed below to unify the gradient derivation methods used by BDOF and PROF. In the first method, it is proposed to align the gradient derivation method of PROF with the gradient derivation method of BDOF. Specifically, with the first method, the integer positions used to generate these prediction samples in the extended area are determined by rounding down the fractional samples, that is, the selected integer sample positions are located to the left of the fractional sample positions (for horizontal gradients) and above the fractional sample positions (for vertical gradients). In the second method, it is proposed to align the gradient derivation method of BDOF with the gradient derivation method of PROF. In more detail, when the second method is applied, the integer reference sample closest to the prediction sample is used for gradient calculation.
[0198] Figure 14 An example of a gradient derivation method using BDOF according to the present application is shown. Figure 14 In , the blank circles represent reference samples at integer positions, the triangles represent fractional prediction samples of the current block, and the gray circles represent integer reference samples used to fill the extended area of the current block.
[0199] Figure 15 An example of a gradient derivation method using PROF according to the present application is shown. Figure 15 In , the blank circles represent reference samples at integer positions, the triangles represent fractional prediction samples of the current block, and the gray circles represent integer reference samples used to fill the extended area of the current block.
[0200] Figure 14 and Figure 15 The results show that when the first method ( Figure 12 ) and the second method ( Figure 13 ) is used for the derivation of the gradients of BDOF and PROF. Figure 14 and Figure 15 In , the open circles represent reference samples at integer positions, the triangles represent the fractional prediction samples of the current block, and the patterned circles represent integer reference samples used to fill the extended area of the current block for gradient derivation.
[0201] In addition, according to the existing BDOF and PROF designs, prediction sample padding is performed at different codec levels. Specifically, for BDOF, padding is applied along the boundary of each sbWidth×sbHeight sub-block, where sbWidth=min(CUWidth,16) and sbHeight=min(CUHeight,16). CU Width and CU Height are the width and height of a CU. On the other hand, PROF padding is always applied at the 4×4 sub-block level. In the above discussion, only the padding method is unified between BDOF and PROF, while the padding sub-block size remains different. This is also not friendly to actual hardware implementation because different modules need to be implemented for the padding process of BDOF and PROF. In order to achieve a more unified design, it is proposed to unify the sub-block padding size of BDOF and PROF. In one embodiment of the present application, it is proposed to apply BDOF prediction sample padding at the 4×4 level. Specifically, with this method, the CU is first divided into multiple 4×4 sub-blocks; after motion compensation is performed on each 4×4 sub-block, the extended samples along the upper / lower and left / right boundaries are filled by copying the corresponding integer sample positions. Figure 18A 、 18B , 18C and 18D show an example of applying the proposed padding method to a 16×16 BDOFCU, where the dotted lines represent 4×4 sub-block boundaries and the gray bands represent the padded samples of each 4×4 sub-block.
[0202] Figure 18A 18 shows the proposed padding method according to the present application applied to a 16×16 BDOF CU, where the dotted line represents the upper left 4×4 sub-block boundary 1820 .
[0203] Figure 18B The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the top right 4×4 sub-block boundary 1840 .
[0204] Figure 18C The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the lower left 4×4 sub-block boundary 1860 .
[0205] Figure 18D The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the bottom-right 4×4 sub-block boundary 1880 .
[0206] Enable / disable advanced signaling syntax for BDOF, PROF, and DMVR
[0207] In the existing BDOF and PROF design, two different flags are signaled in the Sequence Parameter Set (SPS) to control the enablement / disablement of the two coding tools respectively. However, due to the similarities between BDOF and PROF, it is more desirable to enable and / or disable BDOF and PROF from a high level through a common control flag. Based on this consideration, a new flag called sps_bdof_prof_enabled_flag is introduced in the SPS, as shown in Table 1. As shown in Table 1, the enabling and disabling of BDOF depends only on sps_bdof_prof_enabled_flag. When this flag is equal to 1, BDOF is enabled to encode and decode the video content in the sequence. Otherwise, when sps_bdof_prof_enabled_flag is equal to 0, BDOF will not be applied. On the other hand, in addition to sps_bdof_prof_enabled_flag, the SPS-level affine control flag, sps_affine_enabled_flag, is also used to conditionally enable and disable PROF. PROF is enabled for all coded blocks coded in affine mode when both the flags sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are equal to 1. PROF is disabled when the flags sps_bdof_prof_enabled_flag are equal to 1 and sps_affine_enabled_flag are equal to 0.
[0208] Table 1 Modified SPS syntax table with proposed BDOF / PROF enable / disable flags
[0209]
[0210] sps_bdof_prof_enabled_flag specifies whether bidirectional optical flow and optical flow prediction refinement are enabled. When sps_bdof_prof_enabled_flag is 0, both bidirectional optical flow and optical flow prediction refinement are disabled. When sps_bdof_prof_enabled_flag is 1 and sps_affine_enabled_flag is 1, both bidirectional optical flow and optical flow prediction refinement are enabled. Otherwise (sps_bdof_prof_enabled_flag is 1 and sps_affine_enabled_flag is 0), bidirectional optical flow is enabled and optical flow prediction refinement is disabled.
[0211] sps_bdof_prof_dmvr_slice_preset_flag specifies when the flag slice_disable_bdof_prof_dmvr_flag is signaled at the slice level. When this flag is equal to 1, the semantic slice_disable_bdof_prof_dmvr_flag is signaled for each slice that references the current sequence parameter set. Otherwise (when sps_bdof_prof_dmvr_slice_present_flag is equal to 0), the semantic slice_disabled_bdof_prof_dmvr_flag is not signaled at the slice level. When this flag is not signaled, it is inferred to be 0.
[0212] In addition to the above-mentioned SPS BDOF / PROF semantics, it is proposed to introduce another control flag at the slice level, namely, slice_disable_bdof_prof_dmvr_flag to disable BDOF, PROF, and DMVR. The SPS flag sps_bdof_prof_dmvr_slice_present_flag is used to indicate the presence of slice_disable_bdof_prof_dmvr_flag, which is signaled in the SPS when either DMVR or BDOF / PROF SPS level control flags are true. If present, slice_disable_bdof_dmvr_flag is signaled. Table 2 shows the modified slice header semantics table after applying the proposed semantics.
[0213] Table 2 Modified SPS semantic table with proposed BDOF / PROF enable / disable flags
[0214] seq_parameter_set_rbsp(){ if (sps_bdof_prof_dmvr_slice_present_flag) slice_disable_bdof_prof_dmvr_enabled_flag u(1) ……
[0215] Early termination of PROF based on control point MV differences
[0216] According to the current PROF design, PROF is always called when predicting a coding block using affine mode. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from these control point MVs. Therefore, when the difference between the control point MVs is small, the MVs at each sample position should be consistent. In this case, the benefit of applying PROF may be very limited. Therefore, in order to further reduce the average computational complexity of PROF, it is proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between the sample-by-sample MV and the sub-block-by-sub-block MV within a 4×4 sub-block. Since the PROF MV differences of samples within a 4×4 sub-block are symmetric with respect to the sub-block center, the maximum horizontal and vertical PROFMV differences can be calculated based on equation (10) as:
[0217]
[0218]
[0219] Depending on the application, different metrics may be used to determine whether the MV difference is small enough to skip the PROF process.
[0220] In one example, based on equation (19), when the sum of the absolute maximum horizontal MV difference and the absolute maximum vertical MV difference is less than a predefined threshold, the PROF process can be skipped, that is,
[0221]
[0222] In another example, if and If the maximum value is not greater than a threshold, the PROF process can be skipped.
[0223]
[0224] MAX(a,b) is a function that returns the larger value between the input values a and b.
[0225] In addition to the above two examples, the concept of the present application is also applicable to the case where other metrics are used to determine whether the MV difference is small enough to skip the PROF process.
[0226] In the above method, PROF is skipped based on the magnitude of the MV difference. On the other hand, in addition to the MV difference, PROF sample refinement is also calculated based on the local gradient information at each sample position in a motion-compensated block. For prediction blocks containing little high-frequency detail (such as flat areas), the gradient values are often small, so the derived sample refinement value should be small. With this in mind, according to another embodiment of the present application, it is proposed to apply PROF only to prediction samples of blocks containing sufficient high-frequency information.
[0227] Different metrics can be used when determining whether a block contains enough high frequency information to make it worthwhile to call the PROF process for the block. In one example, a decision is made based on the average magnitude (i.e., absolute value) of the gradients of the samples within the prediction block. If the average magnitude is less than a threshold, the prediction block is classified as a flat area and the PROF should not be applied; otherwise, the prediction block is considered to contain enough high frequency details and the PROF is still applicable. In another example, the maximum magnitude of the gradients of the samples within the prediction block can be used. If the maximum magnitude is less than a threshold, the PROF is skipped for the block. In yet another example, the difference between the maximum sample value and the minimum sample value of the prediction block is 1. max -I min It can be used to determine whether to apply the PROF to the block. If the difference is less than a threshold, the PROF is skipped for the block. It is worth noting that the concept of the present application is also applicable to the case where other metrics are used to determine whether a given block contains sufficient high-frequency information.
[0228] Handling the interaction between PROF and LIC for affine mode
[0229] Since the neighboring reconstructed samples (i.e., templates) of the current block are used by LIC to derive the linear model parameters, the decoding of a LIC-coded block depends on the complete reconstruction of its neighboring samples. Due to this interdependence, for practical hardware implementations, it is necessary to perform LIC in the reconstruction phase, where the neighboring reconstructed samples are available for LIC parameter derivation. Because block reconstruction must be performed sequentially (i.e., one after another), throughput (i.e., the amount of work that can be done in parallel per unit time) is an important issue to consider when jointly applying other coding methods to LIC-coded blocks. In this section, two methods are proposed to handle the interaction when PROF and LIC are both enabled for affine mode.
[0230] In the first embodiment of the present application, it is proposed to apply the PROF mode and the LIC mode exclusively to an affine coding block. As previously mentioned, in existing designs, PROF is implicitly applied to all affine blocks without signaling, while a LIC flag is signaled or inherited at the coding block level to indicate whether the LIC mode is applied to an affine block. According to the method of the present application, it is proposed to conditionally apply PROF based on the value of the LIC flag of an affine block. When the flag is equal to 1, only LIC is applied by adjusting the prediction samples of the entire coding block based on the LIC weights and offsets. Otherwise (i.e., the LIC flag is equal to 0), PROF is applied to the affine coding block to refine the prediction samples of each sub-block based on the optical flow model.
[0231] Figure 17A An exemplary flow chart of the decoding process based on the proposed method is shown, where simultaneous application of PROF and LIC is not allowed.
[0232] Figure 17A A diagram illustrates a decoding process based on the proposed method according to the present application, wherein PROF and LIC are disabled. Decoding process 1720 includes steps 1722 (check if LIC flag is on), LIC 1724, and PROF 1726. Step 1722 (check if LIC flag is on) determines whether the LIC flag is set and takes the next step based on this determination. LIC 1724 applies LIC when the LIC flag is set. PROF 1726 applies PROF when the LIC flag is not set.
[0233] In the second embodiment of the present application, it is proposed to apply LIC after PROF to generate prediction samples for an affine block. Specifically, after completing the affine motion compensation based on the sub-block, these prediction samples are refined based on the PROF sample refinement; then, LIC is performed by applying a pair of weights and offsets (derived from the template and its reference samples) to the PROF-adjusted prediction samples to obtain the final prediction samples of the block, as shown below:
[0234] P[x]=α*(P r [x+v]+ΔI[x])+β (22);
[0235] Among them, P r [x+v] is the reference block of the current block indicated by motion vector v; α and β are the LIC weights and offsets; P[x] is the final predicted block; ΔI[x] is the PROF refinement derived in (17).
[0236] Figure 17BA diagram illustrates a decoding process according to the present application in which PROF and LIC are applied. Decoding process 1760 includes affine motion compensation 1762, LIC parameter derivation 1764, PROF 1766, and LIC sample adjustment 1768. Affine motion compensation 1762 applies affine motion and is input to LIC parameter derivation 1764 and PROF 1766. LIC parameter derivation 1764 is used to derive LIC parameters. PROF 1766 applies PROF. LIC sample adjustment 1768 is the LIC weight parameters and offset parameters combined with PROF.
[0237] Figure 17B An exemplary decoding workflow when the second method is applied is shown. Figure 17B As shown in Figure 2, since the LIC uses the template (i.e., adjacent reconstructed samples) to calculate the LIC linear model, the LIC parameters can be derived immediately as soon as the adjacent reconstructed samples are available. This means that PROF refinement and LIC parameter derivation can be performed simultaneously.
[0238] LIC weights and offsets (i.e., α and β) and PROF refinement (i.e., ΔI[x]) are usually floating point numbers. For hardware-friendly implementation, these floating point operations are usually implemented as an integer value multiplication followed by a multiple bit right shift operation. In existing LIC and PROF designs, since these two tools are designed separately, N is applied in two stages. LIC Bits and N PROF Two different right shifts of bits.
[0239] According to the third embodiment of the present application, in order to improve the coding gain when PROF and LIC are jointly applied to affine coded blocks, it is proposed to apply LIC-based and PROF-based sample adjustments with high precision. This is done by combining their two right shift operations into one and applying it at the end to derive the final prediction samples of the current block (as shown in (12)).
[0240] Resolve multiplication overflow issues when combining PROF with weighted prediction and CU-level weighted bidirectional prediction (BCW)
[0241] According to the PROF design in the current VVC working draft, PROF can be applied in conjunction with weighted prediction (WP).
[0242] Figure 10 The present invention provides a method for refining optical flow prediction (PROF) for decoding video signals. For example, the method can be applied to a decoder.
[0243] In step 1010, the decoder may obtain a first reference picture I associated with a video block encoded in an affine mode within a video signal.(0) and the second reference picture I (1) .
[0244] In step 1012, the decoder may generate a decoded image based on the first reference picture I (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) obtains a first horizontal gradient value, a first vertical gradient value, a second horizontal gradient value, and a second vertical gradient value.
[0245] In step 1014, the decoder may generate a decoded image based on the first reference picture I (0) and the second reference picture I (1) The associated control point motion vectors (CPMVs) result in a first horizontal motion refinement, a first vertical motion refinement, a second horizontal motion refinement, and a second vertical motion refinement.
[0246] In step 1016, the decoder may obtain a first prediction refinement and a second prediction refinement based on the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, the second vertical gradient value, the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement, and the second vertical motion refinement.
[0247] In step 1018, the decoder may generate a prediction result based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j) and the first prediction refinement and the second prediction refinement to obtain the first refined sample and the second refined sample.
[0248] In step 1020, the decoder may obtain a final prediction sample of the video block based on the first refinement sample and the second refinement sample by manipulating the first refinement sample, the second refinement sample, and prediction parameters to prevent multiplication overflow. The prediction parameters may include parameters for weighted prediction (WP) and parameters for weighted bidirectional prediction (BCW) at the coding unit (CU) level.
[0249] Specifically, when merging them, a prediction signal of an affine CU can be generated by the following process:
[0250] 1. For each sample at position (x,y), compute the L0 prediction refinement ΔI0(x,y) based on the PROF and add the refinement to the original L0 prediction sample I0(x,y), i.e.,
[0251] ΔI0(x,y)=(g h0 (x,y)·Δv x0 (x,y)+gv0 (x,y)·Δv y0 (x,y)+1)>>1
[0252] I′0(x,y)=I0(x,y)+ΔI0(x,y) (23);
[0253] Among them, I′0(x,y) is the refined sample; g h0 (x,y) and g v0 (x,y) and Δv x0 (x,y) and Δv y0 (x,y) is the L0 horizontal gradient and L0 vertical gradient and L0 horizontal motion refinement and L0 vertical motion refinement at position (x,y).
[0254] 2. For each sample at position (x,y), calculate the L1 prediction refinement ΔI1(x,y) based on the PROF and add this refinement to the original L1 prediction sample I1(x,y), i.e.,
[0255] ΔI1(x,y)=(g h1 (x,y)·Δv x1 (x,y)+g v1 (x,y)·Δv y1 (x,y)+1)>>1
[0256] I′1(x,y)=I1(x,y)+ΔI1(x,y) (24);
[0257] Among them, I′1(x,y) is the refined sample; g h1 (x,y) and g v1 (x,y) and Δv x1 (x,y) and Δv y1 (x,y) is the L1 horizontal gradient and L1 vertical gradient as well as the L1 horizontal motion refinement and L1 vertical motion refinement at position (x,y).
[0258] 3. Combine the refined L0 and L1 prediction samples, i.e.,
[0259] I bi (x,y)=(W0·I′0(x,y)+W1·I′1(x,y)+Offset)>>shift (25);
[0260] Where W0 and W1 are the weights of the WP and BCW; shift and offset are the offset and right shift applied to the weighted average of the L0 and L1 prediction signals, which are used for bidirectional prediction of the WP and BCW. Here, the parameters for the WP include W0, W1, and Offset, while the parameters for the BCW include W0, W1, and shift.
[0261] As can be seen from the equation above, due to the sample-by-sample refinement, namely ΔI0(x,y) and ΔI1(x,y), the predicted samples after PROF (i.e., I′0(x,y) and I′1(x,y)) will have a dynamic range increase of 1 compared to the original predicted samples (i.e., I0(x,y) and I1(x,y)). Since these refined prediction samples will be multiplied by the WP and BCW weighting factors, this will increase the length of the required multipliers. For example, based on the current design, when the intra-coding bit depth is 8 to 12 bits, the dynamic range of the predicted signals I0(x,y) and I1(x,y) is 16 bits. However, after this PROF, the dynamic range of the predicted signals I′0(x,y) and I′1(x,y) is 17 bits. Therefore, when this PROF is applied, 16-bit multiplication overflow may occur.
[0262] Figure 11 The method for obtaining the final prediction sample of the video block according to the present application is shown. For example, the method can be applied to a decoder.
[0263] In step 1112 , the decoder may adjust the first refinement sample and the second refinement sample by right shifting by a first shift value.
[0264] In step 1114 , the decoder may obtain a merged prediction sample by merging the first refinement sample and the second refinement sample.
[0265] In step 1116 , the decoder may obtain the final prediction sample of the video block by left-shifting the merged prediction sample by a first shift value.
[0266] To solve this overflow problem, several methods are proposed:
[0267] 1. In the first approach, it is proposed to disable WP and BCW when applying PROF to an affine CU.
[0268] 2. In the second approach, it is proposed to apply a clipping operation to these derived sample refinements before adding them to the original prediction samples so that the dynamic range of these refined samples I′0(x,y) and I′1(x,y) has the same dynamic bit depth as the original prediction samples I0(x,y) and I1(x,y). Specifically, with this approach, the sample refinements ΔI0(x,y) and ΔI1(x,y) in (23) and (24) are modified by introducing a clipping operation as follows:
[0269] ΔI0(x,y)=clip3(-2 dI-1 ,2 dI-1 -1,ΔI0(x,y))
[0270] ΔI1(x,y)=clip3(-2 dI-1 ,2 dI-1 -1,ΔI1(x,y));
[0271] Where dI = dI base +max(0,BD-12), where BD is the intra-coding bit depth; dI base is the basic bit depth value. In one embodiment, it is proposed to set dI base The value of is set to 14. In another embodiment, it is proposed to set the value to 13.
[0272] 3. In the third method, it is proposed to directly crop these refined prediction samples instead of cropping these sample refinements so that these refined samples have the same dynamic range as the original prediction samples. Specifically, through the third method, these refined L0 and L1 samples will be:
[0273] I′0(x,y)=clip3(-2 dR ,2 dR -1,I0(x,y)+ΔI0(x,y)),
[0274] I′1(x,y)=clip3(-2 dR ,2 dR -1,I1(x,y)+ΔI1(x,y));
[0275] Here, dR=16+max(0, BD-12) (or equivalently max(16, BD+4)), where BD is the intra codec bit depth.
[0276] 4. In the fourth method, it is proposed to apply some right shift to these refined L0 and L1 prediction samples before WP and BCW; then the final prediction samples are adjusted to the original precision by additional left shift. Specifically, the final prediction samples are derived as:
[0277] I bi (x,y)=(W0·(I′0(x,y)>>nb)+W1·(I′1(x,y)>>nb)+Offset)·(shift-nb);
[0278] where nb is the number of additional shifts applied, which can be determined based on the corresponding dynamic range of these PROF sample refinements.
[0279] The methods described above may be implemented using an apparatus comprising one or more circuits, including application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus may use these circuits in combination with other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit disclosed above may be implemented, at least in part, using one or more circuits.
[0280] Figure 19 19 is a diagram illustrating a computing environment coupled with a user interface according to an example of the present application. The computing environment 1910 may be part of a data processing server. The computing environment 1910 includes a processor 1920, a memory 1940, and an input / output (I / O) interface 1950.
[0281] The processor 1920 generally controls the overall operation of the computing environment 1910, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1920 may include one or more processors for executing instructions to perform all or some of the steps in the above-described method. In addition, the processor 1920 may include one or more modules that facilitate interaction between the processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a GPU, etc.
[0282] The memory 1940 is configured to store various types of data to support the operation of the computing environment 1910. The memory 1940 may include predetermined software 1932. Examples of such data include instructions for any application or method operating on the computing environment 1910, video data sets, image data, etc. The memory 1940 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0283] I / O interface 1950 provides an interface between processor 1920 and peripheral interface modules (e.g., keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1950 may be coupled to an encoder and a decoder.
[0284] In some embodiments, a non-transitory computer-readable storage medium including a plurality of programs in, for example, a memory 1940 is also provided, and the plurality of programs can be executed by the processor 1920 in the computing environment 1910 to perform the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0285] The non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, causes the computing device to perform the above-mentioned motion prediction method.
[0286] In some embodiments, the computing environment-1910 can be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-mentioned methods.
[0287] The description of the present application has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the present application. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.
[0288] The examples are chosen and described in order to explain the principles of the present application and to enable others skilled in the art to understand the various embodiments of the present application and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present application is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present application.
Claims
1. A method for refining optical flow prediction (PROF), implemented at the encoder, comprising: Get the video block encoded by affine mode; Obtaining a first reference picture and a second reference picture associated with the video block; Based on first prediction samples I associated with the first reference picture and the second reference picture (0) (i, j) and the second prediction sample I (1) (i, j), obtain the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value and the second vertical gradient value; obtaining a first horizontal motion refinement, a first vertical motion refinement, a second horizontal motion refinement, and a second vertical motion refinement based on control point motion vectors (CPMVs) associated with the first reference picture and the second reference picture; Obtaining a first prediction refinement and a second prediction refinement based on the first horizontal gradient value, the first vertical gradient value, the second horizontal gradient value, and the second vertical gradient value and the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement, and the second vertical motion refinement; Based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), refine the first prediction and refine the second prediction to obtain a first refined sample and a second refined sample; A final prediction sample of the video block is obtained based on the first refined sample, the second refined sample, and prediction parameters, wherein the prediction parameters include parameters for weighted prediction (WP) or parameters for weighted bidirectional prediction (BCW) at a coding unit (CU) level.
2. The method according to claim 1, further comprising: When the PROF is applied to an affine CU, the WP and the BCW are disabled.
3. The method according to claim 1, wherein Obtaining the first prediction refinement and the second prediction refinement includes: obtaining the first prediction refinement and the second prediction refinement based on the first horizontal gradient value, the second vertical gradient value, the second horizontal gradient value, and the second vertical gradient value and the first horizontal motion refinement, the first vertical motion refinement, the second horizontal motion refinement, and the second vertical motion refinement; and The first prediction refinement and the second prediction refinement are clipped based on a prediction refinement threshold.
4. The method according to claim 3, wherein: The prediction refinement threshold is equal to the maximum of the encoding bit depth plus 1 or 13.
5. The method according to claim 1, wherein Obtaining the first refined sample and the second refined sample includes: Based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), refine the first prediction and refine the second prediction, and obtain the first refined sample and the second refined sample; and The first refined samples and the second refined samples are clipped based on a refined sample threshold.
6. The method according to claim 5, wherein: The refinement sample threshold is equal to the maximum of the encoding bit depth plus 4 or 16.
7. The method according to claim 1, wherein Obtaining the final prediction sample of the video block includes: adjusting the first refined sample and the second refined sample by right shifting by a first shift value; obtaining a merged prediction sample by merging the first refined sample and the second refined sample; and The final prediction sample of the video block is obtained by left-shifting the merged prediction sample by the first shift value.
8. The method according to claim 1, wherein Obtaining the final prediction sample of the video block based on the first refined sample, the second refined sample, and the prediction parameter includes: Apply only the WP or only the BCW.
9. A computing device comprising: one or more processors; A non-transitory computer-readable storage medium storing instructions for execution by the one or more processors, wherein when executed by the one or more processors, the instructions cause the computing device to perform the method of any one of claims 1-8.
10. A non-transitory computer-readable storage medium storing a plurality of programs executed by a computing device having one or more processors, wherein: When executed by the one or more processors, the plurality of programs cause the computing device to perform the method of any one of claims 1-8.
Citation Information
Patent Citations
Method, device and terminal for determining pushed video type
CN108521609A
Method and apparatus for deriving VR projection, packing, roi and viewport related tracks in isobmff and supporting viewport roll signaling
WO2018171758A1