Method and apparatus for prediction improvement using optical flow
By harmonizing the bit-depth representation and workflow of BDOF and PROF, the video coding standards address inefficiencies in motion compensation, enhancing coding efficiency and simplifying hardware implementation for improved video compression.
Patent Information
- Application Number
- JP2025239340
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-04-25
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-16
AI Technical Summary
Existing video coding standards like VVC face inefficiencies due to limitations in motion compensation, particularly with small motions in bi-predicted blocks, and the separate pipeline designs of bidirectional optical flow (BDOF) and prediction improvement using optical flow for affine mode (PROF) complicate hardware implementation.
Harmonize the bit-depth representation and workflow of BDOF and PROF to unify their logic, ensuring consistent precision for sample gradients and motion vector differences, and introduce an early termination method for PROF processing to reduce complexity.
Enhances coding efficiency by aligning BDOF and PROF designs, facilitating hardware implementation and improving motion compensation accuracy, thereby optimizing video compression.
Smart Images

Figure 2026026398000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 62 / 838,939, filed April 25, 2019, the entire contents of which are incorporated herein by reference.
[0003] This disclosure relates to video coding and compression. More specifically, this disclosure relates to methods and apparatus based on two inter-prediction tools being investigated in the versatile video coding (VVC) standard: prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF). [Background technology]
[0004] Various video coding techniques may be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, etc. Video coding generally utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit redundancy present in video pictures or video sequences. The focus of video coding techniques is to compress video data into a format that uses a lower bitrate while avoiding or minimizing degradation of video quality. Summary of the Invention [Problem to be solved by the invention]
[0005] The examples of this disclosure provide a method and apparatus for prediction-improved bit-depth representation using optical flow. [Means for solving the problem]
[0006] According to a first aspect of the present disclosure, a bit-depth representation method for prediction improvement using optical flow (PROF) for decoding a video signal is provided. The method may include obtaining a first reference picture associated with a video block in the video signal and a first motion vector (MV) from the video block in a current picture to a reference block in the first reference picture. The first reference picture may include multiple non-overlapping video blocks, and at least one video block may be associated with at least one MV. The method may also include obtaining a first prediction sample I(i,j) of the video block generated from the reference block in the first reference picture. i and j may represent coordinates of a sample comprising the video block. The method may include controlling an internal bit-depth of an internal PROF parameter. The internal PROF parameter may include a horizontal gradient value, a vertical gradient value, a horizontal motion differential, and a vertical motion differential derived for the prediction sample I(i,j). The method may further include obtaining a prediction improvement value for the first prediction sample I(i,j) based on horizontal and vertical gradient values and horizontal and vertical motion differentials. When the video block may include a second MV, the method may further include obtaining a prediction improvement value for the second prediction sample I(i,j) associated with the second MV. The method may include obtaining a final predicted sample for the video block based on a combination of the first predicted sample I'(i,j), the second predicted sample I'(i,j), and the prediction improvement value.
[0007] According to a second aspect of the present disclosure, there is provided a bidirectional optical flow (BDOF) bit depth representation method for decoding a video signal, the method including: (0) and the second reference picture I (1) In display order, the first reference picture I (0) can be the previous picture of the current picture, and the second reference picture I (1) can be after the current picture. This method uses the first reference picture I (0) The first predicted sample of the video block from the reference block in (0) (i, j), where i and j may represent the coordinates of one sample with the current picture. (1) The second predicted sample of the video block from the reference block in (1) (i,j). The method may include obtaining a first prediction sample I (0) (i,j) and the second predicted sample I (1) (i,j). The method may include applying BDOF to the video block based on the padded prediction samples. (0) (i,j) and the second predicted sample I (1) The method may include obtaining horizontal and vertical gradient values of (i,j). The method may further include obtaining a motion refinement of samples in the video block based on the BDOF and the horizontal and vertical gradient values applied to the video block. The method may include obtaining bi-predictive samples of the video block based on the motion refinement.
[0008] According to a third aspect of the present disclosure, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable memory storing instructions executable by the one or more processors. The one or more processors may be configured to obtain a first reference picture associated with a video block in a video signal and a first motion vector (MV) from the video block in a current picture to the reference block in the first reference picture. The first reference picture may include multiple non-overlapping video blocks, and at least one video block may be associated with at least one motion vector. The one or more processors may also be configured to obtain a first prediction sample I(i,j) of the video block generated from the reference block in the first reference picture, where i and j represent coordinates of a sample comprising the video block. The one or more processors may be configured to control an internal bit depth of an internal PROF parameter. The internal PROF parameter may include a horizontal gradient value, a vertical gradient value, a horizontal motion differential, and a vertical motion differential derived for the prediction sample I(i,j). The one or more processors may also be configured to obtain a prediction improvement value for the first predicted sample I(i,j) based on the horizontal and vertical gradient values and the horizontal and vertical motion differentials. When the video block may include a second MV, the one or more processors may also be configured to obtain a second predicted sample I′(i,j) associated with the second MV and a corresponding prediction improvement value for the second predicted sample I′(i,j). The one or more processors may be configured to obtain a final predicted sample of the video block based on a combination of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value.
[0009] According to a fourth aspect of the present disclosure, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable memory that stores instructions executable by the one or more processors. The one or more processors may generate a first reference picture I associated with a video block. (0) and the second reference picture I (1) In the display order, the first reference pixel Kucha I (0) can be the previous picture of the current picture, and the second reference picture I (1) may be after the current picture. The one or more processors may (0) The first predicted sample of the video block from the reference block in (0) The one or more processors may also be configured to obtain the second reference picture I (i, j), where i and j may represent the coordinates of one sample with the current picture. (1) The second predicted sample of the video block from the reference block in (1) The one or more processors may be configured to obtain the first predicted sample I(i,j). (0) (i,j) and the second predicted sample I (1) (i,j). The one or more processors may be configured to apply BDOF to the video block based on the padded predicted samples. (0) (i,j) and the second predicted sample I (1) The one or more processors may be configured to obtain horizontal and vertical gradient values of (i, j). The one or more processors may be further configured to obtain a motion refinement of samples in the video block based on the BDOF and the horizontal and vertical gradient values applied to the video block. The one or more processors may be configured to obtain bi-predictive samples of the video block based on the motion refinement.
[0010] It should be understood that the foregoing summary and the following detailed description are exemplary only and are not limiting of the present disclosure.
[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 2 is a block diagram of an encoder according to an example of the present disclosure. [Figure 2] FIG. 2 is a block diagram of a decoder according to an example of the present disclosure. [Figure 3A] FIG. 1 illustrates block division in a complex tree structure according to an example of the present disclosure. [Figure 3B] FIG. 1 illustrates block division in a complex tree structure according to an example of the present disclosure. [Figure 3C] FIG. 1 illustrates block division in a complex tree structure according to an example of the present disclosure. [Figure 3D] FIG. 1 illustrates block division in a complex tree structure according to an example of the present disclosure. [Figure 3E] FIG. 1 illustrates block division in a complex tree structure according to an example of the present disclosure. [Figure 4] FIG. 1 illustrates a bidirectional optical flow (BDOF) model according to an example of the present disclosure. [Figure 5A] FIG. 1 illustrates an affine model according to an example of the present disclosure. [Figure 5B] FIG. 1 illustrates an affine model according to an example of the present disclosure. [Figure 6] FIG. 1 illustrates an affine model according to an example of the present disclosure. [Figure 7] FIG. 1 illustrates prediction improvement using optical flow (PROF) according to an example of the present disclosure. [Figure 8] 1 is a BDOF workflow according to an example of the present disclosure. [Figure 9] 1 is a workflow of PROF according to an example of the present disclosure. [Figure 10] FIG. 1 is a diagram of a bit depth representation method for PROF according to the present disclosure. [Figure 11] FIG. 1 is a diagram of a bit depth representation method for BDOF according to the present disclosure. [Figure 12] FIG. 10 is a diagram illustrating a workflow of PROF for bi-prediction according to an example of the present disclosure. [Figure 13] FIG. 1 illustrates pipeline stages for BDOF and PROF processing according to the present disclosure. [Figure 14] FIG. 1 illustrates a gradient derivation method for BDOF according to the present disclosure. [Figure 15] FIG. 1 illustrates a gradient derivation method for PROF according to the present disclosure. [Figure 16] FIG. 1 illustrates a computing environment coupled with a user interface according to an example of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which the same numbers in different drawings represent the same or similar elements unless otherwise stated. The implementations set forth in the following description of embodiments do not represent all implementations consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present disclosure as set forth in the appended claims.
[0014] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly dictates otherwise. It is also to be understood that the term "and / or," as used herein, is intended to represent and include any and all possible combinations of the associated one or more listed items.
[0015] As used herein, terms such as "first," "second," and "third" may be used to describe various pieces of information, but it is understood that the information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, first information may be referred to as second information, and similarly, second information may be referred to as first information, without departing from the scope of this disclosure. As used herein, the term "if" may be understood to mean "when," "upon," or "in response to a determination that," depending on the context.
[0016] The first version of the HEVC standard, finalized in October 2013, offers approximately 50% bitrate savings or equivalent perceptual quality compared to the previous generation video coding standard, H.264 / MPEG AVC. While this HEVC standard offers significant coding improvements over its predecessors, it has been demonstrated that superior coding efficiency can be achieved using additional coding tools in HEVC. Based on this, both VCEG and MPEG have begun research into novel coding techniques for future video coding standards. To initiate effective research into advanced technologies that should enable significant improvements in coding efficiency, the Joint Video Exploration Team (JVET) was established by ITU-T VECG and ISO / IEC MPEG in October 2015. The JVET maintains a single reference software, called the Joint Exploration Model (JEM), by integrating several additional coding tools on top of the HEVC Test Model (HM).
[0017] In October 2017, a joint Request for Proposals (CfP) for video compression with capabilities beyond HEVC was issued by ITU-T and ISO / IEC. On April 23, 2018, the 10th JVET meeting received and evaluated the CfP response, which showed an increase in compression efficiency of approximately 40% over HEVC. Based on the results of such evaluation, JVET initiated a new project to develop a new generation video coding standard named Versatile Video Coding (VVC). In the same month, a reference software code base called the VVC Test Model (VTM) was established to demonstrate the VVC standard's reference product.
[0018] Like HEVC, VVC is built on a block-based hybrid video coding framework. Figure 1 shows a general diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows a general encoder 100. The encoder 100 includes a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, and inverse transform 146. 36, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.
[0019] At encoder 100, a video frame is divided into video blocks for processing. For each given video block, a prediction is formed based on either inter-prediction or intra-prediction techniques.
[0020] A prediction residual, which represents the difference between a current video block that is part of video input 110 and a predicted value of the current video block that is part of block predictor 140, is sent from summer 128 to transform 130. The transform coefficients from transform 130 are then sent to quantization 132 for entropy reduction. The quantized coefficients are then provided to entropy coding 138 to generate a compressed video bitstream. As shown in FIG. 1, prediction-related information 142, such as video block partition information, motion vectors (MVs), reference picture indexes, and intra-prediction modes from intra / inter mode decision 116, is also provided through entropy coding 138 and stored in compressed bitstream 144. Compressed bitstream 144 comprises a video bitstream.
[0021] Encoder 100 also requires decoder-related circuitry to reconstruct pixels for prediction. First, a prediction residual is reconstructed by inverse quantization 134 and inverse transform 136. This reconstructed prediction residual is combined with block predictor 140 to generate unfiltered reconstructed pixels for the current video block.
[0022] Spatial prediction (or "intra prediction") predicts the current video block using pixels from samples (called reference samples) of adjacent blocks that have already been coded in the same video frame as the current video block.
[0023] Temporal prediction (also referred to as "inter-prediction") predicts a current video block using pixels reconstructed from an already coded video picture. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically signaled by one or more MVs, which indicate the amount and direction of motion between the current CU and its temporal references. Furthermore, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture storage area the temporal prediction signal comes from.
[0024] Motion estimation 114 takes signals from video input 110 and picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 takes signals from video input 110, picture buffer 120, and motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0025] After spatial prediction and / or temporal prediction are performed, intra / inter mode decision 116 in encoder 100 selects the best prediction mode based, for example, on a rate-distortion optimization technique. Block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using transform 130 and quantization 132. The resulting quantized residual coefficients are inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. The reconstructed CU further undergoes in-loop filtering 122, such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF), before being added to a reference picture storage area in picture buffer 120 for future video block coding. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit 138 to be further compressed and packed to form the bitstream.
[0026] For example, deblocking filters are available in current versions of AVC, HEVC, and VVC. HEVC defines an additional in-loop filter called SAO (Sample Adaptive Offset) to further improve coding efficiency. In the current version of the VVC standard, yet another in-loop filter called ALF (Adaptive Loop Filter) is being actively researched and may be included in the final standard.
[0027] The operation of these in-loop filters is optional. Performing these operations helps improve coding efficiency and visual quality. These in-loop filters may be decided by encoder 100 to be turned off to reduce computational complexity.
[0028] Note that when these filter options are turned on by the encoder 100, intra prediction is typically based on unfiltered reconstructed pixels, while inter prediction is typically based on filtered reconstructed pixels.
[0029] The input video signal is processed block by block (called a coding unit (CU)). In VTM-1.0, a CU can be up to 128x128 pixels. However, unlike HEVC, which only divides blocks based on a quadtree, VVC divides one coding tree unit (CTU) into CUs based on a quadtree / binary tree / ternary tree to fit various local characteristics. In addition, HEVC eliminates the concept of multiple division unit types; that is, in VVC, there is no longer a separation between CUs, prediction units (PUs), and transform units (TUs). Rather, each CU is always used as a basic unit for both prediction and transformation without further division. In the composite tree structure, one CTU is first divided by a quadtree structure. Then, the leaf nodes of each quadtree can be further divided by a binary tree structure and a ternary tree structure.
[0030] As shown in Figures 3A, 3B, 3C, 3D, and 3E (and described below), there are five division types: 4-way division, 2-way horizontal division, 2-way vertical division, 3-way horizontal division, and 3-way vertical division.
[0031] FIG. 3A is a diagram illustrating a quadrant of a block in a composite tree structure according to the present disclosure.
[0032] FIG. 3B is a diagram illustrating a vertical bisection of a block in a composite tree structure according to the present disclosure.
[0033] FIG. 3C is a diagram illustrating a horizontal bisection of a block in a compound tree structure according to the present disclosure.
[0034] FIG. 3D is a diagram illustrating a vertical division of blocks into thirds in a composite tree structure according to the present disclosure.
[0035] FIG. 3E is a diagram illustrating a horizontal division of blocks into thirds in a composite tree structure according to the present disclosure.
[0036] 1, spatial prediction and / or temporal prediction may be performed. Spatial prediction (or "intra prediction") is a prediction method that predicts a block from an already coded neighboring block within the same video picture / slice. Spatial prediction predicts the current video block using pixels from samples of adjacent blocks (called reference samples). Spatial prediction reduces the spatial redundancy inherent in video signals. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") predicts the current video block using pixels reconstructed from previously coded video pictures. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal references. If multiple reference pictures are supported, a reference picture index is additionally sent to identify which reference picture in the reference picture storage area the temporal prediction signal comes from. After spatial prediction and / or temporal prediction, a mode decision block in the encoder selects the best prediction mode based, for example, on a rate-distortion optimization technique. The predicted block is then subtracted from the current video block, and the prediction residual is decorrelated using a transform and quantized. The quantized residual coefficients are inverse quantized and inverse transformed to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. The reconstructed CU is further subjected to in-loop filtering, such as a deblocking filter, sample adaptive offset (SAO), and an adaptive in-loop filter (ALF), before being added to a reference picture store for future video block encoding. To form an output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to an entropy coding unit for further compression and packing to form a bitstream.
[0037] Figure 2 shows an overall block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of a generic decoder 200. The decoder 200 includes a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.
[0038] The decoder 200 is similar to the reconstruction-related parts present in the encoder 100 of FIG. 1. In the decoder 200, an incoming video bitstream 210 is first decoded by entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed by inverse quantization 214 and inverse transform 216 to obtain reconstructed prediction residuals. A block predictor mechanism implemented in an intra / inter mode selector 220 is configured to perform either intra prediction 222 or motion compensation 224 based on the decoded prediction information. A set of unfiltered reconstructed pixels is obtained by summing the reconstructed prediction residual from the inverse transform 216 and the prediction output generated by the block predictor mechanism using an adder 218.
[0039] The reconstructed blocks may further pass through an in-loop filter 228 before being stored in a picture buffer 226, which acts as a reference picture store. The reconstructed video in the picture buffer 226 may be sent to drive a display device as well as used to predict future video blocks. When the in-loop filter 228 is on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 232.
[0040] In Figure 2, the video bitstream is first entropy decoded in the entropy decoding unit. To form a prediction block, the coding mode and prediction information is sent to either a spatial prediction unit (for intra-coding) or a temporal prediction unit (for inter-coding). The residual transform coefficients are sent to an inverse quantization unit and an inverse transform unit to reconstruct a residual block. The predictive block and the residual block are then added together. The reconstructed block may further undergo in-loop filtering before being stored in a reference picture store. The reconstructed video in the reference picture store may then be used to predict future video blocks as well as sent to drive a display device.
[0041] Generally, the basic inter-prediction techniques applied in VVC remain the same as those in HEVC, except that some modules are further extended and / or enhanced. In particular, for all previous video standards, a coding block can only be associated with one MV when the coding block is uni-predicted, and only two MVs when the coding block is bi-predicted. Due to such limitations of traditional block-based motion compensation, small motions still remain in the prediction samples after motion compensation, thus adversely affecting the overall efficiency of motion compensation. To improve both the granularity and accuracy of MVs, two sample-wise improvement methods based on optical flow are currently being researched for the VVC standard: bidirectional optical flow (BDOF) and prediction improvement using optical flow for affine mode (PROF). Below, the main technical aspects of the two inter-coding tools are briefly reviewed.
[0042] Bidirectional Optical Flow
[0043] In VVC, BDOF is applied to refine the predicted samples of bi-predicted coded blocks. Specifically, as shown in Figure 4, BDOF is a motion refinement performed on top of block-based motion compensated prediction for samples when bi-prediction is used. The motion refinement (v x ,vy ) is calculated by minimizing the difference between the predicted samples of L0 and L1 after the BDOF is applied inside one 6×6 window Ω per sub-block. Specifically, (v x , v y ) values are derived as follows.
Number
[0044] Here
Number
Number
[0045] The values of S1, S2, S3, S5, and S6 are calculated as follows.
Number
[0046] Here, the following equation holds.
Number
[0047] Here, I (k) (i, j) is the sample value at the coordinates (i, j) of the predicted signal in the list k (k = 0, 1) generated with intermediate precision (i.e., 16 bits).
[0048]
number
number
[0049] Based on the motion refinement derived in equation (1), the final bi-predicted samples of the CU are calculated by interpolating the L0 / L1 predicted samples along the motion trajectory based on the optical flow model, as instructed by the following equation:
number
[0050] Affine Mode
[0051] In HEVC, only the translational motion model is applied to motion compensated prediction. In the real world, there are various motions, such as zoom in / zoom out, rotation, viewpoint motion, and other irregular motions. In VVC, affine motion compensated prediction is applied by signaling one flag for each inter-coded block, which indicates whether to apply the translational motion model or the affine motion model for inter prediction. In the current VVC design, two affine modes are supported for one affine-coded block, including a 4-parameter affine mode and a 6-parameter affine mode.
[0052] The four-parameter affine model has horizontal translation parameters, vertical translation parameters, zoom parameters, and bidirectional rotation parameters. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. In VVC, to achieve better alignment of the motion vectors and affine parameters, the affine parameters are converted into two motion vectors (also called control point motion vectors (CPMVs)) located at the upper left and upper right corners of the current block. As shown in Figures 5A and 5B, the affine motion field of a block is described by two control point motion vectors (V0, V1).
[0053] Figure 5A shows a four-parameter affine model. Figure 5B shows a four-parameter affine model. 1 illustrates an affine model. Based on the motion of the control points, a motion field (v x ,v y ) is written as follows:
number
[0054] The six-parameter affine mode has a horizontal translation parameter, a vertical translation parameter, a horizontal zoom parameter, a rotation parameter, and a vertical zoom parameter and a rotation parameter. The six-parameter affine motion model is coded using three motion vectors in three CPMVs.
[0055] FIG. 6 is a diagram illustrating a six-parameter affine model. As shown in FIG. 6, three control points of one six-parameter affine block are located at the upper left corner, upper right corner, and lower left corner of the block. The movement of the upper left control point is associated with translational movement, the movement of the upper right control point is associated with rotational and zooming movement in the horizontal direction, and the movement of the lower left control point is associated with rotational and zooming movement in the vertical direction. Compared with the four-parameter affine motion model, the rotational and zooming movement in the horizontal direction of the six-parameter affine model may not be the same as those in the vertical direction. Assuming that the MVs of the upper left corner, upper right corner, and lower left corner of the current block in FIG. 6 are (V0, V1, V2), the MVs of each sub-block (v x ,v y ) is derived using the three MVs at the control points as follows:
number
[0056] PROF for affine modes
[0057] To improve the accuracy of affine motion compensation, PROF is currently being researched in the current VVC, which improves sub-block-based affine motion compensation based on the optical flow model. Specifically, after performing sub-block-based affine motion compensation, the luminance prediction sample of one affine block is corrected by one sample improvement value derived based on the optical flow formula. In detail, the operation of PROF can be summarized as the following four steps:
[0058] Step 1: Subblock-based affine motion compensation is performed to generate the subblock prediction I(i,j) using the subblock MV derived in Equation (6) for the 4-parameter affine model and the subblock MV derived in Equation (7) for the 6-parameter affine model.
[0059] Step 2: The spatial gradient of each prediction sample g x (i,j) and g y (i,j) is the next It is calculated as follows.
number
[0060] To compute the gradient, one additional row / column of prediction samples needs to be generated on each side of one sub-block. To reduce memory bandwidth and complexity, samples on the extended boundary are copied from the nearest integer pixel location in the reference picture, avoiding additional interpolation.
[0061] Step 3: The luminance prediction improvement is calculated as follows:
number
[0062] Step 4: In the current PROF design, after adding the prediction improvement to the original predicted sample, one clipping operation is performed to clip the value of the improved predicted sample to within 15 bits as follows:
number
[0063] FIG. 7 is a diagram showing the PROF process for affine mode.
[0064] Since the parameters of the affine model and the pixel position relative to the sub-block center do not change for each sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks within the same CU. If the horizontal offset from the sample position (i,j) to the center of the sub-block to which the sample belongs is Δx and the vertical offset is Δy, then Δv(i,j) can be derived as follows:
number
[0065] The MV difference Δv(i,j) is derived from the MV of the affine subblocks (6) and (7). Specifically, for a four-parameter affine model,
number
[0066] For the six-parameter affine model, we have:
number
[0067] Affine mode coding efficiency
[0068] Although PROF can improve the coding efficiency of affine mode, its design can be further improved. In particular, considering the fact that both PROF and BDOF are based on the optical flow concept, it is highly desirable to harmonize the design of PROF and BDOF as much as possible so that PROF can make the most of the existing logic circuitry of BDOF to facilitate hardware implementation. Based on such considerations, the following issues regarding the interaction between the current PROF design and BDOF design are identified in this disclosure.
[0069] As explained in the "PROF for Affine Mode" section, in equation (8), the accuracy of the gradient is determined based on the internal bit depth. On the other hand, the MV difference, i.e., Δv x and Δv y is always derived with 1 / 32 pixel precision. Accordingly, based on Equation (9), the precision of the derived PROF improvement depends on the internal bit depth. However, similar to BDOF, PROF is applied to the most significant predicted sample value at an intermediate bit depth (i.e., 16 bits) to maintain higher precision of the PROF derivation. Therefore, the precision of the prediction improvement derived by PROF should match the precision of the intermediate predicted sample, i.e., 16 bits, regardless of the intra-coding bit depth. In other words, the bit depths of the representations of the MV difference and gradient in existing PROF designs are not perfectly matched to derive accurate prediction improvement for the predicted sample precision (i.e., 16 bits). Meanwhile, based on a comparison of Equations (1), (4), and (8), existing PROF and BDOF use different precisions to represent sample gradients and MV differences. As pointed out earlier, such a non-integrated design is undesirable in terms of hardware since it is not possible to reuse existing BDOF logic.
[0070] As discussed in the "PROF for Affine Mode" paragraph, when one of the current affine blocks is bi-predicted, PROF is applied to the prediction samples in lists L0 and L1 separately, and then the improved L0 and L1 prediction signals are averaged to generate the final bi-predictive signal. Rather than deriving a PROF improvement for each prediction direction separately, BDOF derives a prediction improvement once, which is then applied to improve the combined L0 and L1 prediction signal.
[0071] 8 and 9 (described below) compare the current BDOF workflow with PROF for bi-prediction. In the hardware pipeline design of an actual codec, a separate main encoding module / decoding module is usually assigned to each pipeline stage so that more coding blocks can be processed in parallel. However, due to the differences between the BDOF workflow and the PROF workflow, it may be difficult for BDOF and PROF to share the same pipeline design, which is inconvenient for the implementation of an actual codec.
[0072] 8 shows a workflow for BDOF. Workflow 800 includes L0 motion compensation 810, L1 motion compensation 820, and BDOF 830. L0 motion compensation 810 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture that precedes the current picture in the video block. L1 motion compensation 820 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture that follows the current picture in the video block. As described with respect to FIG. 4 above, BDOF 830 takes motion compensation samples from L1 motion compensation 810 and L1 motion compensation 820 and outputs prediction samples.
[0073] FIG. 9 shows the workflow of the existing PROF. Workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0PROF 930, L1PROF 940, and averaging 960. L0 motion compensation 910 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture that precedes the current picture in the video block. L1 motion compensation 920 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is a reference picture that follows the current picture in the video block. L0PROF 930 takes the L0 motion compensation samples from L0 motion compensation 910 and outputs a motion improvement value, as described with reference to FIG. 7 above. L1PROF 940 takes the L1 motion compensation samples from L1 motion compensation 920 and outputs a motion improvement value, as described with reference to FIG. 7 above. Averaging 960 averages the motion improvement value output of L0PROF 930 and the motion improvement value output of L1PROF 940.
[0074] For both BDOF and PROF, the gradient of each sample inside the current coding block needs to be calculated, which requires generating one additional row / column of predicted samples on both sides of the block. To avoid the additional computational complexity of sample interpolation, predicted samples in the extended region around the block are directly copied from reference samples at integer positions (i.e., without interpolation). However, according to existing designs, integer samples at different positions are selected for generating gradient values for BDOF and PROF. Specifically, for BDOF, the integer reference sample to the left of the predicted sample (for horizontal gradient) and the integer reference sample above the predicted sample (for vertical gradient) are used for gradient calculation, while for PROF, the integer reference sample closest to the predicted sample is used for gradient calculation. Similar to the bit depth representation issue, such unified Neither method of gradient calculation is desirable for hardware codec implementations.
[0075] As previously pointed out, the motivation for PROF is to compensate for small MV differences between the MV of each sample and the MV of the sub-block derived at the center of the sub-block to which the sample belongs. According to the current PROF design, PROF is always invoked when a coding block is predicted by the affine mode. However, as indicated in equations (6) and (7), the MVs of the sub-blocks of an affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is relatively small, the MV at each sample position should be stable. In such cases, the benefit of applying PROF is indeed limited, and it may not be worthwhile to implement PROF given the performance / complexity tradeoff.
[0076] Improving Affine Mode Efficiency Using PROF
[0077] This disclosure provides a method for improving and simplifying existing PROF designs to facilitate hardware codec implementation. In particular, special care is taken to harmonize the BDOF design with the PROF design to maximize sharing of existing BDOF logic with PROF. In general, the main aspects of the technology proposed in this disclosure are outlined as follows:
[0078] FIG. 10 illustrates a method for representing the bit depth of a PROF for decoding a video signal according to the present disclosure.
[0079] In step 1010, a first reference picture associated with a video block in the video signal and a first MV from the video block in the current picture to the reference block in the first reference picture are obtained. The first reference picture includes multiple non-overlapping video blocks, and at least one video block is associated with at least one MV. For example, the reference picture may be a video picture adjacent to the picture currently being coded.
[0080] In step 1012, a first predicted sample I(i,j) of the video block generated from a reference block in a first reference picture is obtained. i and j may represent coordinates of one sample with this video block. For example, the predicted sample I(i,j) may be a predicted sample using an MV in the L0 list of the previous reference picture in display order.
[0081] Step 1014 controls the internal bit depth of the internal PROF parameters, which include the horizontal gradient value, vertical gradient value, horizontal motion differential and vertical motion differential derived for the prediction sample I(i,j).
[0082] In step 1016, a prediction improvement value for the first prediction sample I(i,j) is obtained based on the horizontal and vertical gradient values and the horizontal and vertical motion differentials.
[0083] In step 1018, when the video block includes a second MV, obtain a second prediction sample I'(i,j) associated with the second MV and a corresponding prediction improvement value of the second prediction sample I'(i,j).
[0084] In step 1020, a final predicted sample of the video block is calculated based on a combination of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value. Get it.
[0085] First, to improve the coding efficiency of PROF while achieving another integrated design, a method is proposed to unify the bit depth of the representation of sample gradients and MV differences used by BDOF and PROF.
[0086] Second, to facilitate hardware pipeline design, we propose to harmonize the workflow of PROF with that of BDOF for bi-prediction. Specifically, unlike existing PROF, which derives prediction improvements separately for L0 and L1, the proposed method derives prediction improvements once and applies them to the combined L0 and L1 prediction signals.
[0087] Third, two methods are proposed to harmonize the derivation of integer reference samples to calculate the gradient values used by BDOF and PROF.
[0088] Fourth, to reduce the computational complexity, an early termination method is proposed to adaptively inhibit the PROF processing for affine coded blocks when certain conditions are met.
[0089] Improved bit depth representation design for PROF gradients and MV differences
[0090] As analyzed in the section "Improving Affine Mode Efficiency Using PROF," the bit-depth representations of the MV difference and sample gradients in the current PROF are not consistent enough to derive accurate prediction improvements. Furthermore, the bit-depths of the sample gradient and MV difference representations are inconsistent between BDOF and PROF, which is unfavorable for hardware. In this section, a new bit-depth representation method is proposed that improves the BDOF bit-depth representation method by extending it to PROF. Specifically, in the proposed method, the horizontal and vertical gradients at each sample position are calculated as follows:
number
[0091] Additionally, assuming that the horizontal and vertical offsets, expressed with quarter-pixel accuracy, from one sample location to the center of the sub-block to which the sample belongs are Δx and Δy, the MV difference Δv(x,y) of the corresponding PROF at the sample location is derived as follows:
number
[0092] where dMvBits is the bit depth of the gradient values used by the BDOF processing, i.e.
number
[0093] In equations (11) and (12), c, d, e, and f are affine parameters derived based on the affine control points MV. Specifically, for a four-parameter affine model, they are as follows:
number
[0094] For the six-parameter affine model:
number
[0095] Harmonized Workflow of BDOF and PROF for Biprediction
[0096] As previously discussed, when an affine-coded block is bi-predicted, the current PROF is applied in a single manner. More specifically, PROF sample improvements are derived individually and applied to the predicted samples in lists L0 and L1. The improved prediction signals from lists L0 and L1, respectively, are then averaged to generate the final bi-predictive signal for the block. This is in contrast to the BDOF design, in which sample improvements are derived and applied to the bi-predictive signal. Therefore, the difference between the bi-predictive workflows of BDOF and PROF can be disadvantageous for the pipeline design of a practical codec.
[0097] FIG. 11 illustrates a BDOF bit depth representation method for decoding a video signal according to the present disclosure.
[0098] In step 1110, the first reference picture I associated with the video block is (0) and the second reference picture I (1) In display order, the first reference picture I (0) is the picture before the current picture, and the second reference picture I (1) is a picture after the current picture. For example, a reference picture is a picture adjacent to the picture currently being coded. It may be a video picture.
[0099] In step 1112, the first reference picture I (0) The first predicted sample of the video block from the reference block in (0) Get (i,j), where i and j may represent the coordinates of one sample with the current picture.
[0100] In step 1114, the second reference picture I (1) The second predicted sample of the video block from the reference block in (1) Get (i,j).
[0101] In step 1116, the first predicted sample I (0) (i,j) and the second predicted sample I(1) Apply BDOF to the video block based on (i,j).
[0102] In step 1118, a first predicted sample I is calculated based on the padded predicted sample. (0) (i,j) and the second predicted sample I (1) Get the horizontal and vertical gradient values of (i,j).
[0103] At step 1120, a sample motion refinement for the video block is obtained based on the BDOF and horizontal and vertical gradient values applied to the video block.
[0104] At step 1122, bi-predictive samples of the video block are obtained based on the motion refinement.
[0105] According to the present disclosure, one simplification method to facilitate hardware pipeline design is to modify the bi-prediction process of PROF to harmonize the workflow of the two prediction improvement methods. Specifically, instead of applying the improvement to each prediction direction separately, the proposed PROF method derives the prediction improvement once based on the control point MVs of lists L0 and L1, and then applies the derived prediction improvement to the combined L0 and L1 prediction signal to improve quality. Specifically, the proposed method calculates the final bi-prediction sample of one affine-coded block based on the MV difference as derived in Equation (12) as follows:
number
[0106] 12 is a diagram illustrating the PROF process when the proposed bi-predictive PROF method is applied. The PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bi-predictive PROF 1230. The L0 motion compensation 1210 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is a reference picture that is earlier than the current picture in the video block. The L1 motion compensation 1220 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is , are reference pictures after the current picture in the video block. Bi-predictive PROF 1230 takes motion compensation samples from L1 motion compensation 1210 and L1 motion compensation 1220 and outputs bi-predictive samples, as described above.
[0107] Figure 13 illustrates exemplary pipeline stages when both BDOF and the proposed PROF are applied. Figure 13 demonstrates the potential advantages of the proposed method for hardware pipeline design. Pipeline stage 1300 includes analyzing / decoding MV to fetch reference samples 1310, motion compensation 1320, and BDOF / PROF 1330. Pipeline stage 1300 encodes video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block starts with analyzing / decoding MV to fetch reference samples 1310, then moves sequentially to motion compensation 1320, then motion compensation 1320, and BDOF / PROF 1330. This means that BLK0 does not begin processing in pipeline stage 1300 until it moves to motion compensation 1320. This is true for all stages and video blocks in time from T0 to T1, T2, T3, and T4.
[0108] In FIG. 13, the decoding process of one inter block mainly includes the following three steps:
[0109] First, parse / decode the MV of the coded block and take in the reference samples.
[0110] Second, generate L0 and / or L1 prediction signals for the coding block.
[0111] Third, a sample refinement of the generated bi-predictive samples is performed based on BDOF when the coding block is predicted by one non-affine mode, and based on PROF when the coding block is predicted by an affine mode.
[0112] As shown in Figure 13, both BDOF and PROF are directly applied to bi-predictive samples after the proposed harmonization method is applied. If BDOF and PROF are applied to different types of coding blocks (i.e., BDOF is applied to non-affine blocks and PROF is applied to affine blocks), the two coding tools cannot be invoked simultaneously. Therefore, their corresponding decoding processes can be performed by sharing the same pipeline stage. In existing PROF designs, it is difficult to allocate the same pipeline stage to both BDOF and PROF because their bi-predictive workflows are different, whereas the proposed method is more efficient.
[0113] In the above discussion, the proposed method only takes into account the compatibility between the BDOF workflow and the PROF workflow. However, according to the existing design, the basic work units for the two coding tools are also executed with different sizes. Specifically, for BDOF, one coding block is W S ×H S where W is the width of the coding block and W is the number of sub-blocks. S =min(W,16), and H is the height of the coding block. S = min(H,16). BODF operations such as gradient calculation and sample refinement derivation are performed separately for each sub-block. On the other hand, as previously explained, an affine-coded block is divided into 4x4 sub-blocks, and each sub-block is assigned a separate MV derived based on the 4-parameter affine model or the 6-parameter affine model. Since PROF is only applied to affine blocks, the basic operation unit of PROF is a 4x4 sub-block. Similar to the issue of bi-predictive workflow, using a different basic work unit size for PROF than that of BDOF is inconvenient for hardware implementation, making it difficult for BDOF and PROF to share the same pipeline stage of the entire decoding process. In this embodiment, to solve this problem, it is proposed to make the sub-block size of the affine mode the same as that of the BDOF. Specifically, according to the proposed method, when one coding block is coded by the affine mode, W S ×H S where W is the width of the coding block and W is the sub-block size. S =min(W,16), and H is the height of the coding block. S = min(H,16). Each sub-block is assigned one individual MV and is considered as one independent work unit of PROF. It is worth noting that the independent PROF work units ensure that the top-level PROF operation is performed without referring to information from adjacent PROF work units. Specifically, the PROF MV differential at one sample position is calculated as the difference between the MV at the sample position and the MV at the center of the PROF work unit where the sample is located, and the gradient used by the PROF derivation is calculated by padding samples along each PROF work unit. The proposed method claims benefits including three main aspects: 1) simplified pipeline architecture with a unified basic work unit size for both motion compensation and BDOF / PROF improvement, 2) reduced memory bandwidth usage due to the enlarged sub-block size for affine motion compensation, and 3) reduced computational complexity of sample-by-sample fractional sample interpolation.
[0114] In the proposed method, the computational complexity is reduced (i.e., the third aspect), so that the constraints of the existing 6-tap interpolation filter for affine-coded blocks can be overcome. Rather, the default 8-tap interpolation for non-affine-coded blocks is also used for affine-coded blocks. The overall computational complexity in this case can still be advantageous compared to the existing PROF design based on 4x4 sub-blocks using 6-tap interpolation filters.
[0115] Harmonization of gradient derivations for BDOF and PROF
[0116] As previously explained, both BDOF and PROF calculate the gradient of each sample inside the current coding block and access one additional row / column of predicted samples on both sides of the block. To avoid additional interpolation complexity, predicted samples needed in the extension region around the block boundary are directly copied from the integer reference samples. However, as pointed out in the "Problem Statement" paragraph, integer samples at different positions are used to calculate the gradient values of BDOF and PROF.
[0117] To achieve another uniform design, two methods for integrating the gradient derivation method used by BDOF and the gradient derivation method used by PROF are disclosed below. In the first method, it is proposed to make the gradient derivation method of PROF the same as that of BDOF. Specifically, in the first method, the integer positions used to generate prediction samples in the extension region are determined by truncating the fractional sample positions, that is, the integer sample positions to the left of the fractional sample positions (for horizontal gradients) and the integer sample positions above the fractional sample positions (for vertical gradients) are selected.
[0118] In the second method, we propose to make the gradient derivation method for BDOF the same as that for PROF. More specifically, when the second method is applied, the closest integer reference sample to the predicted sample is used for the gradient calculation.
[0119] Figure 14 shows an example of using the BDOF gradient derivation method, where the white circles represent integer-position reference samples 1410, the triangles represent fractional prediction samples 1430 of the current block, and the grey circles represent integer reference samples 1420 used to fill the extension region of the current block.
[0120] Figure 15 shows an example of using the gradient derivation method of PROF, where the white circles represent integer-position reference samples 1510, the triangles represent fractional predicted samples 1530 of the current block, and the grey circles represent integer reference samples 1520 used to fill the extension region of the current block.
[0121] 14 and 15 show the positions of corresponding integer samples used to derive the gradient for BDOF and the gradient for PROF when the first method (FIG. 14) and the second method (FIG. 15) are applied, respectively. In FIGS. 14 and 15, the white circles represent integer-positioned reference samples, the triangles represent fractional predicted samples of the current block, and the gray circles represent integer reference samples used to fill the extension region of the current block for gradient derivation.
[0122] Early termination of PROF based on control point MV difference
[0123] According to the current PROF design, PROF is always invoked when a coding block is predicted by the affine mode. However, as indicated in Equations (6) and (7), the MVs of the sub-blocks of an affine block are derived from the control point MVs. Therefore, when the difference between the control point MVs is relatively small, the MVs at each sample position should be stable. In such a case, the benefit of applying PROF would be limited. Therefore, to further reduce the average calculation complexity of PROF, it is proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between the sample-related MV and the sub-block-related MV within a 4x4 sub-block. Because the MV difference values of the sample PROFs within a 4x4 sub-block are symmetric around the sub-block center, the maximum MV difference of the horizontal PROF and the maximum MV difference of the vertical PROF can be calculated based on Equation (10) as follows:
number
[0124] According to the current disclosure, different metrics may be used to determine whether the MV difference is small enough to skip PROF processing.
[0125] In one example, when the sum of the absolute value of the maximum horizontal MV difference and the absolute value of the maximum vertical MV difference is smaller than a predetermined threshold, that is, when the following equation holds, PROF processing can be skipped based on equation (14).
number
[0126] In another example, |Δv x max | and |Δv y max If the maximum value of | is less than or equal to a threshold, PROF processing may be skipped.
number
[0127] MAX(a,b) is a function that returns the larger of the input values a and b.
[0128] In addition to the above two examples, when other metrics are used to determine whether the MV difference is small enough to skip PROF processing, the spirit of the present disclosure is also applicable in that case.
[0129] In the above method, PROF is skipped based on the magnitude of the MV difference. On the other hand, in addition to the MV difference, PROF sample improvement is also calculated based on local gradient information at each sample position in one motion-compensated block. For predicted blocks with less high-frequency details (e.g., flat areas), the gradient value tends to be small, and the derived sample improvement value tends to be small. Taking this into consideration, according to another aspect of the present disclosure, it is proposed to apply PROF only to predicted samples of blocks that contain sufficiently high-frequency information.
[0130] Various metrics can be used to determine whether a block contains enough high frequency information to merit invoking PROF processing. In one example, the determination is based on the average magnitude (i.e., absolute value) of the gradient of the samples in the predicted block. If the average magnitude is less than a threshold, the predicted block is classified as flat and PROF should not be applied; otherwise, the predicted block is deemed to contain sufficient high frequency detail and PROF can still be applied. In another example, the maximum magnitude of the gradient of the samples in the predicted block can be used. If the maximum magnitude is less than a threshold, PROF for this block is skipped. In yet another example, the difference I between the maximum and minimum sample values of the predicted block is used. max -I min may be used to determine whether PROF should be applied to this block. If such difference value is less than a threshold, PROF for this block is skipped. It should be noted that the spirit of this disclosure is also applicable when some other metric is used to determine whether a given block contains sufficient high frequency information.
[0131] 16 shows a computing environment 1610 coupled with a user interface 1660. The computing environment 1610 may be part of a data processing server. The computing environment 1610 includes a processor 1620, a memory 1640, and an input / output interface 1650.
[0132] The processor 1620 generally controls the overall operation of the computing environment 1610, such as operations related to display, data collection, data communication, and image processing. The processor 1620 may include one or more processors for executing instructions to perform all or some of the steps in the methods described above. Additionally, the processor 1620 may include one or more modules that facilitate interaction between the processor 1620 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0133] The memory 1640 is configured to store various types of data to support the operation of the computing environment 1610. The memory 1640 may include certain software. The memory 1640 may include a memory device 1642. Examples of such data include instructions for any application or method operating on the computing environment 1610, video data sets, image data, etc. The memory 1640 may be implemented using any type of volatile or non-volatile memory device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.
[0134] The input / output interface 1650 provides an interface between the processor 1620 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The input / output interface 1650 may be coupled to an encoder and a decoder.
[0135] In one embodiment, a non-transitory computer-readable storage medium is also provided that includes a plurality of programs executable by the processor 1620 in the computing environment 1610, such as those included in the memory 1640, to perform the methods described above. For example, the non-transitory computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0136] A non-transitory computer-readable storage medium has stored thereon a plurality of programs for execution by a computing device having one or more processors, the plurality of programs, when executed by the one or more processors, causing the computing device to perform the aforementioned method for motion prediction.
[0137] In one embodiment, the computing environment 1610 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0138] The description of the present disclosure has been provided for purposes of illustration and is not intended to be exhaustive or to limit the disclosure. Many modifications, variations and alternative implementations will become apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0139] The above examples have been selected and described to explain the principles of the present disclosure and to enable those skilled in the art to understand various implementations of the present disclosure and to best utilize the underlying principles and various implementations with various modifications suited to the particular applications contemplated. Therefore, it is understood that the scope of the present disclosure is not limited to the specific implementations disclosed, and that modifications and other implementations are intended to be included within the scope of the present disclosure. I want to be understood.
Claims
1. 1. A bit depth representation method for prediction improvement using optical flow (PROF) for decoding a video signal, comprising: obtaining a first reference picture associated with a video block in the video signal and a first motion vector (MV) from a video block in a current picture to a reference block in the first reference picture, the first reference picture including a plurality of non-overlapping video blocks, at least one video block being associated with at least one MV; obtaining a first predicted sample I(i,j) of a video block generated from the reference block in the first reference picture, where i and j represent coordinates of a sample comprising the video block; controlling an internal bit depth of the internal PROF parameters, the internal PROF parameters including a horizontal gradient value, a vertical gradient value, a horizontal motion differential, and a vertical motion differential derived for the prediction sample I(i,j); obtaining a prediction improvement value for the first prediction sample I(i,j) based on horizontal and vertical gradient values and horizontal and vertical motion differentials; when the video block includes a second motion vector, obtaining a second prediction sample I′(i,j) associated with the second motion vector and a corresponding prediction improvement value of the second prediction sample I′(i,j); obtaining a final predicted sample for the video block based on a combination of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value; A method comprising:
2. controlling the internal bit depth of the internal PROF parameters, obtaining a horizontal gradient value of the first predicted sample I(i,j) based on a difference between the first predicted sample I(i+1,j) and the first predicted sample I(i-1,j); obtaining a vertical gradient value of the first predicted sample I(i,j) based on a difference between the first predicted sample I(i,j+1) and the first predicted sample I(i,j-1); right-shifting the horizontal gradient value by a first shift value; right-shifting the vertical gradient value by the first shift value; The method of claim 1 , comprising:
3. The method of claim 2 , wherein the first shift value is equal to the greater of 6 and an encoded bit depth value minus 6.
4. obtaining control points MVs of the video block, the control points MVs including a top left corner block MV, a top right corner block MV, and a bottom left corner block MV of a block including the video block; obtaining affine model parameters derived based on the control points MV; determining horizontal and vertical offsets based on the affine model parameters; a horizontal MV difference Δv for the first predicted sample I(i,j) based on the affine parameters, the horizontal offset, and the vertical offset x obtaining (i, j); a vertical MV differential Δv for a first predicted sample I(i,j) based on the affine parameters, the horizontal offset, and the vertical offset. y Get (i, j) and The horizontal MV difference Δv x right-shifting (i,j) by a second shift value; The vertical MV difference Δv y right-shifting (i, j) by the second shift value; The method of claim 2 further comprising:
5. The method of claim 4 , wherein the second shift value is equal to 13 minus the precision bit depth of the gradient value.
6. The method of claim 5 , wherein the fine bit depth of the gradient values is equal to the greater of 6 and the encoded bit depth minus 6.
7. When the video block includes the second MV, obtaining a final predicted sample for the video block comprises: The horizontal gradient value generated for the first prediction sample I(i,j), the horizontal MV difference Δv x (i, j), the vertical gradient value, and the vertical MV difference Δv y obtaining a first predicted improvement ΔI(i,j) based on (i,j); The horizontal gradient value generated for the second prediction sample I′(i,j), the horizontal motion differential Δv x (i,j), the vertical gradient value, and the vertical motion differential Δv y obtaining a second prediction improvement ΔI′(i,j) based on (i,j); obtaining the prediction improvement by averaging the first prediction improvement ΔI(i,j) and the second prediction improvement ΔI′(i,j); obtaining a bi-predictive sample based on the sum of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value; right-shifting the sum by a third shift value; The method of claim 4, comprising:
8. The method of claim 6 , wherein the third shift value is equal to 15 minus the encoding bit depth.
9. obtaining the horizontal gradient value and the vertical gradient value of the first prediction sample I(i,j), 3. The method of claim 2, further comprising padding additional rows and columns of prediction samples to each of a top boundary, a left boundary, a bottom boundary, and a right boundary of the reference block for the first prediction sample I(i,j).
10. padding the additional rows and columns of the predicted samples copying padded prediction samples of the left and right boundaries from the integer reference samples to the left of the fractional sample positions; copying padded prediction samples of upper and lower boundaries from integer reference samples to upper sides of said fractional sample positions; 10. The method of claim 9, further comprising:
11. padding the additional rows / columns of the predicted samples, copying padded prediction samples for the left and right boundaries from the nearest integer reference sample horizontally to the fractional sample position; copying padded prediction samples for the upper and lower boundaries from the nearest integer reference sample vertically to said fractional sample position; 10. The method of claim 9, further comprising:
12. obtaining the prediction improvement value for the first prediction sample I(i,j), obtaining the largest absolute value of the horizontal direction MV difference values; obtaining the second largest absolute value of the horizontal MV difference values; avoiding calculation of the prediction improvement value when the sum of the largest value and the second largest value is less than a predetermined threshold; The method of claim 1 , comprising:
13. obtaining the prediction improvement value for the first prediction sample I(i,j), obtaining the largest absolute value of the horizontal MV difference values of the video block; obtaining a second largest absolute value of the vertical MV difference values of the video block; avoiding calculation of the prediction improvement value when the larger of the largest value and the second largest value is less than a predetermined threshold; The method of claim 1 , comprising:
14. obtaining the prediction improvement value for the first prediction sample I(i,j), obtaining a largest absolute value of the horizontal gradient values of the video block; obtaining a second largest absolute value of the vertical gradient values for the video block; avoiding said calculation of said prediction improvement value when the larger of said largest value and said second largest value is less than a predetermined threshold; The method of claim 1 , comprising:
15. A bidirectional optical flow (BDOF) bit depth representation method for decoding a video signal, comprising: The first reference picture I associated with the video block (0) and the second reference picture I (1) , in display order, the first reference picture I (0) is the picture before the current picture, and the second reference picture I (1) is after the current picture; The first reference picture I (0) a first predicted sample I of the video block from a reference block in (0) obtaining (i, j), where i and j represent the coordinates of one sample comprising the current picture; The second reference picture I (1) a second predicted sample I of the video block from a reference block in (1) obtaining (i, j); The first predicted sample I (0) (i, j) and the second predicted sample I (1) applying a BDOF to the video block based on (i, j); The first predicted sample I is calculated based on the padded predicted sample. (0) (i, j) and the second predicted sample I (1) obtaining horizontal and vertical gradient values of (i,j); obtaining a sample motion improvement for the video block based on the BDOF and the horizontal and vertical gradient values applied to the video block; obtaining bi-predictive samples for the video block based on the motion refinement; A method comprising:
16. The first predicted sample I (0) (i, j) and the second predicted sample I (1) obtaining the horizontal gradient value and the vertical gradient value of (i,j), The first predicted sample I (0) padding additional rows and columns of predicted samples for each of the upper, left, lower, and right boundaries of (i,j); The second prediction sample I (1) padding additional rows and columns of predicted samples for each of the upper, left, lower, and right boundaries of (i,j); 16. The method of claim 15, comprising:
17. padding the additional rows and columns of the predicted samples copying padded prediction samples of the left and right boundaries from the integer reference samples to the left of the fractional sample positions; copying padded prediction samples of upper and lower boundaries from integer reference samples to upper sides of said fractional sample positions; 17. The method of claim 16, further comprising:
18. padding the additional rows and columns of the predicted samples copying padded prediction samples for the left and right boundaries from the nearest integer reference sample horizontally to the fractional sample position; copying padded prediction samples for the upper and lower boundaries from the nearest integer reference sample vertically to said fractional sample position; 17. The method of claim 16, further comprising:
19. one or more processors; a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors: a process for obtaining a first reference picture associated with a video block in the video signal and a first motion vector (MV) from the video block in a current picture to a reference block in the first reference picture, the first reference picture including a plurality of non-overlapping video blocks, at least one video block being associated with at least one MV; obtaining a first predicted sample I(i,j) of a video block generated from the reference block in the first reference picture, where i and j represent coordinates of a sample comprising the video block; a process for controlling an internal bit depth of the internal PROF parameters, the internal PROF parameters including a horizontal gradient value, a vertical gradient value, a horizontal motion differential, and a vertical motion differential derived for a prediction sample I(i,j); obtaining a prediction improvement value for the first prediction sample I(i,j) based on horizontal and vertical gradient values and horizontal and vertical motion differentials; when the video block includes a second motion vector, obtaining a second prediction sample I′(i,j) associated with the second motion vector and a corresponding prediction improvement value of the second prediction sample I′(i,j); obtaining a final predicted sample for the video block based on a combination of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value; 1. A computing device configured to:
20. the one or more processors configured to control an internal bit depth of the internal PROF parameter; obtaining a horizontal gradient value of the first predicted sample I(i,j) based on a difference between the first predicted sample I(i+1,j) and the first predicted sample I(i-1,j); Between the first predicted sample I(i,j+1) and the first predicted sample I(i,j-1) obtaining a vertical gradient value of the first predicted sample I(i,j) based on the difference; right-shifting the horizontal gradient value by a first shift value; right-shifting the vertical gradient value by the first shift value; 20. The computing device of claim 18, further configured to:
21. 20. The computing device of claim 19, wherein the first shift value is equal to the greater of 6 and an encoded bit depth value minus 6.
22. the one or more processors: a process for obtaining control points MVs of the video block, the control points MVs including a top left corner block MV, a top right corner block MV, and a bottom left corner block MV of a block including the video block; obtaining affine model parameters derived based on the control points MV; determining horizontal and vertical offsets based on the affine model parameters; a horizontal MV difference Δv for the first predicted sample I(i,j) based on the affine parameters, the horizontal offset, and the vertical offset x A process of obtaining (i, j); a vertical MV differential Δv for the first predicted sample I(i,j) based on the affine parameters, the horizontal offset, and the vertical offset y A process of obtaining (i, j); The horizontal MV difference Δv x right-shifting (i, j) by a second shift value; The vertical MV difference Δv y 20. The computing device of claim 19, further configured to: right-shift (i, j) by the second shift value.
23. 22. The computing device of claim 21, wherein the second shift value is equal to 13 minus the precision bit depth of the gradient value.
24. 23. The computing device of claim 22, wherein a precise bit depth of the gradient values is equal to the greater of 6 and the encoded bit depth minus 6.
25. the one or more processors configured to obtain a final predicted sample for the video block when the video block includes the second MV; The horizontal gradient value generated for the first prediction sample (i,j), the horizontal MV difference Δv x (i, j), the vertical gradient value, and the vertical MV difference Δv y obtaining a first predicted improvement ΔI(i,j) based on (i,j); The horizontal gradient value generated for the second prediction sample I′(i,j), the horizontal motion differential Δv x (i,j), the vertical gradient value, and the vertical motion differential Δv y obtaining a second prediction improvement ΔI′(i,j) based on (i,j); obtaining the prediction improvement value by averaging the first prediction improvement value ΔI(i,j) and the second prediction improvement value ΔI′(i,j); obtaining a bi-predictive sample based on the sum of the first predicted sample I(i,j), the second predicted sample I′(i,j), and the prediction improvement value; right-shifting the sum by a third shift value; 22. The computing device of claim 21, further configured to:
26. 24. The computing device of claim 23, wherein the third shift value is equal to 15 minus the encoding bit depth.
27. the horizontal gradient value and the vertical gradient value of the first prediction sample I(i,j) the one or more processors configured to obtain 21. The computing device of claim 20, further configured to pad additional rows and columns of predictive samples to each of a top boundary, a left boundary, a bottom boundary, and a right boundary of the reference block for the first predictive sample I(i,j).
28. the one or more processors configured to pad the additional rows and columns of the prediction samples; copying padded predicted samples of the left and right boundaries from the integer reference samples to the left of the fractional sample positions; copying padded prediction samples of the upper and lower boundaries from the integer reference samples to the upper side of said fractional sample positions; 28. The computing device of claim 27, further configured to:
29. the one or more processors configured to pad the additional rows and columns of the prediction samples; copying padded prediction samples for the left and right boundaries from the nearest integer reference sample horizontally to the fractional sample position; copying padded prediction samples for the upper and lower boundaries from the nearest integer reference sample vertically to said fractional sample position; 28. The method of claim 27, further configured to:
30. the one or more processors configured to obtain the prediction improvement value for the first prediction sample I(i,j), A process of obtaining the largest absolute value of the horizontal direction MV difference values; A process of obtaining the second largest absolute value of the vertical direction MV difference values; avoiding calculation of the prediction improvement value when the sum of the largest value and the second largest value is less than a predetermined threshold; 20. The computing device of claim 19, further configured to:
31. the one or more processors configured to obtain the prediction improvement value for the first prediction sample I(i,j), A process of obtaining the largest absolute value of the horizontal MV difference values of the video block; obtaining the second largest absolute value of the vertical MV difference values of the video block; avoiding the calculation of the prediction improvement value when the larger of the largest value and the second largest value is less than a predetermined threshold; 20. The computing device of claim 19, further configured to:
32. the one or more processors configured to obtain the prediction improvement value for the first prediction sample I(i,j), obtaining the largest absolute value of the horizontal gradient values of the video block; obtaining a second largest absolute value of the vertical gradient values for the video block; avoiding calculation of the prediction improvement value when the larger of the largest value and the second largest value is smaller than a predetermined threshold; 20. The computing device of claim 19, further configured to:
33. one or more processors; a non-transitory computer-readable storage medium storing instructions executable by the one or more processors, wherein the one or more processors: The first reference picture I associated with the video block (0) and the second reference picture I (1) In display order, the first reference picture I (0) is the picture before the current picture, and the second reference picture I (1) is after the current picture; The first reference picture I (0) a first predicted sample I of the video block from a reference block in (0) obtaining (i, j), where i and j represent the coordinates of one sample comprising the current picture; The second reference picture I (1) a second predicted sample I of the video block from a reference block in (1) A process of obtaining (i, j); The first predicted sample I (0) (i, j) and the second predicted sample I (1) applying a BDOF to the video block based on (i,j); The first predicted sample I is calculated based on the padded predicted sample. (0) (i, j) and the second predicted sample I (1) obtaining horizontal and vertical gradient values of (i,j); obtaining a motion improvement of samples for the video block based on the BDOF and the horizontal and vertical gradient values applied to the video block; obtaining bi-predictive samples of the video block based on the motion improvement; 1. A computing device configured to:
34. The first predicted sample I (0) (i, j) and the second predicted sample I (1) the one or more processors configured to obtain the horizontal gradient value and the vertical gradient value of (i, j), The first predicted sample I (0) padding additional rows and columns of predicted samples for each of the upper, left, lower, and right boundaries of (i,j); The second prediction sample I (1) padding additional rows and columns of predicted samples for each of the upper, left, lower, and right boundaries of (i,j); 34. The computing device of claim 33, further configured to:
35. the one or more processors configured to pad the additional rows and columns of the prediction samples; copying padded predicted samples of the left and right boundaries from the integer reference samples to the left of the fractional sample positions; copying padded prediction samples of the upper and lower boundaries from the integer reference samples to the upper side of said fractional sample positions; 35. The computing device of claim 34, further configured to:
36. the one or more processors configured to pad the additional rows and columns of the prediction samples; copying padded prediction samples for the left and right boundaries from the nearest integer reference sample horizontally to the fractional sample position; copying padded prediction samples for the upper and lower boundaries from the nearest integer reference sample in the vertical direction to said fractional sample position; 35. The computing device of claim 34, further configured to: