Method and device for optical flow prediction refinement (PROF)
By introducing optical flow prediction and bidirectional optical flow technology, the problem of insufficient motion compensation accuracy in VVC is solved, the video encoding efficiency and quality are improved, and complex motion modes are adapted to.
Patent Information
- Application Number
- CN202310217548.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-23
- Filing Date
- 2020-09-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-09-17
AI Technical Summary
The existing video encoding and decoding technology has the problem of insufficient motion compensation accuracy in inter-frame prediction, especially in the high-efficiency video encoding and decoding standard VVC. The traditional block-based motion compensation method cannot effectively handle complex motion modes, resulting in limited encoding efficiency.
Optical flow prediction refinement (PROF) and bidirectional optical flow (BDOF) technologies are introduced to improve prediction accuracy by calculating the gradient value and motion refinement of video blocks. PROF refines the predicted samples of video blocks through affine mode, and BDOF improves the accuracy of motion compensation through sample-by-sample motion refinement.
It improves the encoding efficiency of video encoding, enhances the adaptability to complex motion modes, and improves video quality and compression performance.
Smart Images

Figure CN116233466B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application number 202080064927.5 (application date September 17, 2020, invention name “Method and device for optical flow prediction refinement (PROF)”). Technical Field
[0002] The present application relates to video coding and compression. More specifically, the present application relates to methods and apparatus for two inter-frame prediction tools studied in the Versatile Video Coding (VVC) standard, namely, Optical Flow Prediction Refinement (PROF) and Bidirectional Optical Flow (BDOF). Background Art
[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Codec (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Codec (HEVC / H.265), Advanced Video Codec (AVC / H.264), Moving Picture Experts Group (MPEG) codec, etc. Video coding typically uses a prediction method (e.g., inter-frame prediction, intra-frame prediction, etc.) that utilizes redundancy in video images or sequences. An important goal of video coding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. Summary of the Invention
[0004] Embodiments of the present application provide methods and devices for optical flow prediction refinement (PROF) and bidirectional optical flow (BDOF) in video coding and decoding.
[0005] According to a first aspect of the present application, a PROF method is provided. The method may include a decoder obtaining a first reference picture I associated with a video block encoded by an affine mode in a video signal. (0) and the second reference picture I (1) The decoder can also be based on the first reference picture I of the video block (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) to obtain a first horizontal gradient value, a second horizontal gradient value, a first vertical gradient value, and a second vertical gradient value. The decoder can also be based on the first reference picture I of the video block. (0) and the second reference picture I (1)The decoder may further obtain a first prediction refinement ΔI based on the first horizontal gradient value, the second horizontal gradient value, the first vertical gradient value, the second vertical gradient value and the first horizontal motion refinement, the second horizontal motion refinement, the first vertical motion refinement, and the second vertical motion refinement. (0) (i, j) and the second prediction refinement ΔI (1) (i, j). The decoder can also be based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), first prediction refinement ΔI (0) (i, j), second prediction refinement ΔI (1) The final prediction sample of the video block is obtained by adding (i, j) and prediction parameters. These prediction parameters may include weighting parameters and offset parameters for weighted prediction (WP) and weighted bidirectional prediction (BCW) at the coding unit (CU) level.
[0006] According to a second aspect of the present application, a method for PROF is provided, which is implemented by an encoder. The method may include sending two general constraint information (GCI) level control flags by signaling. The two GCI level control flags may include a first GCI level control flag and a second GCI level control flag. The first GCI level control flag indicates whether the BDOF is enabled for the current video sequence. The second GCI level control flag indicates whether the PROF is enabled for the current video sequence. The encoder may also send two sequence parameter set (SPS) level control flags by signaling. The two SPS level control flags indicate whether the BDOF and the PROF are enabled for the current video block in the current video sequence. Among them, the first SPS level control flag indicates that the BDOF is enabled for the current video block, which is based on determining that the BDOF is applied to the first prediction sample I when the video block is not encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives motion refinement of the video block. The second SPS level control flag indicates that PROF is enabled for the current video block, based on determining that PROF is applied to the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives the motion refinement for this video block.
[0007] According to a third aspect of the present application, a computing device is provided. The computing device may include one or more processors and a non-transitory computer-readable storage medium storing instructions executable by the one or more processors. The one or more processors may be configured to obtain a first reference picture I associated with a video block encoded in an affine mode within a video signal. (0) and the second reference picture I (1) The one or more processors may also be configured to generate a first reference image based on the first reference image I (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) obtains a first horizontal gradient value, a second horizontal gradient value, a first vertical gradient value, and a second vertical gradient value. The one or more processors may also be configured to obtain a first horizontal gradient value, a second horizontal gradient value, a first vertical gradient value, and a second vertical gradient value based on the first reference picture I (0) and the second reference picture I (1) The one or more processors may be further configured to obtain a first prediction refinement ΔI based on the first horizontal gradient value, the second horizontal gradient value, the first vertical gradient value, the second vertical gradient value and the first horizontal motion refinement, the second horizontal motion refinement, the first vertical motion refinement, and the second vertical motion refinement. (0) (i, j) and the second prediction refinement ΔI (1) (i, j). The one or more processors may also be configured to calculate the prediction result based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), the first prediction refinement ΔI (0) (i, j), the second prediction refinement ΔI (1) (i, j) and prediction parameters to obtain the final prediction sample of the video block, wherein the prediction parameters include weighting parameters and offset parameters for weighted prediction (WP) and weighted bidirectional prediction (BCW) at the coding unit (CU) level.
[0008] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium having a plurality of instructions stored therein is provided. When executed by one or more processors of a device, the instructions may cause the device to send two general constraint information (GCI) level control flags by a signal. The two GCI level control flags include a first GCI level control flag and a second GCI level control flag. The first GCI level control flag indicates whether the BDOF is enabled for the current video sequence. The second GCI level control flag indicates whether the PROF is enabled for the current video sequence. The instructions may also cause the device to send two SPS level control flags by a signal. The two SPS level control flags indicate whether the BDOF and the PROF are enabled for the current video block. Among them, the first SPS level control flag indicates that BDOF is enabled for the current video block, based on determining that BDOF is applied to the first prediction sample I when the video block is not encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives motion refinement of the video block. Wherein, the second SPS level control flag indicates that PROF is enabled for the current video block, based on determining that PROF is applied to the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives the motion refinement for this video block.
[0009] It is to be understood that both the foregoing general description and the following detailed description are examples only and are not intended to restrict the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples according to the application and, together with the description, serve to explain the principles of the application.
[0011] Figure 1 is a block diagram of an encoder according to an example of the present application.
[0012] Figure 2 is a block diagram of a decoder according to an example of the present application.
[0013] Figure 3A is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0014] Figure 3B is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0015] Figure 3C is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0016] Figure 3D is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0017] Figure 3E is a diagram illustrating block partitioning in a multi-type tree structure according to an example of the present application.
[0018] Figure 4 is an example diagram of a bidirectional optical flow (BDOF) model according to an example of the present application.
[0019] Figure 5A is an example diagram of an affine model according to an example of the present application.
[0020] Figure 5B is an example diagram of an affine model according to an example of the present application.
[0021] Figure 6 is an example diagram of an affine model according to an example of the present application.
[0022] Figure 7 is an example diagram of optical flow prediction refinement (PROF) according to an example of the present application.
[0023] Figure 8 This is the BDOF workflow according to an example of this application.
[0024] Figure 9 This is the workflow of PROF according to the example of this application.
[0025] Figure 10 is a BDOF method according to an example of the present application.
[0026] Figure 11 is the BDOF and PROF method according to the examples of this application.
[0027] Figure 12 is an example diagram of the workflow of PROF for bidirectional prediction according to an example of the present application.
[0028] Figure 13 is an example diagram of pipeline stages of the BDOF and PROF processes according to the present application.
[0029] Figure 14 is an example diagram of the gradient derivation method of BDOF according to the present application.
[0030] Figure 15 This is an example diagram of the gradient derivation method of PROF according to the present application.
[0031] Figure 16A is an example diagram of a derivation template sample for an affine mode according to an example of the present application.
[0032] Figure 16B is an example diagram of a derivation template sample for an affine mode according to an example of the present application.
[0033] Figure 17A is an example diagram of exclusively enabling PROF and LIC for affine mode according to an example of the present application.
[0034] Figure 17B is an example diagram of jointly enabling PROF and LIC for affine mode according to an example of the present application.
[0035] Figure 18A is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0036] Figure 18B is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0037] Figure 18C is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0038] Figure 18D is a diagram illustrating a proposed padding method applied to a 16×16 BDOF CU according to an example of the present application.
[0039] Figure 19 is a diagram illustrating a computing environment coupled with a user interface according to examples of the present application. DETAILED DESCRIPTION
[0040] Reference will now be made in detail to the specific embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numbers in different figures represent the same or similar elements, unless otherwise specified. The embodiments set forth in the following description of exemplary embodiments are not intended to represent all embodiments consistent with the present application. Instead, they are merely examples of apparatus and methods consistent with aspects related to the present application, as described in the appended claims.
[0041] The terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0042] It should be understood that although the terms "first," "second," "third," etc. may be used herein to describe various types of information, such information should not be limited by these terms. These terms are merely used to distinguish one type of information from another. For example, first information may be referred to as second information without departing from the scope of this application. Similarly, second information may be referred to as first information. As used herein, the term "if" may be understood to mean "when," "upon," or "in response to a determination," depending on the context.
[0043] The first version of the HEVC standard was completed in October 2013, offering approximately 50% bitrate savings or equivalent perceptual quality compared to the previous-generation video codec standard, H.264 / MPEG AVC. While the HEVC standard offers significant codec improvements over its predecessor, there is evidence that even higher coding efficiency can be achieved using additional codec tools compared to HEVC. Building on this momentum, both VCEG and MPEG began exploring new codec technologies for future video codec standardization. In October 2015, ITU-T VECG and ISO / IEC MPEG established a Joint Video Exploration Team (JVET) to initiate significant research into advanced technologies that could significantly improve coding efficiency. JVET maintains a reference software called the Joint Exploration Model (JEM) by integrating several additional coding tools on top of the HEVC Test Model (HM).
[0044] In October 2017, ITU-T and ISO / IEC issued a joint call for proposals (CfP) for video compression with capabilities exceeding HEVC. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating compression efficiency improvements of approximately 40% over HEVC. Based on these evaluation results, JVET launched a new project to develop a next-generation video codec standard, called the Versatile Video Codec (VVC). That same month, a reference software codebase, called the VVC Test Model (VTM), was established to demonstrate a reference implementation of the VVC standard.
[0045] Like HEVC, VVC is built on a block-based hybrid video codec framework.
[0046] Figure 1 An overall diagram of a block-based video encoder for VVC is shown. Specifically, Figure 1A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter mode decision 116, a block predictor 140, an adder 128, a transform 130, quantization 132, prediction related information 142, intra prediction 118, a picture buffer 120, inverse quantization 134, an inverse transform 136, an adder 126, a memory 124, an in-loop filter 122, entropy coding 138, and a bitstream 144.
[0047] In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each given video block, a prediction is formed based on either an inter-frame prediction method or an intra-frame prediction method.
[0048] The prediction residual representing the difference between the current video block (part of the video input 110) and its predictor (part of the block predictor 140) is sent from the adder 128 to the transform 130. The transform coefficients are then sent from the transform 130 to the quantization 132 for entropy reduction. The quantized coefficients are then fed to the entropy coding 138 to generate the compressed video bitstream. Figure 1 As shown, prediction related information 142 from the intra / inter mode decision 116, such as video block segmentation information, motion vector (MV), reference picture index, and intra prediction mode, is also fed through entropy coding 138 and saved into a compressed bitstream 144. The compressed bitstream 144 includes a video bitstream.
[0049] In encoder 100, decoder-related circuitry is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed through inverse quantization 134 and inverse transformation 136. This reconstructed prediction residual is combined with a block predictor 140 to generate unfiltered reconstructed pixels for the current video block.
[0050] Spatial prediction (or "intra-prediction") uses pixels from samples of already-encoded neighboring blocks in the same video frame as the current video block (called reference samples) to predict the current video block.
[0051] Temporal prediction (also known as "inter prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in video signals. The temporal prediction signal for a given coding unit (CU) or coding block is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, a reference picture index is additionally sent, which identifies which reference picture in the reference picture library the temporal prediction signal comes from.
[0052] Motion estimation 114 receives the video input 110 and a signal from the picture buffer 120 and outputs a motion estimation signal to motion compensation 112. Motion compensation 112 receives the video input 110, a signal from the picture buffer 120, and a motion estimation signal from motion estimation 114 and outputs a motion compensation signal to intra / inter mode decision 116.
[0053] After performing spatial and / or temporal prediction, the intra / inter mode decision 116 in the encoder 100 selects the optimal prediction mode based on, for example, a rate-distortion optimization method. The block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using a transform 130 and a quantizer 132. The resulting quantized residual coefficients are inversely quantized by an inverse quantizer 134 and inversely transformed by an inverse transform 136 to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal for the CU. Furthermore, in-loop filtering 122 (e.g., a deblocking filter, sample adaptive offset (SAO), and / or an adaptive in-loop filter (ALF)) may be applied to the reconstructed CU before it is placed in the reference picture memory of the picture buffer 120 and used to encode and decode future video blocks. To form the output video bitstream 144, the coding mode (inter or intra), prediction mode information, motion information, and the quantized residual coefficients are all sent to an entropy encoder 138 for further compression and packaging to form the bitstream.
[0054] Figure 1 A block diagram of a general block-based hybrid video coding system is given. The input video signal is processed block by block (called CU). In VTM-1.0, a CU can be up to 128×128 pixels. However, unlike HEVC, which only partitions blocks based on quadtree, in VVC, a coding tree unit (CTU) is split into multiple CUs to adapt to different local characteristics based on quadtree / binary tree / ternary tree. In addition, the concept of multiple partitioning unit types in HEVC is removed, that is, there is no separation of CU, prediction unit (PU) and transform unit (TU) in VVC; instead, each CU is always used as a basic unit for prediction and transformation without further partitioning. In the multi-type tree structure, a CTU is first partitioned by a quadtree structure. Then, the leaf nodes of each quadtree can be further partitioned by binary tree and ternary tree structures. As Figure 3A 、 3B As shown in , 3C, 3D, and 3E, there are five types of segmentation, namely, four-way segmentation, horizontal binary segmentation, vertical binary segmentation, horizontal three-way segmentation, and vertical three-way segmentation.
[0055] Figure 3A is a diagram illustrating block quad partitioning in a multi-type tree structure according to the present application.
[0056] Figure 3B is a diagram illustrating vertical binary partitioning of blocks in a multi-type tree structure according to the present application.
[0057] Figure 3C is a diagram illustrating block-level binary partitioning in a multi-type tree structure according to the present application.
[0058] Figure 3D is a diagram illustrating vertical ternary partitioning of blocks in a multi-type tree structure according to the present application.
[0059] Figure 3E is a diagram illustrating block-level ternary partitioning in a multi-type tree structure according to the present application.
[0060] exist Figure 1 In the video codec, spatial prediction and / or temporal prediction can be performed. Spatial prediction (or "intra-frame prediction") uses pixels from samples of already coded neighboring blocks in the same video picture / slice (called reference samples) to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also known as "inter-frame prediction" or "motion compensated prediction") uses reconstructed pixels from already coded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically represented by one or more motion vectors (MVs), which indicate the amount and direction of motion between the current CU and its temporal reference. In addition, if multiple reference pictures are supported, a reference picture index is additionally sent, which is used to identify which reference picture in the reference picture memory the temporal prediction signal comes from. After spatial and / or temporal prediction, the mode decision block in the encoder selects the best prediction mode based on, for example, rate-distortion optimization methods. The prediction block is then subtracted from the current video block, and the prediction residual is decorrelated and quantized using a transform. The quantized residual coefficients are inversely quantized and inversely transformed to form the reconstructed residual, which is then added back to the prediction block to form the reconstructed signal of the CU. Further, in-loop filtering (such as a deblocking filter, sample adaptive offset (SAO), and / or adaptive in-loop filter (ALF)) can be applied to the reconstructed CU before the reconstructed CU is placed in the reference picture memory and used for encoding and decoding future video blocks. In order to form the output video bitstream, the coding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy coding unit to be further compressed and packaged to form the bitstream.
[0061] Figure 2 The overall block diagram of the video decoder for VVC is shown. Specifically, Figure 2A block diagram of a typical decoder 200 is shown. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction related information 234, and video output 232.
[0062] Decoder 200 is similar to the one residing in Figure 1 The decoder 200 reconstructs the relevant parts of the encoder 100. In the decoder 200, the input video bitstream 210 is first decoded by entropy decoding 212 to derive quantized coefficient levels and prediction-related information. These quantized coefficient levels are then processed by inverse quantization 214 and inverse transform 216 to obtain reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter mode selector 220 is configured to perform intra prediction 222 or motion compensation 224 based on the decoded prediction information. The unfiltered reconstructed pixel set is obtained by adding the reconstructed prediction residual from the inverse transform 216 to the prediction output generated by the block predictor mechanism using an adder 218.
[0063] Before the reconstructed block is stored in a picture buffer 226, which serves as a reference picture memory, the reconstructed block may also be passed through an in-loop filter 228. The reconstructed video in the picture buffer 226 may be sent to drive a display device and may be used to predict future video blocks. With the in-loop filter 228 turned on, filtering operations are performed on the reconstructed pixels to derive the final reconstructed video output 232.
[0064] Figure 2 The overall block diagram of a block-based video decoder is presented. The video bitstream is first entropy decoded in the entropy decoding unit. The codec mode and prediction information are sent to the spatial prediction unit (if intra-frame coding) or the temporal prediction unit (if inter-frame coding) to form the prediction block. The residual transform coefficients are sent to the inverse quantization unit and the inverse transform unit to reconstruct the residual block. The prediction block and the residual block are then added. The reconstructed block may also be in-loop filtered before being stored in the reference picture memory. The reconstructed video in the reference picture memory is then sent to drive the display device and used to predict future video blocks.
[0065] In general, the basic inter prediction techniques applied in VVC remain the same as in HEVC, except that several modules are further extended and / or enhanced. Specifically, for all previous video standards, a coding block can only be associated with a single MV when it is unidirectionally predicted, or with two MVs when it is bidirectionally predicted. Due to this limitation of traditional block-based motion compensation, small motions are still left in the predicted samples after motion compensation, which has a negative impact on the overall efficiency of motion compensation. In order to improve the granularity and accuracy of these MVs, two optical flow-based sample-by-sample refinement methods are currently studied for the VVC standard, namely Bidirectional Optical Flow (BDOF) and Optical Flow Prediction Refinement for Affine Mode (PROF). The following briefly reviews the main technical aspects of these two inter codec tools.
[0066] Bidirectional optical flow
[0067] In VVC, BDOF is applied to refine the prediction samples of bidirectionally predicted coding blocks. Specifically, Figure 4 As shown, when bidirectional prediction is used, BDOF is a sample-by-sample motion refinement performed on the basis of block-based motion compensated prediction.
[0068] Figure 4 An example diagram of a BDOF model according to the present application is shown.
[0069] Motion refinement for each 4×4 sub-block (v x , v y ) is calculated by minimizing the difference between L0 and L1 predicted samples after applying the BDOF within a 6×6 window Ω around the sub-block. Specifically, (v x , v y ) is derived as:
[0070]
[0071]
[0072] in, is a floor function; clip3(min, max, x) is a function that limits the given value x to the range [min, max]; the symbol >> represents a bitwise right shift operation; the symbol << represents a bitwise left shift operation; th BDOF is a motion refinement threshold for preventing propagation errors due to irregular local motion, which is equal to 1<<max(5, bit-depth-7), where bit-depth is the internal bit depth.
[0073] In (1),
[0074] The values of S1, S2, S3, S5, and S6 are calculated as:
[0075]
[0076] in,
[0077]
[0078]
[0079] θ(i, j)=(I (1) (i,j)>>max(4,bit-depth-8))-(I (0) (i, j)>>max(4, bit-depth-8)) (3);
[0080] Among them, I (k) (i, j) is the sample value of the prediction signal at coordinate (i, j) in list k (k=0, 1), which is generated with medium-high precision (i.e., 16 bits); and The horizontal and vertical gradients of a sample are obtained by directly calculating the difference between two adjacent samples of the sample, that is,
[0081]
[0082]
[0083] Based on the motion refinement derived in (1), the final bidirectional prediction samples of the CU are calculated by interpolating the L0 / L1 prediction samples along the motion trajectory based on the optical flow model, as shown in the following formula:
[0084] pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b + o offset )>>shift
[0085]
[0086] Among them, shift and o offset are the right shift and offset values used to merge the L0 and L1 prediction signals for bidirectional prediction, equal to 15-bit-depth and 1<<(14-bit-depth)+2·(1<<13), respectively. Based on the above bit-depth control method, the maximum bit depth of the intermediate parameters in the entire BDOF process is ensured to be no more than 32 bits, and the maximum input of the multiplication is within 15 bits. That is, a 15-bit multiplier is sufficient for BDOF implementation.
[0087] Affine mode
[0088] In HEVC, only the translational motion model is applied to motion-compensated prediction. However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, affine motion-compensated prediction is applied by signaling a flag for each inter-frame coding block to indicate whether the translational motion or affine motion model is applied to inter-frame prediction. In the current VVC design, one affine coding block supports two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode.
[0089] The 4-parameter affine model has the following parameters: two parameters for translation in the horizontal and vertical directions, one parameter for scaling, and one parameter for rotation in both directions. The horizontal scaling parameter is equal to the vertical scaling parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To better adapt the MV and affine parameters, in VVC, these affine parameters are converted into two MVs (also called control point motion vectors (CPMVs)) located at the top left and top right corners of the current block. Figure 5A and Figure 5B As shown, the affine motion field of the block is described by two control points MV(V0, V1).
[0090] Figure 5A A diagram showing a 4-parameter affine model according to the present application is shown.
[0091] Figure 5B A diagram showing a 4-parameter affine model according to the present application is shown.
[0092] Based on the control point motion, an affine encoding block motion field (v x , v y ) is described as:
[0093]
[0094] The 6-parameter affine mode has the following parameters: two parameters for horizontal and vertical translation, one parameter for horizontal scaling, and one parameter for rotation, and one parameter for vertical scaling and one parameter for rotation. The 6-parameter affine motion model is encoded and decoded using three MVs at three CPMVs.
[0095] Figure 6 A diagram showing a 6-parameter affine model according to the present application is shown.
[0096] like Figure 6As shown, the three control points of a 6-parameter affine block are located at the top left, top right, and bottom left corners of the block. The motion at the top left control point is associated with translation, the motion at the top right control point is associated with horizontal rotation and scaling, and the motion at the bottom left control point is associated with vertical rotation and scaling. Compared to the 4-parameter affine motion model, the 6-parameter rotation and scaling motion in the horizontal direction may be different from those in the vertical direction. Assume that (V0, V1, V2) is Figure 6 The MVs of the upper left corner, upper right corner and lower left corner of the current block in the MV, then the three MVs at the control point are used to convert the MV of each sub-block (v x , v y ) is derived as:
[0097]
[0098]
[0099] Refinement of optical flow prediction for affine mode
[0100] To improve the accuracy of affine motion compensation, PROF is currently being studied in VVC, which refines sub-block-based affine motion compensation based on the optical flow model. Specifically, after performing sub-block-based affine motion compensation, the brightness prediction sample of an affine block is modified by a sample refinement value derived based on the optical flow equation. Specifically, the operation of PROF can be summarized in the following four steps:
[0101] Step 1: Perform sub-block based affine motion compensation to generate sub-block predictions I(i, j) using the sub-block MVs derived in (6) for the 4-parameter affine model and in (7) for the 6-parameter affine model.
[0102] Step 2: Spatial gradient g of each prediction sample x (i, j) and g y (i, j) is calculated as follows:
[0103] g x (i,j)=(I(i+1,j)-I(i-1,j))>>(max(2,14-bit-depth)-4)
[0104] g y (i,j)=(I(i,j+1)-I(i,j-1))>>(max(2,14-bit-depth)-4) (8);
[0105] To compute these gradients, an additional row / column of prediction samples is generated on each side of a subblock. To reduce memory bandwidth and complexity, samples on the extended boundary are copied from the nearest integer pixel position in the reference picture to avoid additional interpolation.
[0106] Step 3: The brightness prediction refinement value is calculated by the following formula:
[0107] ΔI(i, j)=g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j) (9);
[0108] Where Δv(i, j) is the difference between the pixel MV calculated for the sample position (i, j) (denoted by v(i, j)) and the sub-block MV of the sub-block where the pixel (i, j) is located. In addition, in the current PROF design, after the prediction refinement is added to the original prediction sample, a clipping operation is performed to clip the value of the refined prediction sample to within 15 bits, that is,
[0109] I r (i, j) = I (i, j) + ΔI (i, j)
[0110] I r (i, j) = clip3(-2 14 , 2 14 -1,I r (i, j))
[0111] Among them, I(i, j) and I r (i, j) are the original prediction sample and the refined prediction sample at position (i, j), respectively.
[0112] Figure 7 A diagram showing the PROF process for affine mode according to the present application is shown.
[0113] Since these affine model parameters and the pixel positions relative to the sub-block center do not change from sub-block to sub-block, Δv(i, j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let Δx and Δy be the horizontal and vertical offsets from the sample position (i, j) to the center of the sub-block to which the sample belongs, then Δv(i, j) can be derived as:
[0114]
[0115] Based on the affine sub-block MV derivation equations (6) and (7), the MV difference Δv(i, j) can be derived. Specifically, for the 4-parameter affine model,
[0116]
[0117] For the 6-parameter affine model,
[0118]
[0119] Among them, (v 0x , v 0y )、(v 1x , v 1y )、(v 2x , v 2y ) are the upper left, upper right and lower left control points MV of the current coding block, w and h are the width and height of the block. In the existing PROF design, the MV difference Δv is always derived with an accuracy of 1 / 32 pixel. x and Δv y .
[0120] Local brightness compensation
[0121] Local illumination compensation (LIC) is a codec tool used to address local illumination variations between temporally adjacent images. A pair of weight parameters and offset parameters are applied to reference samples to obtain a predicted sample for the current block. The overall mathematical model is as follows:
[0122] P[x]=α*P r [x+v]+β (11);
[0123] Among them, P r [x+v] is the reference block indicated by the motion vector v, [α, β] is the corresponding weight parameter and offset parameter pair for the reference block, and P[x] is the final prediction block. This pair of weight parameters and offset parameters is estimated using the minimum linear mean square error (LLMSE) algorithm based on the template of the current block (i.e., the adjacent reconstructed samples) and the reference block of the template (the reference block is derived using the motion vector of the current block). By minimizing the mean square error between these template samples and the reference samples of the template, the mathematical representation of α and β can be derived as follows:
[0124]
[0125]
[0126] Where I represents the number of samples in the template, P c [x i ] is the i-th sample of the template of the current block, P r [x i ] is the reference sample of the i-th template sample based on the motion vector v.
[0127] In addition to being applied to regular inter blocks that contain at most one motion vector per prediction direction (L0 or L1), LIC is also applied to affine mode coded blocks, where a coded block is further split into multiple smaller sub-blocks, each of which can be associated with different motion information. In order to derive the reference samples for LIC for affine mode coded blocks, as described below, Figure 16A and 16B As shown in (12), the reference samples in the top template of an affine coded block are obtained by using the motion vectors of each sub-block in the top sub-block row, while the reference samples in the left template are obtained by using the motion vectors of the sub-blocks in the left sub-block column. The same LLMSE derivation method is then applied to derive these LIC parameters based on the composite template, as shown in (12).
[0128] Figure 16A 16 shows a diagram for deriving template samples for affine mode according to the present application. The diagram includes CurFrame 1620 and CurCU 1622. CurFrame 1620 is the current frame, and CurCU 1622 is the current coding unit.
[0129] Figure 16B A diagram for deriving template samples for affine mode is shown. The diagram contains RefFrame 1640, ColCU 1642, ARef 1643, BRef 1644, CRef 1645, DRef 1646, ERef 1647, FRef 1648, and GRef 1649. RefFrame 1640 is the reference frame, ColCU 1642 is the co-located coding unit, and ARef 1643, BRef 1644, CRef 1645, DRef 1646, ERef 1647, FRef 1648, and GRef 1649 are reference samples.
[0130] Shortcomings of optical flow prediction refinement for affine models
[0131] While PROF can improve the encoding and decoding efficiency of affine modes, its design still needs further improvement. In particular, given that both PROF and BDOF are built on the concept of optical flow, it is highly desirable to coordinate the designs of PROF and BDOF as much as possible so that PROF can maximize the use of BDOF's existing logic to facilitate hardware implementation. Based on this consideration, this application identifies the following deficiencies in the interaction between the current PROF and BDOF designs.
[0132] First, as described in the “Optical flow prediction refinement for affine mode” section, in Equation (8), the accuracy of the gradient is determined based on the internal bit depth. On the other hand, the MV difference is always derived with an accuracy of 1 / 32 pixel, i.e., Δv x and Δv y . Accordingly, based on equation (9), the accuracy of the derived PROF refinement depends on the internal bit depth. However, similar to BDOF, PROF is applied on the prediction sample values of medium and high bit depths (i.e., 16 bits) to maintain higher PROF derivation accuracy. Therefore, regardless of the internal coding bit depth, the prediction refinement accuracy derived by PROF should match the accuracy of the intermediate prediction samples, i.e., 16 bits. In other words, the representation bit depth of MV differences and gradients in the existing PROF design is not fully matched, and accurate prediction refinement relative to the prediction sample accuracy (i.e., 16 bits) cannot be derived. At the same time, based on the comparison of equations (1), (4) and (8), the existing PROF and BDOF use different precisions to represent sample gradients and MV differences. As pointed out earlier, this non-uniform design is not desirable for hardware because the existing BDOF logic cannot be reused.
[0133] Second, as discussed in the "Optical Flow Prediction Refinement for Affine Mode" section, when bidirectionally predicting a current affine block, PROF is applied to the prediction samples in lists L0 and L1 separately; then, the enhanced L0 and L1 prediction signals are averaged to generate the final bidirectional prediction signal. In contrast, BDOF does not derive PROF refinements separately for each prediction direction, but derives prediction refinements once and then applies them to enhance the merged L0 and L1 prediction signals. (As described below) Figure 8 and Figure 9 The current BDOF and PROF workflows for bidirectional prediction are compared. In actual codec hardware pipeline designs, different main encoding / decoding modules are usually assigned to each pipeline stage to enable parallel processing of more coding blocks. However, due to the differences between the BDOF and PROF workflows, this can make it difficult to have a single pipeline design that can be shared by both BDOF and PROF, which is not friendly to actual codec implementations.
[0134] Figure 8The workflow of BDOF according to the present application is shown. The workflow 800 includes L0 motion compensation 810, L1 motion compensation 820 and BDOF 830. L0 motion compensation 810 can be, for example, a list of motion compensated samples from a previous reference picture. The previous reference picture is a reference picture that precedes the current picture in the video block. L1 motion compensation 820 can be, for example, a list of motion compensated samples from a next reference picture. The next reference picture is a reference picture that follows the current picture in the video block. BDOF 830 obtains motion compensated samples from L0 motion compensation 810 and L1 motion compensation 820 and outputs predicted samples, as previously described. Figure 4 As described in .
[0135] Figure 9 The workflow of the existing PROF according to the present application is shown. The workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0 PROF 930, L1 PROF 940 and averaging 960. L0 motion compensation 910 can be, for example, a list of motion compensated samples from a previous reference picture. The previous reference picture is a reference picture that precedes the current picture in the video block. L1 motion compensation 920 can be, for example, a list of motion compensated samples from a next reference picture. The next reference picture is a reference picture that follows the current picture in the video block. L0 PROF 930 obtains the L0 motion compensated samples from L0 motion compensation 910 and outputs a motion refinement value, as previously described. Figure 7 L1 PROF 940 takes L1 motion compensation samples from L1 motion compensation 920 and outputs motion refinement values, as previously described in Figure 7 Average 960 averages the motion refinement values output by L0PROF 930 and L1 PROF 940.
[0136] Third, for BDOF and PROF, gradients need to be calculated for each sample within the current coding block, which requires generating an additional row / column of prediction samples on each side of the block. To avoid the additional computational complexity of sample interpolation, the prediction samples in the extended area around the block are copied directly from the reference samples at integer positions (i.e., no interpolation). However, according to the existing design, integer samples at different positions are selected to generate the gradient values for BDOF and PROF. Specifically, for BDOF, integer reference samples located to the left of the prediction sample (horizontal gradient) and above the prediction sample (vertical gradient) are used; for PROF, the integer reference sample closest to the prediction sample is used for gradient calculation. Similar to the bit depth representation problem, this non-uniform gradient calculation method is also undesirable for hardware codec implementation.
[0137] Fourth, as pointed out earlier, the motivation of PROF is to compensate for small MV differences between the MV of each sample and the sub-block MV derived at the center of the sub-block to which the sample belongs. According to the current PROF design, PROF is always called when an affine mode is used to predict a coding block. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from the control point MVs. Therefore, when the differences between the control point MVs are small, the MVs of each sample position should be consistent. In this case, since the benefits of applying PROF may be very limited, it may not be worthwhile to perform PROF when considering the performance / complexity trade-off.
[0138] Improved refinement of optical flow prediction for affine models
[0139] This application presents a method for improving and simplifying existing PROF designs to facilitate hardware codec implementations. Specifically, special attention is paid to coordinating the design of BDOF and PROF to maximize sharing of existing BDOF logic with PROF. In general, the main aspects of the techniques presented in this application are summarized below.
[0140] First, to improve the coding efficiency of PROF while achieving a more unified design, a method is proposed to unify the representation bit depth of sample gradients and MV differences used by BDOF and PROF.
[0141] Second, to facilitate hardware pipeline design, we propose a method that coordinates the PROF workflow with the BDOF workflow to enable bidirectional prediction. Specifically, unlike the existing PROF method, which derives prediction refinements for L0 and L1 separately, the proposed method derives a single prediction refinement, which is applied to the combined L0 and L1 prediction signals.
[0142] Third, two methods are proposed to coordinate the derivation of integer reference samples to compute the gradient values used by BDOF and PROF.
[0143] Fourth, to reduce the computational complexity, an early termination method is proposed to adaptively disable the PROF process for affine coded blocks when certain conditions are met.
[0144] Improved bit depth representation design for PROF gradient and MV difference
[0145] As analyzed in the "Problem Statement" section, the representation bit depths of MV differences and sample gradients in the current PROF are not aligned to achieve accurate prediction refinement. In addition, the representation bit depths of sample gradients and MV differences between BDOF and PROF are inconsistent, which is not hardware-friendly. In this section, an improved bit depth representation method is proposed by extending the bit depth representation method of BDOF to PROF. Specifically, in the proposed method, the horizontal and vertical gradients at each sample position are calculated as:
[0146] g x (i,j)=(I(i+1,j)-I(i-1,j))>>max(6, bit-depth-6)
[0147] g y (i, j)=(I(i, j+1)-I(i, j-1))>>max(6, bit-depth-6) (13);
[0148] Furthermore, assuming that Δx and Δy are the horizontal and vertical offsets from a sample position to the center of the sub-block to which the sample belongs, expressed in 1 / 4 pixel precision, the corresponding PROF MV difference Δv(x, y) at the sample position is derived as:
[0149] Δv x (i,j)=(c*Δx+d*Δy)>>(13-dMvBits)
[0150] Δv y (i,j)=(e*Δx+f*Δy)>>(13-dMvBits) (14);
[0152] Where dMvBits is the bit depth of the gradient value used by the BDOF process, that is, dMvBits = max(5, (bit-depth-7)) + 1. In equations (13) and (14), c, d, e, and f are affine parameters derived based on the affine control point MV. Specifically, for a 4-parameter affine model,
[0153]
[0154] For the 6-parameter affine model,
[0155]
[0156] Among them, ((v 0x , v 0y )、(v 1x , v 1y ) and (v 2x , v 2y) are the upper left, upper right, and lower left control point MVs of the current coding block, expressed with 1 / 16 pixel precision, and w and h are the width and height of the block.
[0157] In the previous discussion, as shown in equations (13) and (14), a pair of fixed right shifts are applied to calculate these gradient values and MV differences. In practice, for different trade-offs between the intermediate calculation accuracy and the bit width of the internal PROF derivation process, different bit-by-bit right shifts can be applied to (13) and (14) to achieve different representation accuracy of these gradients and MV differences. For example, when the input video contains a lot of noise, the derived gradients may not reliably represent the true local horizontal / vertical gradient values at each sample. In this case, it makes more sense to use more bits to represent the MV differences than the gradients. On the other hand, when the input video shows stable motion, the MV differences derived from the affine model should be very small. If so, using high-precision MV differences does not provide additional benefits to improve the accuracy of the derived PROF refinement. In other words, in this case, it is more advantageous to use more bits to represent the gradient values. Based on the above considerations, in one or more embodiments of the present application, a general method for calculating gradients and MV differences for PROF is proposed below. Specifically, it is assumed that the horizontal and vertical gradients at each sample position are obtained by performing n-scaling on the differences of adjacent prediction samples. a It is calculated by right shift, that is,
[0158] g x (i,j)=(I(i+1,j)-I(i-1,j))>>n a
[0159] g y (i,j)=(I(i,j+1)-I(i,j-1))>>n a (15);
[0160] The corresponding PROF MV difference Δv(x, y) at this sample location should be calculated as:
[0161] Δv x (i,j)=(c*Δx+d*Δy)>>(13-n a )
[0162] Δv y (i,j)=(e*Δx+f*Δy)>>(13-n a ) (16);
[0163] Where Δx and Δy are the horizontal and vertical offsets from a sample position to the center of the sub-block to which the sample belongs, expressed in 1 / 4 pixel precision, and c, d, e, and f are parameters derived based on the 1 / 16 pixel affine control point MV. Finally, the final PROF refinement of the sample is calculated as:
[0164] ΔI(i, j)=(g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j)+1)>>1 (17);
[0165] In some embodiments of the present application, another PROF bit depth control method is proposed as follows. In this method, the difference between adjacent prediction samples is still n-th a The horizontal and vertical gradients at each sample position are calculated by right shifting as in (15). The corresponding PROF MV difference Δv(x, y) at this sample position should be calculated as:
[0166] Δv x (i,j)=(c*Δx+d*Δy)>>(14-n a ),
[0167] Δv y (i,j)=(e*Δx+f*Δy)>>(14-n a ).
[0168] Furthermore, to keep the entire PROF derivation within the appropriate internal bit depth, the derived MV differences are cropped as follows:
[0169] Δv x (i, j) = Clip3 (-limit, limit, Δv x (i, j)),
[0170] Δv y (i, j) = Clip3 (-limit, limit, Δv y (i, j));
[0171] Among them, limit is equal to clip3(min, max, x) is a function that clips a given value x to the range [min, max]. In an example, n b The value is set to 2 max(5,bit-depth-7) Finally, the final PROF refinement for this sample is calculated as:
[0172] ΔI(i, j)=g x (i, j)*Δvx (i, j)+g y (i, j)*Δv y (i, j).
[0173] In addition, in one or more embodiments of the present application, a PROF bit depth control scheme is proposed. In this method, the horizontal and vertical PROF motion refinements at each sample position (i, i) are derived as follows:
[0174] Δv x (i, j)=(c*Δx+d*Δy)>>(13-max(5, bit-depth-7)),
[0175] Δv y (i,j)=(e*Δx+f*Δy)>>(13-max(5, bit-depth-7)).
[0176] Additionally, the exported horizontal and vertical motions are refined and cropped to:
[0177] Δv x (i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv x (i, j))
[0178] Δv y (i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv y (i, j)).
[0179] Here, given the motion refinement derived above, the final PROF sample refinement at position (i, i) is calculated as:
[0180] ΔI(i, j)=g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j).
[0181] In another embodiment, another PROF bit depth control scheme is proposed. In the second method, the horizontal and vertical PROF motion at the sample position (i, j) are refined and derived as:
[0182] Δv x (i, j)=(c*Δx+d*Δy)>>(13-max(6, bit-depth-6))
[0183] Δv y(i,j)=(e*Δx+f*Δy)>>(13-max(6, bit-depth-6)).
[0184] These exported motions are then refined and cropped to:
[0185] Δv x (i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv x (i, j))
[0186] Δv y (i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv y (i, j)).
[0187] Therefore, given the motion refinement derived above, the final PROF sample refinement at position (i, i) is calculated as:
[0188] ΔI(i, j)=(g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j)+1)>>1
[0189] In one or more embodiments of the present application, it is proposed to combine the motion refinement accuracy control method in the solution with the PROF sample refinement derivation method in the second solution. Specifically, through this method, the horizontal and vertical PROF motion refinements at the sample position (i, j) are derived as follows:
[0190] Δv x (i, j)=(c*Δx+d*Δy)>>(13-max(5, bit-depth-7)),
[0191] Δv y (i,j)=(e*Δx+f*Δy)>>(13-max(5, bit-depth-7)).
[0192] Furthermore, these derived horizontal and vertical motions are refined and cropped to:
[0193] Δv x (i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv x (i, j))
[0194] Δv y(i, j)=Clip3(-max(5, bit-depth-7), max(5, bit-depth-7)-1, Δv y (i, j)).
[0195] Finally, given the motion refinement derived above, the final PROF sample refinement at position (i, i) is calculated as:
[0196] ΔI(i, j)=(g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j)+1)>>1.
[0197] BDOF and PROF coordinated workflow for bidirectional prediction
[0198] As discussed previously, when an affine-coded block is bidirectionally predicted, the current PROF is applied unilaterally. More specifically, these PROF sample refinements are derived and applied to the prediction samples in lists L0 and L1, respectively. The refined prediction signals from lists L0 and L1 are then averaged to generate the final bidirectional prediction signal for the block. This contrasts with the BDOF design, where these sample refinements are derived and applied to the bidirectional prediction signal. This difference between the bidirectional prediction workflows of BDOF and PROF may not be friendly to practical codec pipeline designs.
[0199] To facilitate hardware pipeline design, according to the present application, a simplified approach is to modify the bidirectional prediction process of PROF so that the workflows of the two prediction refinement methods are consistent. Specifically, instead of applying refinement to each prediction direction separately, the proposed PROF method derives a prediction refinement based on the control point MVs of lists L0 and L1; these derived prediction refinements are then applied to the merged L0 and L1 prediction signals to improve quality. Specifically, based on the MV differences derived in equation (14), the final bidirectional prediction samples of an affine coded block are calculated by the proposed method as:
[0200] pred PROF (i, j) = (I (0) (i, j)+I (1) (i, j) + ΔI(i, j) + o offset )>>shift,
[0201] ΔI(i, j)=(g x (i, j)*Δv x (i, j)+g y (i, j)*Δvy (i, j)+1)>>1
[0202] I r (i,j)=I(i,j)+ΔI(i,j) (18);
[0203] Among them, shift and o offset are the right shift value and offset value used to merge the L0 and L1 prediction signals for bidirectional prediction, which are equal to (15-bit-depth) and 1<<(14-bit-depth)+(2<<13), respectively. In addition, as shown in (18), the cropping operation in the existing PROF design (as shown in (9)) is removed in the proposed method.
[0204] Figure 12 The corresponding PROF process when applying the proposed bidirectional prediction PROF method is shown. PROF process 1200 includes L0 motion compensation 1210, L1 motion compensation 1220, and bidirectional prediction PROF 1230. L0 motion compensation 1210 can be, for example, a list of motion compensated samples from a previous reference picture. This previous reference picture is a reference picture that precedes the current picture in the video block. L1 motion compensation 1220 can be, for example, a list of motion compensated samples from a next reference picture. This next reference picture is a reference picture that follows the current picture in the video block. As described above, bidirectional prediction PROF 1230 receives motion compensated samples from L0 motion compensation 1210 and L1 motion compensation 1220 and outputs bidirectional prediction samples.
[0205] To demonstrate the potential benefits of the proposed approach for hardware pipeline design, Figure 13 An example is shown to illustrate the pipeline stages when BDOF and the proposed PROF are applied simultaneously. Figure 13 In , the decoding process of an inter-frame block mainly includes three steps:
[0206] First, parse / decode the MV of the coded block and obtain the reference sample.
[0207] Second, generate L0 and / or L1 prediction signals for the coding block.
[0208] Third, based on the BDOF when the coding block is predicted by a non-affine mode or the PROF when the coding block is predicted by an affine mode, sample-by-sample refinement is performed on the generated bidirectional prediction samples.
[0209] Figure 13 A diagram showing exemplary pipeline stages when applying BDOF and the proposed PROF according to the present application is shown. Figure 13This paper demonstrates the potential benefits of the proposed approach for hardware pipeline design. Pipeline stage 1300 includes parsing / decoding MV and acquiring reference samples 1310, motion compensation 1320, and BDOF / PROF 1330. Pipeline stage 1300 will encode video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block will begin by parsing / decoding MV and acquiring reference samples 1310 and move sequentially to motion compensation 1320, motion compensation 1320, and BDOF / PROF 1330. This means that BLK0 does not begin processing in pipeline stage 1300 until it moves to motion compensation 1320. This is true for all stages and video blocks as time passes from T0 to T1, T2, T3, and T4.
[0210] Figure 13 In , the decoding process of an inter-frame block mainly includes three steps:
[0211] First, parse / decode the MV of the coded block and obtain the reference sample.
[0212] Second, generate L0 and / or L1 prediction signals for the coding block.
[0213] Third, based on the BDOF when the coding block is predicted by a non-affine mode or the PROF when the coding block is predicted by an affine mode, sample-by-sample refinement is performed on the generated bidirectional prediction samples.
[0214] like Figure 13 As shown in Figure 3, after applying the proposed coordination method, both BDOF and PROF are directly applied to bidirectionally predicted samples. Given that BDOF and PROF are applied to different types of coding blocks (i.e., BDOF is applied to non-affine blocks and PROF is applied to affine blocks), the two encoding tools cannot be invoked simultaneously. Therefore, their corresponding decoding processes can be performed by sharing the same pipeline stage. This is more efficient than existing PROF designs, in which it is difficult to assign BDOF and PROF to the same pipeline stage due to their different bidirectional prediction workflows.
[0215] In the above discussion, the proposed method only considers the coordination of BDOF and PROF workflows. However, according to the existing design, the basic operation units for these two coding tools are also performed at different sizes. For example, for BDOF, a coding block is split into multiple blocks of size W. s ×H s sub-blocks, where W s =min(W, 16), H s=min(H, 16), where W and H are the width and height of the coding block, respectively. BODF operations such as gradient calculation and sample refinement derivation are performed independently for each sub-block. On the other hand, as mentioned above, the affine coding block is divided into 4×4 sub-blocks, and each sub-block is assigned a separate MV derived based on a 4-parameter or 6-parameter affine model. Since PROF is only applied to the affine block, its basic operation unit is a 4×4 sub-block. Similar to the bidirectional prediction workflow problem, using different basic operation unit sizes from BDOF to PROF is not friendly to hardware implementation, and makes it difficult for BDOF and PROF to share the same pipeline stage of the entire decoding process. In order to solve such problems, in one or more embodiments, it is proposed to align the sub-block size of the affine mode to be the same as the sub-block size of BDOF.
[0216] Here, according to the proposed method, if a coding block is affine-coded, it will be split into blocks of size W s ×H s sub-blocks, where W s =min(W, 16), H s =min(H, 16), where W and H are the width and height of the coding block. Each sub-block is assigned a separate MV and is regarded as an independent PROF operation unit. It is worth mentioning that the independent PROF operation unit ensures that the PROF operation performed on it does not need to refer to information from adjacent PROF operation units. Specifically, the PROF MV difference at a sample position is calculated as the difference between the MV at the sample position and the MV at the center of the PROF operation unit where the sample is located; the gradient used for PROF derivation is calculated by filling samples along each PROF operation unit. The benefits of the proposed method mainly include the following aspects: 1) a simplified pipeline architecture with a unified basic operation unit size for motion compensation and BDOF / PROF refinement; 2) reduced memory bandwidth usage due to the increase in sub-block size for affine motion compensation; 3) reduced per-sample calculation complexity of fractional sample interpolation.
[0217] It should also be mentioned that due to the reduced computational complexity using the proposed method (i.e., item 3), the existing 6-tap interpolation filter constraint for affine coded blocks can be removed. Instead, the default 8-tap interpolation used for non-affine coded blocks is also used for affine coded blocks. In this case, the overall computational complexity is still advantageous over the existing PROF design (i.e., based on 4×4 sub-blocks with 6-tap interpolation filters).
[0218] Harmonization of gradient derivation for BDOF and PROF
[0219] As mentioned earlier, both BDOF and PROF compute the gradient for each sample within the current coding block, accessing one additional row / column of prediction samples on each side of the block. To avoid additional interpolation complexity, the required prediction samples in the extended area around the block boundary are copied directly from the integer reference samples. However, as pointed out in the "Problem Statement" section, integer samples at different locations are used to calculate the gradient values for BDOF and PROF.
[0220] In order to achieve a more unified design, two methods are proposed below to unify the gradient derivation methods used by BDOF and PROF. In the first method, it is proposed to align the gradient derivation method of PROF with the gradient derivation method of BDOF. Specifically, with the first method, the integer positions used to generate these prediction samples in the extended area are determined by rounding down the fractional samples, that is, the selected integer sample positions are located to the left of the fractional sample positions (for horizontal gradients) and above the fractional sample positions (for vertical gradients). In the second method, it is proposed to align the gradient derivation method of BDOF with the gradient derivation method of PROF. In more detail, when the second method is applied, the integer reference sample closest to the prediction sample is used for gradient calculation.
[0221] Figure 14 An example of a gradient derivation method using BDOF according to the present application is shown. Figure 14 In , the blank circles represent reference samples at integer positions, the triangles represent fractional prediction samples of the current block, and the black circles represent integer reference samples used to fill the extended area of the current block.
[0222] Figure 15 An example of a gradient derivation method using PROF according to the present application is shown. Figure 15 In , the blank circles represent reference samples at integer positions, the triangles represent fractional prediction samples of the current block, and the black circles represent integer reference samples used to fill the extended area of the current block.
[0223] Figure 14 and Figure 15 The results show that when the first method ( Figure 14 ) and the second method ( Figure 15 ) is used for the derivation of the gradients of BDOF and PROF. Figure 14 and Figure 15 In , the open circles represent reference samples at integer positions, the triangles represent the fractional prediction samples of the current block, and the patterned circles represent integer reference samples used to fill the extended area of the current block for gradient derivation.
[0224] In addition, according to the existing BDOF and PROF designs, prediction sample padding is performed at different codec levels. Specifically, for BDOF, padding is applied along the boundary of each sbWidth×sbHeight sub-block, where sbWidth=min(CUWidth, 1 6) and sbHeight=min(CUHeight, 1 6). CUWidth and CUHeight are the width and height of a CU. On the other hand, PROF padding is always applied at the 4×4 sub-block level. In the above discussion, only the padding method is unified between BDOF and PROF, while the padding sub-block size remains different. This is also not friendly to actual hardware implementation because different modules need to be implemented for the padding process of BDOF and PROF. In order to achieve a more unified design, it is proposed to unify the sub-block padding size of BDOF and PROF. In one or more embodiments of the present application, it is proposed to apply BDOF prediction sample padding at the 4×4 level. Specifically, with this method, the CU is first divided into multiple 4×4 sub-blocks; after motion compensation is performed on each 4×4 sub-block, the extended samples along the upper / lower and left / right boundaries are filled by copying the corresponding integer sample positions. Figure 18A 、 18B , 18C and 18D show an example of applying the proposed padding method to a 16×16 BDOF CU, where the dotted lines represent 4×4 sub-block boundaries and the black bands represent the padded samples of each 4×4 sub-block.
[0225] Figure 18A 18 shows the proposed padding method according to the present application applied to a 16×16 BDOF CU, where the dotted line represents the upper left 4×4 sub-block boundary 1820 .
[0226] Figure 18B The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the top right 4×4 sub-block boundary 1840 .
[0227] Figure 18C The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the lower left 4×4 sub-block boundary 1860 .
[0228] Figure 18D The proposed padding method according to the present application applied to a 16×16 BDOF CU is shown, where the dashed line represents the bottom-right 4×4 sub-block boundary 1880 .
[0229] Enable / disable advanced signaling semantics for BDOF, PROF, and DMVR
[0230] In the existing BDOF and PROF design, there are two different flags in the sequence parameter set (SPS) to control the enable / disable of the two coding tools respectively. However, due to the similarity between BDOF and PROF, it is more desirable to enable and / or disable BDOF and PROF from a high level through a same control flag. Based on this consideration, a new flag called sps_bdof_prof_enabled_flag is introduced in the SPS, as shown in Table 1. As shown in Table 1, the enabling and disabling of BDOF depends only on sps_bdof_prof_enabled_flag. When this flag is equal to 1, BDOF is enabled to encode and decode the video content in the sequence. Otherwise, when sps_bdof_prof_enabled_flag is equal to 0, BDOF will not be applied. On the other hand, in addition to sps_bdof_prof_enabled_flag, the SPS level affine control flag, sps_affine_enabled_flag, is also used to conditionally enable and disable PROF. PROF is enabled for all coded blocks coded in affine mode when both the flags sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are equal to 1. PROF is disabled when the flags sps_bdof_prof_enabled_flag are equal to 1 and sps_affine_enabled_flag are equal to 0.
[0231] Table 1 Modified SPS semantic table with proposed BDOF / PROF enable / disable flags
[0232]
[0233] sps_bdof_prof_enabled_flag indicates whether bidirectional optical flow and optical flow prediction refinement are enabled. When sps_bdof_prof_enabled_flag is equal to 0, both bidirectional optical flow and optical flow prediction refinement are disabled. When sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 1, both bidirectional optical flow and optical flow prediction refinement are enabled. Otherwise (sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0), bidirectional optical flow is enabled and optical flow prediction refinement is disabled.
[0234] sps_bdof_prof_dmvr_slice_preset_flag indicates when the flag slice_disable_bdof_prof_dmvr_flag is signaled at the slice level. When this flag is equal to 1, the syntax slice_disable_bdof_prof_dmvr_flag is signaled for each slice that references the current sequence parameter set. Otherwise (when sps_bdof_prof_dmvr_slice_present_flag is equal to 0), the syntax slice_disabled_bdof_prof_dmvr_flag is not signaled at the slice level. When this flag is not signaled, it is inferred to be 0.
[0235] In addition, when the recommended SPS level BDOF and PROF control flags are used, the corresponding control flag no_bdof_constraint_flag in the general constraint information semantics should also be modified by the following table:
[0236]
[0237] no_bdof_prof_constraint_flag equal to 1 indicates that sps_bdof_prof_enabled_flag shall be equal to 0. no_bdof_constraint_flag equal to 0 imposes no constraint.
[0238] In addition to the above-mentioned SPSBDOF / PROF semantics, it is proposed to introduce another control flag at the slice level, namely, slice_disable_bdof_prof_dmvr_flag to disable BDOF, PROF and DMVR. The SPS flag sps_bdof_prof_dmvr_slice_present_flag is used to indicate the presence of slice_disable_bdof_prof_dmvr_flag, which is signaled in the SPS when either DMVR or BDOF / PROF SPS level control flag is true. If present, slice_disable_bdof_dmvr_flag is signaled. Table 2 illustrates the modified slice header semantics table after applying the proposed semantics. In another embodiment, it is proposed to still use these two control flags in the slice header to control the enable / disable of BDOF and DMVR and the enable / disable of PROF, respectively. Specifically, the method uses two flags in the slice header: a flag slice_disable_bdof_dmvr_slice_flag is used to control the opening / closing of BDOF and DMVR and another flag disable_prof_slice_flag is used to control the opening / closing of PROF alone.
[0239] Table 2 Modified SPS semantic table with proposed BDOF / PROF enable / disable flags
[0240]
[0241] In another embodiment, two separate SPS flags are used to control BDOF and PROF. Specifically, two separate SPS flags, sps_bdof_enable_flag and sps_prof_enable, are introduced to enable / disable the two tools, respectively. In addition, a high-level control flag, no_prof_constraint_flag, is added to the general_constrain_info() semantic table to forcibly disable the PROF tool.
[0242]
[0243] sps_bdof_enabled_flag indicates whether bidirectional optical flow is enabled. When sps_bdof_enabled_flag is equal to 0, bidirectional optical flow is disabled. When sps_bdof_enabled_flag is equal to 1, bidirectional optical flow is enabled.
[0244] sps_prof_enabled_flag indicates whether optical flow prediction refinement is enabled. When sps_prof_enabled_flag is equal to 0, optical flow prediction refinement is disabled. When sps_prof_enabled_flag is equal to 1, optical flow prediction refinement is enabled.
[0245]
[0246] no_prof_constraint_flag equal to 1 indicates that sps_prof_enabled_flag should be equal to 0. no_prof_constraint_flag equal to 0 does not impose constraints.
[0247] At the slice level, in one or more embodiments of the present application, it is proposed to introduce another control flag at the slice level, namely, slice_disable_bdof_prof_dmvr_flag for disabling BDOF, PROF and DMVR together. In another embodiment, it is proposed to add two separate flags at the slice level, namely, slice_disable_bdof_dmvr_flag and slice_disable_prof_flag. The first flag (i.e., slice_disable_bdof_dmvr_flag) is used to adaptively turn on / off BDOF and DMVR for a slice, and the second flag (i.e., slice_disable_prof_flag) is used to control the enabling and disabling of the PROF tool at the slice level. In addition, when the second method is applied, the flag slice_disable_bdof_dmvr_flag needs to be sent by signal only when the SPSBDOF or SPSDMVR flag is enabled, and the flag needs to be sent by signal only when the SPS PROF flag is enabled.
[0248] Figure 11 The method of BDOF and PROF is shown. For example, the method can be applied to a decoder.
[0249] In step 1110, the decoder may receive two general constraint information (GCI) level control flags. The two GCI level control flags are signaled by the encoder and may include a first GCI level control flag and a second GCI level control flag. The first GCI level control flag indicates whether the BDOF is allowed to decode the current video sequence. The second GCI level control flag indicates whether the PROF is allowed to decode the current video sequence.
[0250] In step 1112, the decoder may receive two SPS level control flags. These two SPS level control flags are signaled by the SPS encoder and indicate whether BDOF and PROF are enabled for the current video block.
[0251] In step 1114, when the first SPS level control flag is enabled, the decoder may apply BDOF to generate a prediction image based on the first prediction sample I when the video block is not encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives the motion refinement for this video block.
[0252] In step 1116, when the second SPS level control flag is enabled, the decoder may apply PROF to generate a prediction image based on the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derives the motion refinement for this video block.
[0253] In step 1118, the decoder may obtain prediction samples for the video block based on the motion refinements.
[0254] Early termination of PROF based on control point MV differences
[0255] According to the current PROF design, PROF is always called when predicting a coding block using affine mode. However, as shown in equations (6) and (7), the sub-block MVs of an affine block are derived from these control point MVs. Therefore, when the difference between the control point MVs is small, the MVs at each sample position should be consistent. In this case, the benefit of applying PROF may be very limited. Therefore, in order to further reduce the average computational complexity of PROF, it is proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between the sample-by-sample MV and the sub-block-by-sub-block MV within a 4×4 sub-block. Since the PROF MV differences of samples within a 4×4 sub-block are symmetric with respect to the sub-block center, the maximum horizontal and vertical PROF MV differences can be calculated based on equation (10) as:
[0256]
[0257]
[0258] Depending on the application, different metrics may be used to determine whether the MV difference is small enough to skip the PROF process.
[0259] In one example, based on equation (19), when the sum of the absolute maximum horizontal MV difference and the absolute maximum vertical MV difference is less than a predefined threshold, the PROF process can be skipped, that is,
[0260]
[0261] In another example, if and If the maximum value is not greater than a threshold, the PROF process can be skipped.
[0262]
[0263] Here, MAX(a, b) is a function that returns the larger value between input values a and b.
[0264] In addition to the two examples above, the concepts of this application are also applicable to situations where other metrics are used to determine whether the MV difference is small enough to skip the PROF process. In the above method, PROF is skipped based on the magnitude of the MV difference. On the other hand, in addition to the MV difference, the PROF sample refinement is calculated based on the local gradient information at each sample position in a motion-compensated block. For prediction blocks containing less high-frequency details (such as flat areas), the gradient values are often small, so the derived sample refinement value should be small. With this in mind, according to another embodiment, it is proposed to apply PROF only to the prediction samples of the block containing sufficient high-frequency information.
[0265] Different metrics can be used when determining whether a block contains enough high frequency information to make it worthwhile to call the PROF process for the block. In one example, the decision is made based on the average magnitude (i.e., absolute value) of the gradients of the samples within the prediction block. If the average magnitude is less than a threshold, the prediction block is classified as a flat area and the PROF should not be applied; otherwise, the prediction block is considered to contain enough high frequency details and the PROF is still applicable. In another example, the maximum magnitude of the gradients of the samples within the prediction block can be used. If the maximum magnitude is less than a threshold, the PROF is skipped for the block. In yet another example, the difference between the maximum sample value and the minimum sample value of the prediction block is 1. max -I min It can be used to determine whether to apply the PROF to the block. If the difference is less than a threshold, the PROF is skipped for the block. It is worth noting that the concept of the present application is also applicable to the case where other metrics are used to determine whether a given block contains sufficient high-frequency information.
[0266] Handling the interaction between PROF and LIC for affine mode
[0267] Since the neighboring reconstructed samples (i.e., templates) of the current block are used by LIC to derive the linear model parameters, the decoding of a LIC-coded block depends on the complete reconstruction of its neighboring samples. Due to this interdependence, for practical hardware implementations, it is necessary to perform LIC in the reconstruction phase, where the neighboring reconstructed samples are available for LIC parameter derivation. Because block reconstruction must be performed sequentially (i.e., one after another), throughput (i.e., the amount of work that can be done in parallel per unit time) is an important issue to consider when jointly applying other coding methods to LIC-coded blocks. In this section, two methods are proposed to handle the interaction when PROF and LIC are both enabled for affine mode.
[0268] In the first embodiment of the present application, it is proposed to apply the PROF mode and the LIC mode exclusively to an affine coding block. As previously mentioned, in existing designs, PROF is implicitly applied to all affine blocks without signaling, while a LIC flag is marked or inherited at the coding block level to indicate whether the LIC mode is applied to an affine block. According to the method of the present application, it is proposed to conditionally apply PROF based on the value of the LIC flag of an affine block. When the flag is equal to 1, only LIC is applied by adjusting the prediction samples of the entire coding block based on the LIC weights and offsets. Otherwise (i.e., the LIC flag is equal to 0), PROF is applied to the affine coding block to refine the prediction samples of each sub-block based on the optical flow model.
[0269] Figure 17A An exemplary flow chart of the decoding process based on the proposed method is shown, where simultaneous application of PROF and LIC is not allowed.
[0270] Figure 17A A diagram illustrates a decoding process based on the proposed method according to the present application, wherein PROF and LIC are disabled. Decoding process 1720 includes steps 1722 (check if LIC flag is on), LIC 1724, and PROF 1726. Step 1722 (check if LIC flag is on) determines whether the LIC flag is set and takes the next step based on this determination. LIC 1724 applies LIC when the LIC flag is set. PROF 1726 applies PROF when the LIC flag is not set.
[0271] In the second embodiment of the present application, it is proposed to apply LIC after PROF to generate prediction samples for an affine block. Specifically, after completing the affine motion compensation based on the sub-block, these prediction samples are refined based on the PROF sample refinement; then, LIC is performed by applying a pair of weights and offsets (derived from the template and its reference samples) to the PROF-adjusted prediction samples to obtain the final prediction samples of the block, as shown below:
[0272] P[x]=α*(P r [x+v]+ΔI[x])+β (22);
[0273] Among them, P r [x+v] is the reference block of the current block indicated by motion vector v; α and β are the LIC weights and offsets; P[x] is the final prediction block; ΔI[x] is the PROF refinement derived in (15).
[0274] Figure 17B A diagram illustrates a decoding process according to the present application in which PROF and LIC are applied. Decoding process 1760 includes affine motion compensation 1762, LIC parameter derivation 1764, PROF 1766, and LIC sample adjustment 1768. Affine motion compensation 1762 applies affine motion and is an input to LIC parameter derivation 1764 and PROF 1766. LIC parameter derivation 1764 is used to derive LIC parameters. PROF 1766 applies PROF. LIC sample adjustment 1768 is the LIC weight parameters and offset parameters combined with PROF.
[0275] Figure 17B An exemplary decoding workflow when the second method is applied is shown. Figure 17B As shown in Figure 2, since the LIC uses the template (i.e., adjacent reconstructed samples) to calculate the LIC linear model, the LIC parameters can be derived immediately as soon as the adjacent reconstructed samples are available. This means that PROF refinement and LIC parameter derivation can be performed simultaneously.
[0276] LIC weights and offsets (i.e., α and β) and PROF refinement (i.e., ΔI[x]) are usually floating point numbers. For hardware-friendly implementation, these floating point operations are usually implemented as an integer value multiplication followed by a multiple bit right shift operation. In existing LIC and PROF designs, since these two tools are designed separately, N is applied in two stages. LIC Bits and N PROF Two different right shifts of bits.
[0277] According to the third embodiment, in order to improve the coding gain when PROF and LIC are jointly applied to affine coded blocks, it is proposed to apply LIC-based and PROF-based sample adjustments with high accuracy. This is done by combining their two right shift operations into one and applying it at the end to derive the final prediction samples of the current block (as shown in (12)).
[0278] Resolve multiplication overflow issues when combining PROF with weighted prediction and CU-level weighted bidirectional prediction (BCW)
[0279] According to the PROF design in the current VVC working draft, PROF can be used in conjunction with weighted prediction (WP). Specifically, when they are combined, the prediction signal of an affine CU will be generated through the following process:
[0280] First, for each sample at position (x, y), an L0 prediction refinement ΔI0(x, y) is computed based on the PROF and added to the original L0 prediction sample I0(x, y), i.e.,
[0281] ΔI0(x, y)=(g h0 (x, y)·Δv x0 (x, y) + g v0 (x, y)·Δv y0 (x, y)+1)>>1
[0282] I0′(x, y)=I0(x, y)+ΔI0(x, y) (23);
[0283] Among them, I0′(x, y) is the refined sample; g h0 (x, y) and g v0 (x, y) is the L0 horizontal / vertical gradient and L0 horizontal / vertical motion refinement at position (x, y).
[0284] Second, for each sample at position (x, y), calculate the L1 prediction refinement ΔI1(x, y) based on the PROF and add the refinement to the original L1 prediction sample I1(x, y), i.e.,
[0285] ΔI1(x, y)=(g h1 (x, y)·Δv x1 (x, y) + g v1 (x, y)·Δv y1 (x, y)+1)>>1
[0286] I1′(x, y)=I1(x, y)+ΔI1(x, y) (24);
[0287] Among them, I1′(x, y) is the refined sample; g h1 (x, y) and g v1 (x, y) and Δv x1 (x, y) and Δv y1 (x, y) is the L1 horizontal / vertical gradient and L1 horizontal / vertical motion refinement at position (x, y).
[0288] Third, combine the refined L0 and L1 prediction samples, i.e.,
[0289] I bi(x, y)=(W0·I0′(x, y)+W1·I1′(x, y)+Offset)>>shift (25);
[0290] Where W0 and W1 are the WP and BCW weights; shift and offset are the offset and right shift applied to the weighted average of the L0 and L1 prediction signals for bidirectional prediction of WP and BCW. Here, the parameters for WP include W0, W1 and Offset, while the parameters for BCW include W0, W1 and shift.
[0291] As can be seen from the above equations, due to the sample-by-sample refinement, i.e., ΔI0(x, y) and ΔI1(x, y), the predicted samples after PROF (i.e., I0′(x, y) and I1′(x, y)) will have an increased dynamic range compared to the original predicted samples (i.e., I0(x, y) and I1(x, y)). Given that these refined prediction samples will be multiplied by the WP and BCW weighting factors, this will increase the length of the required multiplier. For example, based on the current design, when the intra-coding bit depth is 8-12 bits, the dynamic range of the predicted signals I0(x, y) and I1(x, y) is 16 bits. However, after this PROF, the dynamic range of the predicted signals I0′(x, y) and I1′(x, y) is 17 bits. Therefore, when this PROF is applied, a 16-bit multiplication overflow problem may occur. In order to solve this overflow problem, several methods are proposed below:
[0292] First, in the first approach, it is proposed to disable WP and BCW when applying the PROF to an affine CU.
[0293] Second, in the second method, it is proposed to apply a clipping operation to the derived sample refinements before adding them to the original prediction samples, so that the dynamic ranges I0′(x, y) and I1′(x, y) of these refined prediction samples have the same dynamic bit depth as the original prediction samples I0(x, y) and I1(x, y). Specifically, in this method, the sample refinements ΔI0(x, y) and ΔI1(x, y) in (23) and (24) are modified by introducing a clipping operation as follows:
[0294] ΔI0(x, y)=clip3(-2 dI-1 , 2 dI-1 -1, ΔI0(x, y)),
[0295] ΔI1(x, y)=clip3(-2 dI-1 , 2 dI-1 -1, ΔI1(x, y));
[0296] Where dI = dI base+max(0, BD-12), where BD is the intra-coding bit depth; dI base is the basic bit depth value. In one or more embodiments, it is proposed to set dI base The value of dI is set to 14. In another embodiment, it is proposed to set the value to 13. In one or more embodiments, it is proposed to directly set the value of dI to a fixed value. In one example, it is proposed to set the value of dI to 13, that is, the sample is thinned and clipped to the range of [-4096, 4095]. In another example, it is proposed to set the value of dI to 14, that is, the sample is thinned and clipped to the range of [-8192, 8191].
[0297] Figure 10 The method of PROF is shown. For example, the method can be applied to a decoder.
[0298] In step 1010, the decoder may obtain a first reference picture I associated with a video block encoded in an affine mode within a video signal. (0) and the second reference picture I (1) .
[0299] In step 1012, the decoder may generate a decoded image based on the first reference picture I (0) and the second reference picture I (1) The first prediction sample I associated with (0) (i, j) and the second prediction sample I (1) (i, j) obtain a first horizontal gradient value, a second horizontal gradient value, a first vertical gradient value, and a second vertical gradient value.
[0300] In step 1014, the decoder may generate a decoded image based on the first reference picture I (0) and the second reference picture I (1) The associated CPMV obtains a first horizontal motion refinement, a second horizontal motion refinement, a first vertical motion refinement, and a second vertical motion refinement.
[0301] In step 1016, the decoder may obtain a first prediction refinement ΔI based on the first horizontal gradient value, the second horizontal gradient value, the first vertical gradient value, the second vertical gradient value, the first horizontal motion refinement, the second horizontal motion refinement, the first vertical motion refinement, and the second vertical motion refinement. (0) (i, j) and the second prediction refinement ΔI (1) (i, j).
[0302] In step 1018, the decoder may generate a prediction result based on the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), first prediction refinement ΔI (0)(i, j), second prediction refinement ΔI (1) The final prediction sample of the video block is obtained by adding (i, j) and prediction parameters. These prediction parameters may include weighting parameters and offset parameters for WP and BCW.
[0303] First, in the third method, it is proposed to directly crop these refined prediction samples instead of cropping these sample refinements so that these refined samples have the same dynamic range as the original prediction samples. Specifically, through the third method, these refined L0 and L1 samples will be:
[0304] I0′(x, y)=clip3(-2 dR , 2 dR -1,I0(x,y)+ΔI0(x,y)),
[0305] I1′(x, y)=clip3(-2 dR , 2 dR -1,I1(x,y)+ΔI1(x,y));
[0306] Where dR = 16 + max(0, BD - 12) (or equivalently max(16, BD + 4)), where BD is the intra codec bit depth. In one or more embodiments, it is proposed to hard-clip these refined PROF prediction samples to 16 bits, i.e., set the value of dR to 15. In another embodiment, it is proposed to clip these refined PROF sample values to the same dynamic range [na, nb] of the initial prediction samples I0 and I1 before PROF, where na and nb are the minimum and maximum extreme values that these initial prediction samples can reach.
[0307] Second, in the fourth method, it is proposed to apply some right shifts to these refined L0 and L1 prediction samples before WP and BCW; then the final prediction samples are adjusted to the original precision through additional left shifts. Specifically, the final prediction samples are derived as:
[0308] I bi (x, y)=(W0·(I0′(x,y)>>nb)+W1·(I1′(x,y)>>nb)+Offset)>>(shift-nb);
[0309] where nb is the number of additional shifts applied, which can be determined based on the corresponding dynamic range of these PROF sample refinements.
[0310] Third, in the fifth method, it is proposed to divide each multiplication of the L0 / L1 prediction sample and the corresponding WP / BCW weight in formula (25) into two multiplications, and both multiplications do not exceed 16 bits, which can be described as:
[0311] I bi (x, y)=(W0·I0(x, y)+W0·ΔI0(x, y)+W1·I1(x, y)+W1·ΔI1(x, y)+Offset)>>shift.
[0312] The methods described above may be implemented using an apparatus comprising one or more circuits, including application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components. The apparatus may use these circuits in combination with other hardware or software components to perform the methods described above. Each module, submodule, unit, or subunit disclosed above may be implemented, at least in part, using one or more circuits.
[0313] Figure 19 19 is a diagram illustrating a computing environment coupled with a user interface according to an example of the present application. The computing environment 1910 may be part of a data processing server. The computing environment 1910 includes a processor 1920, a memory 1940, and an input / output (I / O) interface 1950.
[0314] The processor 1920 generally controls the overall operation of the computing environment 1910, such as operations associated with display, data acquisition, data communication, and image processing. The processor 1920 may include one or more processors for executing instructions to perform all or some of the steps in the above-described method. In addition, the processor 1920 may include one or more modules that facilitate interaction between the processor 1920 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip microcomputer, a GPU, etc.
[0315] The memory 1940 is configured to store various types of data to support the operation of the computing environment 1910. The memory 1940 may include predetermined software 1932. Examples of such data include instructions for any application or method operating on the computing environment 1910, video data sets, image data, etc. The memory 1940 may be implemented using any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0316] I / O interface 1950 provides an interface between processor 1920 and peripheral interface modules (e.g., keyboard, click wheel, buttons, etc.). Buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. I / O interface 1950 may be coupled to an encoder and a decoder.
[0317] In some embodiments, a non-transitory computer-readable storage medium including a plurality of programs in, for example, a memory 1940 is also provided, and the plurality of programs can be executed by the processor 1920 in the computing environment 1910 to perform the above-described method. For example, the non-transitory computer-readable storage medium can be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0318] The non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, causes the computing device to perform the above-mentioned motion prediction method.
[0319] In an embodiment, the computing environment 1910 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0320] The description of the present application has been presented for purposes of illustration and is not intended to be exhaustive or limited to the present application. Many modifications, variations, and alternative embodiments will be apparent to one of ordinary skill in the art having the benefit of the teachings presented in the foregoing description and the associated drawings.
[0321] The examples are chosen and described in order to explain the principles of the present application and to enable others skilled in the art to understand the various embodiments of the present application and to best utilize the basic principles and various embodiments with various modifications as are suited to the particular use contemplated. Therefore, it will be understood that the scope of the present application is not limited to the specific examples of the embodiments disclosed and that modifications and other embodiments are intended to be included within the scope of the present application.
Claims
1. A method for bidirectional optical flow (BDOF) and optical flow prediction refinement (PROF), implemented by an encoder, comprising: signaling two general constraint information (GCI) level control flags, wherein the two GCI level control flags include a first GCI level control flag and a second GCI level control flag, wherein the first GCI level control flag indicates whether the BDOF is enabled for a current video sequence, and the second GCI level control flag indicates whether the PROF is enabled for the current video sequence; signaling two sequence parameter set (SPS) level control flags, wherein the two SPS level control flags include a first SPS level control flag and a second SPS level control flag, wherein the first SPS level control flag indicates whether BDOF is enabled for a current video block and the second SPS level control flag indicates whether PROF is enabled for the current video block; The first SPS level control flag indicates that BDOF is enabled for the current video block, based on determining that BDOF is applied based on the first prediction sample I when the video block is not encoded in the affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) deriving a motion refinement of the video block; The second SPS level control flag indicates that PROF is enabled for the current video block, based on determining that PROF is applied to the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derive a motion refinement for the video block.
2. The method according to claim 1, further comprising: When the first GCI level control flag disables the BDOF, applying a first bitstream consistency constraint requiring the first SPS level control flag to be zero; as well as When the second GCI level control flag disables the PROF, a second bitstream conformance constraint requiring the second SPS level control flag to be zero applies.
3. The method according to claim 1, further comprising: When the first SPS level control flag is enabled, signaling a first control flag in a slice header, wherein the first control flag indicates whether the BDOF is disabled for a video block in the slice header; and When the second SPS level control flag is enabled, a second control flag in the slice header is signaled, wherein the second control flag indicates whether the PROF is enabled for the video block in the slice header.
4. A computing device comprising: one or more processors; a memory coupled to the one or more processors; as well as A plurality of computer programs stored in the memory, which, when executed by the one or more processors, cause the computing device to perform the following operations: signaling two general constraint information (GCI) level control flags, wherein the two GCI level control flags include a first GCI level control flag and a second GCI level control flag, wherein the first GCI level control flag indicates whether BDOF is enabled for a current video sequence, and the second GCI level control flag indicates whether PROF is enabled for the current video sequence; signaling two sequence parameter set (SPS) level control flags, wherein the two SPS level control flags include a first SPS level control flag and a second SPS level control flag, wherein the first SPS level control flag indicates whether BDOF is enabled for a current video block and the second SPS level control flag indicates whether PROF is enabled for the current video block; The first SPS level control flag indicates that BDOF is enabled for the current video block, based on determining that BDOF is applied based on the first prediction sample I when the video block is not encoded in the affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) deriving a motion refinement of the video block; The second SPS level control flag indicates that PROF is enabled for the current video block, based on determining that PROF is applied to the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derive a motion refinement for the video block.
5. The computing device of claim 4, wherein: The plurality of computer programs further cause the computing device to: When the first GCI level control flag disables the BDOF, applying a first bitstream conformance constraint requiring the first SPS level control flag to be zero; and When the second GCI level control flag disables the PROF, a second bitstream conformance constraint requiring the second SPS level control flag to be zero applies. The computing device according to claim 4 , wherein: The plurality of computer programs further cause the computing device to: When the first SPS level control flag is enabled, signaling a first control flag in a slice header, wherein the first control flag indicates whether the BDOF is disabled for a video block in the slice header; and When the second SPS level control flag is enabled, a second control flag in the slice header is signaled, wherein the second control flag indicates whether the PROF is enabled for the video block in the slice header.
7. A non-transitory computer-readable storage medium storing a bitstream including video data, causing an encoding device to perform the following operations: Two General Constraint Information (GCI) level control flags are signaled, where The two GCI level control flags include a first GCI level control flag and a second GCI level control flag, wherein the first GCI level control flag indicates whether BDOF is enabled for the current video sequence, and the second GCI level control flag indicates whether PROF is enabled for the current video sequence; signaling two sequence parameter set (SPS) level control flags, wherein the two SPS level control flags include a first SPS level control flag and a second SPS level control flag, wherein the first SPS level control flag indicates whether BDOF is enabled for a current video block and the second SPS level control flag indicates whether PROF is enabled for the current video block; The first SPS level control flag indicates that BDOF is enabled for the current video block, based on determining that BDOF is applied based on the first prediction sample I when the video block is not encoded in the affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) deriving a motion refinement of the video block; The second SPS level control flag indicates that PROF is enabled for the current video block, based on determining that PROF is applied to the first prediction sample I when the video block is encoded in affine mode. (0) (i, j) and the second prediction sample I (1) (i, j) derive a motion refinement for the video block.
8. The non-transitory computer-readable storage medium of claim 7, wherein: The encoding device is further configured to perform the following operations: When the first GCI level control flag disables the BDOF, applying a first bitstream conformance constraint requiring the first SPS level control flag to be zero; and When the second GCI level control flag disables the PROF, a second bitstream conformance constraint requiring the second SPS level control flag to be zero applies.
9. The non-transitory computer-readable storage medium of claim 7, wherein: The encoding device is further configured to perform the following operations: When the first SPS level control flag is enabled, signaling a first control flag in a slice header, wherein the first control flag indicates whether the BDOF is disabled for a video block in the slice header; and When the second SPS level control flag is enabled, a second control flag in the slice header is signaled, wherein the second control flag indicates whether the PROF is enabled for the video block in the slice header.
Citation Information
Patent Citations
Video cover selection method and device, computer equipment and storage medium
CN108833942A
Advanced motion vector prediction speedups for video coding
US20190230376A1