Method and apparatus for predictive refinement using optical flow, bidirectional optical flow, and decoder-side motion vector refinement.

By controlling bit depth and integrating decoder-side motion vector refinement, the inefficiencies in BDOF and PROF are addressed, enhancing coding efficiency and accuracy in video coding standards.

JP2026048758APending Publication Date: 2026-03-17BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing video coding standards like VVC face inefficiencies in motion compensation due to limitations in bit depth control for bidirectional optical flow (BDOF) and predictive refinement with optical flow (PROF), leading to suboptimal hardware implementation and coding accuracy.

Method used

Implement bit depth control mechanisms for BDOF and PROF by applying right shifts to internal parameters, using precision adjustments for gradient and motion difference values, and integrating decoder-side motion vector refinement to align with intermediate high bit depth accuracy.

Benefits of technology

Enhances coding efficiency and accuracy by harmonizing BDOF and PROF designs, allowing for shared hardware pipelines and improved motion compensation in video coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026048758000001_ABST
    Figure 2026048758000001_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, and non-temporary computer-readable storage medium for bit depth control of bidirectional optical flow (BDOF). [Solution] The method includes the decoder acquiring a reference picture I associated with a video block having a video signal; acquiring predicted samples I(i,j) of the video block from the reference block in the reference picture I; controlling the internal PROF parameters of the internal predictive refinement (PROF) process by applying a right shift to the internal predictive refinement (PROF) parameters based on a bit shift value to achieve a preset accuracy; acquiring predictive refinement values ​​for the samples in the video block based on the PROF derivation process being applied to the video block based on the predicted samples I(i,j); and acquiring predicted samples of the video block based on a combination of the predicted samples and the predictive refinement values.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to Provisional Patent No. 62 / 913,141, filed on 9 October 2019, which in its entirety is incorporated herein by reference for all purposes.

[0002] This disclosure relates to video coding and compression. More specifically, this disclosure relates to methods and apparatus for two interprediction tools being studied in the Multipurpose Video Coding (VVC) standard: prediction refinement with optical flow (PROF) and bi-directional optical flow (BDOF). [Background technology]

[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include General Purpose Video Coding (VVC), Joint Search Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Expert Group (MPEG) coding. Video coding generally utilizes predictive methods (e.g., inter-prediction, intra-prediction, etc.) that leverage the redundancy present in video images or sequences. A key goal of video coding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. [Overview of the project] [Problems that the invention aims to solve]

[0004] Examples of the present disclosure provide methods and apparatus for bit depth control for bidirectional optical flow. [Means for solving the problem]

[0005] A first aspect of this disclosure provides a method for representing the bit depth of a PROF. This method may include a decoder obtaining a reference picture I associated with a video block in a video signal. The decoder may also obtain a predicted sample I(i,j) of the video block from the reference block in the reference picture I, where i and j represent the coordinates of a single sample having the video block. The decoder may further control the internal PROF parameters of the PROF derivation process by applying a right shift to the internal PROF parameters based on a bit shift value to achieve a pre-set precision. The internal PROF parameters may include a horizontal gradient value, a vertical gradient value, a horizontal motion difference value, and a vertical motion difference value derived for the predicted sample I(i,j). The decoder may also obtain a predicted refined value for the sample in the video block based on the PROF derivation process being applied to the video block based on the predicted sample I(i,j). The decoder may further obtain a predicted sample of the video block based on a combination of the predicted sample and the predicted refined value.

[0006] A second aspect of this disclosure provides a method for controlling the bit depth of a BDOF. The method involves a decoder associated with a first reference picture I video block. (0) and second reference picture I (1) This may include obtaining the first reference picture I in the display order. (0) It is before the current picture, and is the second reference picture I (1) It is located after the current picture. The decoder also references the first reference picture I (0) First predicted sample I of the video block from the reference block within (0) The decoder can obtain (i,j), where i and j represent the coordinates of one sample with the current picture. (1) Second predicted sample I of the video block from the reference block within (1)(i,j) can be further obtained. The decoder can also control the internal BDOF parameters of the BDOF derivation process by applying a shift to the internal BDOF parameters. The internal BDOF parameters include the first prediction sample I (0) (i,j), the second prediction sample I (1) (i,j), the first prediction sample I (0) (i,j) and the second prediction sample I (1) (i,j), the sample differences between them, and the horizontal and vertical gradient values derived based on the intermediate BDOF derivation parameters. The intermediate BDOF derivation parameters can include sGxdI, sGydI, sGx2, sGxGy, and sGy2 parameters. sGxdI and sGydI can include the cross-correlation values between the horizontal gradient value and the sample difference value and between the vertical gradient value and the sample difference value. sGx2 and sGy2 can include the auto-correlation values of the horizontal and vertical gradient values. sGxGy can include the cross-correlation value between the horizontal gradient value and the vertical gradient value.

[0007] A third aspect of this disclosure provides BDOF, PROF, and DMVR methods. The method may include a decoder receiving three control flags in a sequence parameter set (SPS). A first control flag indicates whether BDOF is enabled to decode video blocks in the current video sequence. A second control flag indicates whether PROF is enabled to decode video blocks in the current video sequence. A third control flag indicates whether DMVR is enabled to decode video blocks in the current video sequence. The decoder may also receive a first presence flag in the SPS when the first control flag is true, a second presence flag in the SPS when the second control flag is true, and a third presence flag in the SPS when the third control flag is true. The decoder may further receive a first picture control flag in the picture header of each picture when the first presence flag in the SPS indicates that BDOF is disabled for video blocks in the picture. The decoder may also receive a second picture control flag in the picture header of each picture when a second presence flag in the SPS indicates that PROF is disabled for the video block in the picture. The decoder may further receive a third picture control flag in the picture header of each picture when a third presence flag in the SPS indicates that DMVR is disabled for the video block in the picture.

[0008] A computing device is provided according to a fourth aspect of the present disclosure. The computing device may include one or more processors and non-temporary computer-readable memory for storing instructions executable by one or more processors. One or more processors may be configured to acquire a reference picture I associated with a video block in a video signal. One or more processors may also be configured to acquire a predicted sample I(i,j) of a video block from a reference block in the reference picture I, where i and j represent the coordinates of a single sample in the video block. One or more processors may be further configured to control internal PROF parameters of a PROF derivation process by applying a right shift to internal PROF parameters based on a bit shift value to achieve a preset precision. Internal PROF parameters may include a horizontal gradient value, a vertical gradient value, a horizontal motion difference value, and a vertical motion difference value derived for the predicted sample I(i,j). One or more processors may also be configured to acquire a predicted refined value for a sample in a video block based on the fact that the PROF derivation process has been applied to the video block based on the predicted sample I(i,j). One or more processors may be further configured to acquire predicted samples of video blocks based on a combination of predicted samples and predicted refined values.

[0009] A computing device is provided according to a fifth aspect of this disclosure. The computing device may include one or more processors and non-temporary computer-readable memory that stores instructions executable by one or more processors. One or more processors have a first reference picture I associated with a video block. (0) and second reference picture I (1) It can be configured to obtain the first reference picture I in the display order. (0) It is before the current picture, and is the second reference picture I (1) It is after the current picture. One or more processors refer to the first reference picture I (0)First predicted sample I of the video block from the reference block within (0) It may be further configured to obtain (i,j), where i and j represent the coordinates of one sample having the current picture. One or more processors may also be configured to obtain predicted refined values ​​for samples in a video block, based on the PROF derivation process being applied to the video block based on the predicted sample I(i,j). One or more processors may obtain a second reference picture I (1) Second predicted sample I of the video block from the reference block within (1) The system may be further configured to obtain (i,j). One or more processors may be further configured to control the internal BDOF parameters of the BDOF derivation process by applying a shift to the internal BDOF parameters. The internal BDOF parameters are the first predicted sample I (0) (i,j), Second prediction sample I (1) (i,j), First prediction sample I (0) (i,j) and the second predicted sample I (1) The intermediate BDOF derivation parameters include the sample difference between (i,j) and the horizontal and vertical gradient values ​​derived based on the intermediate BDOF derivation parameters. The intermediate BDOF derivation parameters may include the parameters sGxdI, sGydI, sGx2, sGxGy, and sGy2. sGxdI and sGydI may include the cross-correlation values ​​between the horizontal gradient value and the sample difference value, and between the vertical gradient value and the sample difference value. sGx2 and sGy2 may include the autocorrelation values ​​between the horizontal and vertical gradient values. sGxGy may include the cross-correlation value between the horizontal and vertical gradient values. One or more processors process the first predicted sample I (0) (i,j) and the second predicted sample I (1) Based on the fact that BDOF has been applied to the video block based on (i,j), it may be further configured to obtain motion refinement for the samples in the video block. One or more processors may be further configured to obtain bipredicted samples of the video block based on the motion refinement.

[0010] According to a sixth aspect of this disclosure, a non-temporary computer-readable storage medium storing instructions is provided. When an instruction is executed by one or more processors of the device, the instruction may cause the device to receive three control flags in a sequence parameter set (SPS). A first control flag indicates whether BDOF is enabled to decode video blocks in the current video sequence. A second control flag indicates whether PROF is enabled to decode video blocks in the current video sequence. A third control flag indicates whether DMVR is enabled to decode video blocks in the current video sequence. The instruction may also cause the device to receive a first presence flag in the SPS when the first control flag is true, a second presence flag in the SPS when the second control flag is true, and a third presence flag in the SPS when the third control flag is true. The instruction may further cause the device to receive a first picture control flag in the picture header of each picture when the first presence flag in the SPS indicates that BDOF is disabled for video blocks in the picture. The instruction may also cause the device to receive a second picture control flag in the picture header of each picture when a second presence flag in the SPS indicates that PROF is disabled for the video blocks in the picture. The instruction may further cause the device to receive a third picture control flag in the picture header of each picture when a third presence flag in the SPS indicates that DMVR is disabled for the video blocks in the picture.

[0011] Please understand that both the above general explanation and the following detailed explanation are for illustrative purposes only and do not limit this disclosure.

[0012] The accompanying drawings incorporated herein and constituting part of herein illustrate examples consistent with this disclosure and, together with the descriptions, help to illustrate the principles of this disclosure. [Brief explanation of the drawing]

[0013] [Figure 1]This is a block diagram of an encoder relating to an example of this disclosure. [Figure 2] This is a block diagram of a decoder relating to an example of this disclosure. [Figure 3A] This figure illustrates an example of block division in a multi-type tree structure related to this disclosure. [Figure 3B] This figure illustrates an example of block division in a multi-type tree structure related to this disclosure. [Figure 3C] This figure illustrates an example of block division in a multi-type tree structure related to this disclosure. [Figure 3D] This figure illustrates an example of block division in a multi-type tree structure related to this disclosure. [Figure 3E] This figure illustrates an example of block division in a multi-type tree structure related to this disclosure. [Figure 4] This is an example diagram of a bidirectional optical flow (BDOF) model relating to an example of this disclosure. [Figure 5A] This is an example of an affine model related to an example of this disclosure. [Figure 5B] This is an example of an affine model related to an example of this disclosure. [Figure 6] This is an example of an affine model related to an example of this disclosure. [Figure 7] This is an example of predictive refinement (PROF) using optical flow related to an example of this disclosure. [Figure 8] This is a BDOF workflow related to an example of this disclosure. [Figure 9] This is a PROF workflow related to an example of this disclosure. [Figure 10] This is an example of a BDOF method related to this disclosure. [Figure 11] This is an example of a BDOF and PROF method relating to this disclosure. [Figure 12] This disclosure provides an example of a BDOF, PROF, and DMVR method. [Figure 13] This is an example of a PROF workflow for dual prediction related to an example of this disclosure. [Figure 14] This is an example of the pipeline stages of the BDOF and PROF processes related to this disclosure. [Figure 15] This is an example of a method for deriving the gradient of the BDOF related to this disclosure. [Figure 16] This is an example of a method for deriving the gradient of PROF related to this disclosure. [Figure 17A] This is an example of deriving a template sample for affine mode related to an example of this disclosure. [Figure 17B] This is an example illustrating the derivation of a template sample for affine mode related to an example of this disclosure. [Figure 18A] This is an example of exclusively enabling PROF and LIC for affine mode according to an example of this disclosure. [Figure 18B] This is an example of enabling PROF and LIC together for affine mode as described in this disclosure. [Figure 19A] This figure illustrates a proposed padding method applicable to a 16×16 BDOF CU as an example of the present disclosure. [Figure 19B] This figure illustrates a proposed padding method applicable to a 16×16 BDOF CU as an example of the present disclosure. [Figure 19C] This figure illustrates a proposed padding method applicable to a 16×16 BDOF CU as an example of the present disclosure. [Figure 19D] This figure illustrates a proposed padding method applicable to a 16×16 BDOF CU as an example of the present disclosure. [Figure 20] This figure illustrates a computing environment coupled with a user interface related to an example of this disclosure. [Modes for carrying out the invention]

[0014] Next, references to exemplary embodiments illustrated in the accompanying drawings are made in detail. The following description refers to the accompanying drawings, where, unless otherwise indicated, the same number in different drawings represents the same or similar elements. The implementations described below in the description of exemplary embodiments do not represent all implementations consistent with the present disclosure. Rather, these implementations are merely examples of apparatus and methods consistent with aspects relating to the present disclosure, such as those enumerated in the accompanying claims.

[0015] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit this disclosure. Where used in this disclosure and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural form unless the context otherwise explicitly indicates. The terms "and / or" as used herein are intended to mean and include any or all possible combinations of one or more of the related enumerated items.

[0016] Terms such as “first,” “second,” and “third” may be used herein to describe various types of information, but it should be understood that the information should not be limited by these terms. These terms are used solely to distinguish one category of information from another category of information. For example, without departing from the scope of this disclosure, first information may be referred to as second information, and similarly, second information may be referred to as first information. Where used herein, the term “if” may be understood, depending on the context, to mean “when,” “upon,” or “in response to a judgment.”

[0017] The first version of the HEVC standard, which offers approximately 50% bitrate savings or equivalent perceived quality compared to the previous generation video encoding standard H.264 / MPEG AVC, was finalized in October 2013. While the HEVC standard offers significant encoding improvements over its predecessor, there is evidence that even better encoding efficiency can be achieved with additional encoding tools. Based on this, both VECG and MPEG have begun exploring new encoding technologies for future video encoding standardization. One Joint Video Exploration Team (JVET) was formed in October 2015 by ITU-T VECG and ISO / IEC MPEG to initiate key research into cutting-edge technologies that could enable significant improvements in encoding efficiency. One reference software called the Joint Exploration Model (JEM) was maintained by JVET by integrating several additional encoding tools on top of the HEVC Test Model (HM).

[0018] In October 2017, a joint call for proposals (CfP) for video compression with capabilities exceeding HEVC was issued by the ITU-T and ISO / IEC[9]. In April 2018, 23 CfP responses were received and evaluated at the 10th JVET meeting, demonstrating compression efficiency gains approximately 40% higher than HEVC. Based on these evaluation results, JVET launched a new project to develop a new generation video coding standard called Versatile Video Coding (VVC). In the same month, a reference software codebase called the VVC Test Model (VTM) was established to demonstrate a reference implementation of the VVC standard.

[0019] Like HEVC, VVC is built on a block-based hybrid video encoding framework.

[0020] Figure 1 shows a schematic diagram of a block-based video encoder for VVC. Specifically, Figure 1 shows a typical encoder 100. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-mode determination 116, block predictor 140, adder 128, transform 130, quantization 132, predict-related information 142, intra-prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, in-loop filter 122, entropy coding 138, and bitstream 144.

[0021] In encoder 100, a video frame is divided into multiple video blocks for processing. For each given video block, a prediction is formed based on either an interpretation method or an intrapretation method.

[0022] A prediction residual, representing the difference between the current video block, which is part of the video input 110, and its predictor, which is part of the block predictor 140, is sent from the adder 128 to the transform 130. The transform coefficients are then sent from the transform 130 to the quantizer 132 for entropy reduction. The quantized coefficients are then fed to the entropy encoder 138 to generate a compressed video bitstream. As shown in Figure 1, prediction-related information 142 from the intra / inter-mode determination 116, such as video block segmentation information, motion vectors (MV), reference picture index, and intra-prediction mode, is also fed through the entropy encoder 138 and stored in the compressed bitstream 144. The compressed bitstream 144 contains the video bitstream.

[0023] In encoder 100, decoder-related circuit configurations are also required to reconstruct pixels for prediction purposes. First, the prediction residuals are reconstructed through inverse quantization 134 and inverse transform 136. These reconstructed prediction residuals are combined with block predictor 140 to generate unfiltered reconstructed pixels for the current video block.

[0024] Spatial prediction (or "intra prediction") uses pixels from already encoded samples of adjacent blocks (called reference samples) within the same video frame as the current video block to predict the current video block.

[0025] Time prediction (also called "interpretation") uses reconstructed pixels from already encoded video pictures to predict the current video block. Time prediction reduces the time redundancy inherent in video signals. A time prediction signal for a given encoding unit (CU) or encoding block is typically signaled by one or more MVs indicating the amount and direction of movement between the current CU and its time reference. In addition, if multiple reference pictures are supported, one reference picture index is transmitted, which is used to identify which reference picture in the reference picture storage the time prediction signal is coming from.

[0026] The motion estimation unit 114 takes in signals from the video input 110 and the picture buffer 120 and outputs a motion estimation signal to the motion compensation unit 112. The motion compensation unit 112 takes in signals from the video input 110, the picture buffer 120, and the motion estimation signal from the motion estimation unit 114 and outputs a motion compensation signal to the intra / inter-mode determination unit 116.

[0027] After spatial and / or temporal predictions are performed, an intra / inter mode determination 116 within the encoder 100 selects the best prediction mode, for example, based on a rate-distortion optimization method. The block predictor 140 is then subtracted from the current video block, and the resulting prediction residual is decorrelated using a transform 130 and quantization 132. The resulting quantized residual coefficients are inversely quantized by inverse quantization 134 and inversely transformed by inverse transform 136 to form a reconstructed residual, which is then added back to the prediction block to form a reconstructed signal for the CU. Furthermore, in-loop filtering 122, such as a deblocking filter, sample-adaptive offset (SAO), and / or adaptive in-loop filter (ALF), may be applied on the reconstructed CU before being placed into the reference picture storage of the picture buffer 120, and used to encode future video blocks. To form the output video bitstream 144, the encoding mode (inter or intra), prediction mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit 138 to form the bitstream, where they are further compressed and packed.

[0028] Figure 1 shows a block diagram of a typical block-based hybrid video coding system. The input video signal is processed block by block (called CUs). In VTM-1.0, a CU can be up to 128 x 128 pixels. However, unlike HEVC, which partitions blocks based solely on quadtrees, in VVC, a single coding tree unit (CTU) is split into CUs based on quadtrees / binary / ternary trees to fit varying local characteristics. In addition, the concept of multiple partitioned unit types in HEVC is removed; that is, the distinction between CUs, prediction units (PUs), and transformation units (TUs) no longer exists in VVC, and instead, each CU is always used as the basic unit for both prediction and transformation without further partitioning. In a multi-type tree structure, a single CTU is first partitioned by a quadtree structure. Then, each quadtree leaf node can be further partitioned by binary and ternary tree structures. As shown in Figures 3A, 3B, 3C, 3D, and 3E, there are five split types: quarter division, horizontal bisection, vertical bisection, horizontal third division, and vertical third division.

[0029] Figure 3A shows an example of block quartion in a multi-type tree structure according to this disclosure.

[0030] Figure 3B shows an example of block vertical bisection in a multi-type tree structure according to this disclosure.

[0031] Figure 3C illustrates an example of horizontal division of a block in a multi-type tree structure according to this disclosure.

[0032] Figure 3D illustrates an example of a block vertical tertiary division in a multi-type tree structure according to this disclosure.

[0033] Figure 3E illustrates an example of a block horizontal three-part division in a multi-type tree structure according to this disclosure.

[0034] In Figure 1, spatial and / or temporal predictions can be performed. Spatial prediction (or "intra-prediction") uses pixels from already encoded adjacent block samples (called reference samples) within the same video picture / slice to predict the current video block. Spatial prediction reduces the spatial redundancy inherent in the video signal. Temporal prediction (also called "inter-prediction" or "motion-compensated prediction") uses reconstructed pixels from already encoded video pictures to predict the current video block. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a given CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its time reference. Additionally, if multiple reference pictures are supported, one reference picture index is sent, used to identify which reference picture in the reference picture store the temporal prediction signal is coming from. After spatial and / or temporal predictions, a mode determination block within the encoder selects the best prediction mode, for example, based on a rate-distortion optimization method. Next, the predicted block is subtracted from the current video block, and the predicted residual is decorrelated using transformation and quantization. The quantized residual coefficients are inversely quantized and inversely transformed to form the reconstructed residual, which is then added back to the predicted block to form the reconstructed signal of the CU. Furthermore, in-loop filtering such as deblocking filters, sample-adaptive offsets (SAO), and adaptive in-loop filters (ALF) may be applied to the reconstructed CU before being placed in the reference picture store and used to encode future video blocks. To form the output video bitstream, the encoding mode (inter or intra), predicted mode information, motion information, and quantized residual coefficients are all sent to the entropy encoding unit to form the bitstream, where they are further compressed and packed.

[0035] Figure 2 shows a schematic block diagram of a video decoder for VVC. Specifically, Figure 2 shows a block diagram of a typical decoder 200. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter-mode selection 220, intra-prediction 222, memory 230, in-loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.

[0036] Decoder 200 is similar to the reconstruction-related section in encoder 100 in Figure 1. In decoder 200, the incoming video bitstream 210 is first decoded through entropy decoding 212 to derive quantized coefficient levels and prediction-related information. The quantized coefficient levels are then processed through inverse quantization 214 and inverse transform 216 to obtain reconstructed prediction residuals. The block predictor mechanism implemented in intra / inter-mode selector 220 is configured to perform either intra-prediction 222 or motion compensation 224 based on the decoded prediction information. The unfiltered reconstructed set of pixels is obtained by summing the reconstructed prediction residuals from the inverse transform 216 with the prediction output generated by the block predictor mechanism using an adder (summer) 218.

[0037] The reconstructed blocks may further pass through the in-loop filter 228 before being stored in the picture buffer 226, which functions as a reference picture store. The reconstructed video in the picture buffer 226 may be sent to drive the display device and may also be used to predict future video blocks. When the in-loop filter 228 is turned on, filtering operations are performed on these reconstructed pixels to derive the final reconstructed video output 232.

[0038] Figure 2 shows a schematic block diagram of a block-based video decoder. The video bitstream is first entropically decoded in the entropy decoding unit. The encoding mode and prediction information are sent to either the spatial prediction unit (if intra-encoded) or the temporal prediction unit (if inter-encoded) to form prediction blocks. The residual transformation coefficients are sent to the inverse quantization unit and the inverse transformation unit to reconstruct the residual blocks. The prediction blocks and residual blocks are then summed. The reconstructed blocks may further pass through in-loop filtering before being stored in the reference picture storage. The reconstructed video in the reference picture storage is then sent out to drive the display device and used to predict future video blocks.

[0039] In general, the basic inter-prediction techniques applied in VVC remain the same as those in HEVC, except that some modules are further extended and / or enhanced. In particular, for all preceding video standards, a single encoded block can only be associated with one MV when the encoded block is single-predicted, or with two MVs when the encoded block is bi-predicted. Such limitations of conventional block-based motion compensation mean that small motions may still remain in the predicted samples after motion compensation, thus negatively impacting the overall efficiency of motion compensation. To improve both the granularity and accuracy of MVs, two-sample-level refinement methods based on optical flow, namely bidirectional optical flow (BDOF) for affine modes and predictive refinement by optical flow (PROF), are currently being studied for the VVC standard. Below, the main technical aspects of the two inter-coding tools are briefly discussed.

[0040] Bidirectional optical flow In VVC, BDOF is applied to refine the predicted samples of the bipredicted coded blocks. Specifically, as shown in Figure 4, BDOF is a sample-by-sample motion refinement performed on top of block-based motion-compensated predictions when biprediction is used.

[0041] Figure 4 shows an example of a BDOF model related to this disclosure.

[0042] Refinement of the movement of each 4x4 subblock (v x ,v y ) is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample after BDOF has been applied within a single 6x6 window Ω around the subblock. Specifically, (v x ,v y The value of ) is

number

[0043] The values ​​of S1, S2, S3, S5, and S6 are:

number

number

number

[0044] Based on the motion refinement derived in (1), the final biprediction sample of CU is:

number

[0045] Affine Mode In HEVC, only translational motion models are applied to motion-compensated predictions. In contrast, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, affine motion-compensated predictions are applied by signaling one flag per inter-encoded block to indicate whether a translational motion model or an affine motion model is applied to the inter-prediction. In the current VVC design, two affine modes, including a 4-parameter affine mode and a 6-parameter affine mode, are supported for one affine-encoded block.

[0046] The four-parameter affine model has the following parameters: two parameters for translational motion in the horizontal and vertical directions, one parameter for zoom motion in both directions, and one parameter for rotational motion. The horizontal zoom parameter is equal to the vertical zoom parameter. The horizontal rotation parameter is equal to the vertical rotation parameter. To achieve better adaptation of motion vectors and affine parameters, in VVC, these affine parameters are transformed into two MVs (also called control point motion vectors (CPMVs)) located at the upper left and upper right corners of the current block. As shown in Figures 5A and 5B, the affine motion field of the block is described by the two control point MVs (V0,V1).

[0047] Figure 5A shows an example of a four-parameter affine model related to this disclosure.

[0048] Figure 5B shows an example of a four-parameter affine model related to this disclosure.

[0049] Based on the control point movement, the motion field of one affine-encoded block (v x ,v y )teeth,

number

[0050] The 6-parameter affine mode has the following parameters: two parameters for translational motion in the horizontal and vertical directions, one parameter for zoom motion and one parameter for rotational motion in the horizontal direction, and one parameter for zoom motion and one parameter for rotational motion in the vertical direction. The 6-parameter affine motion model is encoded using three MVs in three CPMVs.

[0051] Figure 6 shows an example of a six-parameter affine model related to this disclosure.

[0052] As shown in Figure 6, the three control points of a six-parameter affine block are located at the upper left, upper right, and lower left corners of the block. The motion at the upper left control point relates to translational motion, the motion at the upper right control point relates to horizontal rotational and zoom motion, and the motion at the lower left control point relates to vertical rotational and zoom motion. Compared to a four-parameter affine motion model, the horizontal rotational and zoom motions of a six-parameter model may not be the same as their vertical motions. Assuming (V0, V1, V2) are the motion vectors (MV) of the upper left, upper right, and lower left corners of the current block in Figure 6, the motion vectors (v) of each subblock are given by x ,v y ) uses three MVs at the control point

number

[0053] Predictive refinement using optical flow for affine mode To improve the accuracy of affine motion compensation, PROF, which refines subblock-based affine motion compensation based on an optical flow model, is currently being studied in VVC. Specifically, after performing subblock-based affine motion compensation, the luma prediction sample of one affine block is corrected by a single sample-refined value derived based on the optical flow equation. In detail, the operation of PROF can be summarized as follows:

[0054] Step 1: Subblock-based affine motion compensation is performed to generate subblock predictions I(i,j) using the subblock MV derived in (6) for 4-parameter affine models and in (7) for 6-parameter affine models.

[0055] Step 2: Spatial gradient g for each predicted sample x (i,j) and g y (i,j) is

number

[0056] To calculate the gradient, it is necessary to generate one additional row / column of predicted samples on each side of a single subblock. To reduce memory bandwidth and complexity, samples on the extended boundary are copied from the nearest integer pixel position in the reference picture to avoid additional interpolation processes.

[0057] Step 3: The refined value of the rumor prediction is,

number

[0058] Figure 7 illustrates a PROF process for affine mode according to the present disclosure. Figure 7 includes blocks 710, 720, and 730. Block 730 is a rotated block of block 720.

[0059] Since the affine model parameters and pixel locations relative to the subblock center do not change with each subblock, Δv(i,j) is calculated for the first subblock and can be reused for other subblocks within the same CU. If Δx and Δy are the horizontal and vertical offsets from the sample location (i,j) to the center of the subblock to which the sample belongs, then Δv(i,j) is:

number

[0060] Based on the affine subblock MV derivation equations (6) and (7), the MV difference Δv(i,j) can be derived. Specifically, in the case of a four-parameter affine model, The file is JPEG2026048758000016.jpg24155.

[0061] In the case of a 6-parameter affine model, JPEG2026048758000017.jpg43155, where, (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the control points MV of the top-left, top-right, and bottom-left of the current coded block, and w and h are the width and height of the block. In existing PROF designs, the MV difference Δv x and Δv y This is always derived with an accuracy of 1 / 32 Pell.

[0062] Local lighting compensation Local illumination compensation (LIC) is an encoding tool used to address the problem of local illumination changes that exist between time-adjacent pictures. A pair of weight and offset parameters is applied to a reference sample to obtain a predicted sample for one current block. A common mathematical model is:

number

number

[0063] In addition to being applied to normal interblocks containing at most one motion vector for each prediction direction (L0 or L1), LIC is also applied to affine-mode encoded blocks, where one encoded block is further split into multiple smaller subblocks, each subblock potentially associated with different motion information. To derive reference samples for LIC of affine-mode encoded blocks, as shown in Figures 17A and 17B (described below), reference samples in the upper template of one affine-coded block are fetched using the motion vectors of each subblock in the upper subblock row, while reference samples in the left template are fetched using the motion vectors of the subblocks in the left subblock column. The same LLMSE derivation method as shown in (12) is then applied to derive LIC parameters based on the composite template.

[0064] Figure 17A shows an example for deriving a template sample for affine mode according to the present disclosure. This example includes Cur Frame 1720 and Cur CU 1722. Cur Frame 1720 is the current frame. Cur CU 1722 is the current coding unit.

[0065] Figure 17B shows an example for deriving template samples for affine mode. This example includes Ref Frame 1740, Col CU 1742, A Ref 1743, B Ref 1744, C Ref 1745, D Ref 1746, E Ref 1747, F Ref 1748, and G Ref 1749. Ref Frame 1740 is the reference frame. Col CU 1742 is the collated coding unit. A Ref 1743, B Ref 1744, C Ref 1745, D Ref 1746, E Ref 1747, F Ref 1748, and G Ref 1749 are reference samples.

[0066] Inefficiency of predictive refinement using optical flow for affine modes While PROFs can improve the coding efficiency of affine modes, their design still has room for further improvement. In particular, given the fact that both PROFs and BDOFs are built on the concept of optical flow, it is highly desirable to harmonize the designs of PROFs and BDOFs as much as possible so that PROFs can make the most of the existing logic of BDOFs to facilitate hardware implementation. Based on such considerations, the following inefficiencies regarding the interaction between current PROF and BDOF designs are identified in this disclosure.

[0067] Firstly, as explained in the section “Predictive refinement by optical flow for affine modes,” in equation (8), the accuracy of the gradient is determined based on the internal bit depth. On the other hand, the MV difference, i.e., Δv x and Δv yIt is always derived with an accuracy of 1 / 32 Pell. Correspondingly, based on equation (9), the accuracy of the derived PROF refinement depends on the internal bit depth. However, similar to BDOF, in order to maintain higher PROF derivation accuracy, PROF is applied on the predicted sample values ​​at an intermediate high bit depth (i.e., 16 bits). Therefore, regardless of the internal coding bit depth, the accuracy of the predictive refinement derived by PROF should match the accuracy of the intermediate predicted sample, i.e., 16 bits. In other words, the representation bit depth of the MV difference and gradient in existing PROF designs does not perfectly fit to derive accurate predictive refinements compared to the predicted sample accuracy (i.e., 16 bits). On the other hand, based on a comparison of equations (1), (4), and (8), existing PROF and BDOF use different accuracies to represent the sample gradient and MV difference. As previously noted, such a non-integrated design is undesirable for hardware because the existing BDOF logic cannot be reused.

[0068] Secondly, as discussed in the section "Predictive Refinement by Optical Flow for Affine Mode," when a single current affine block is bipredicted, PROF is applied separately to the predicted samples in lists L0 and L1, and then the augmented L0 and L1 predicted signals are averaged to produce the final bipredictive signal. Conversely, instead of deriving PROF refinement separately for each prediction direction, BDOF derives predictive refinement all at once, and then the predictive refinement is applied to augment the combined L0 and L1 predicted signals. Figures 8 and 9 (described below) compare the current BDOF and PROF workflows for biprediction. In actual codec hardware pipeline designs, different primary encoding / decoding modules are typically assigned to each pipeline stage so that more encoded blocks can be processed in parallel. However, differences between the BDOF and PROF workflows can lead to the difficulty of having one same pipeline design that can be shared by BDOF and PROF, which is undesirable for actual codec implementations.

[0069] Figure 8 shows the BDOF workflow relating to this disclosure. Workflow 800 includes L0 motion compensation 810, L1 motion compensation 820, and BDOF 830. L0 motion compensation 810 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is the reference picture before the current picture in the video block. L1 motion compensation 820 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is the reference picture after the current picture in the video block. BDOF 830 takes motion compensation samples from L1 motion compensation 810 and L1 motion compensation 820 and outputs predicted samples, as described above with respect to Figure 4.

[0070] Figure 9 shows the workflow of an existing PROF relating to this disclosure. Workflow 900 includes L0 motion compensation 910, L1 motion compensation 920, L0 PROF 930, L1 PROF 940, and average 960. L0 motion compensation 910 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is the reference picture before the current picture in the video block. L1 motion compensation 920 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is the reference picture after the current picture in the video block. L0 PROF 930 takes L0 motion compensation samples from L0 motion compensation 910 and outputs a motion refinement value, as described above with respect to Figure 7. L1 PROF 940 takes L1 motion compensation samples from L1 motion compensation 920 and outputs a motion refinement value, as described above with respect to Figure 7. The average of 960 is the average of the motion refined value outputs of L0 PROF 930 and L1 PROF 940.

[0071] Thirdly, for both BDOF and PROF, the gradient must be calculated for each sample within the current encoded block, which requires generating one additional row / column of predicted samples on each side of the block. To avoid the additional computational complexity of sample interpolation, predicted samples in the extended region around the block are copied directly from the reference sample at integer positions (i.e., without interpolation). However, according to existing designs, integer samples at different locations are selected to generate the gradient values ​​for BDOF and PROF. Specifically, for BDOF, the integer reference samples to the left of the predicted sample (for horizontal gradients) and above the predicted sample (for vertical gradients) are used for gradient calculation, while for PROF, the integer reference sample closest to the predicted sample is used. Similar to the bit depth representation problem, such a non-integrated gradient calculation method is undesirable for hardware codec implementations.

[0072] Fourth, as previously noted, the motivation for PROF is to compensate for the small MV difference between the MV of each sample and the subblock MV derived at the center of the subblock to which the sample belongs. According to the current PROF design, PROF is always invoked when a single coded block is predicted by an affine mode. However, as shown in equations (6) and (7), the subblock MV of a single affine block is derived from the control point MV. Therefore, when the difference between the control point MVs is relatively small, the MV at each sample position should be consistent. In such cases, the benefits of applying PROF can be very limited, and considering the performance / complexity trade-off, it may not be worthwhile to perform PROF.

[0073] Improved predictive refinement using optical flow for affine mode This disclosure provides methods for improving and simplifying existing PROF designs to facilitate hardware codec implementation. Particular attention is paid to harmonizing the designs of BDOF and PROF in order to maximize the sharing of existing BDOF logic with PROF. In general, the main aspects of the technology proposed in this disclosure can be summarized as follows:

[0074] Firstly, in order to improve the coding efficiency of PROF while achieving a single, more integrated design, a method is proposed for integrating the representation bit depths of the sample gradient and MV difference used by BDOF and PROF.

[0075] Secondly, to facilitate hardware pipeline design, it is proposed to harmonize the PROF workflow with the BDOF workflow for dual prediction. Specifically, unlike existing PROFs that derive predictive refinements separately for L0 and L1, the proposed method derives the predictive refinements applied to the combined L0 and L1 predictive signals in a single step.

[0076] Thirdly, two methods are proposed to harmonize the derivation of integer reference samples for calculating the gradient values ​​used by BDOF and PROF.

[0077] Fourth, in order to reduce computational complexity, an early termination method is proposed to adaptively disable the PROF process for affine-coded blocks when certain conditions are met.

[0078] Improved bit-depth representation design for PROF gradients and MV differences As analyzed in the "Statement of Problem" section, the current bit depth representations of MV differences and sample gradients in PROF are not consistent enough to derive accurate predictive refinements. Furthermore, the bit depth representations of sample gradients and MV differences are inconsistent between BDOF and PROF, which is undesirable for hardware. This section proposes an improved bit depth representation method by extending the BDOF bit depth representation method to PROF. Specifically, in the proposed method, the horizontal and vertical gradients at each sample position are:

number

[0079] In addition, assuming that Δx and Δy are the horizontal and vertical offsets expressed in 1 / 4 Pell precision from one sample location to the center of the subblock to which the sample belongs, the corresponding PROF MV difference Δv(x,y) at the sample location is:

number

[0080] In the case of a 6-parameter affine model, JPEG2026048758000023.jpg42159, where, (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the top-left, top-right, and bottom-left control points MV of the current coded block, expressed with 1 / 16 Pell precision, where w and h are the width and height of the block.

[0081] In the above discussion, a fixed pair of right shifts is applied to compute the gradient and MV difference values, as shown in equations (13) and (14). In practice, different bitwise right shifts that can be applied to (13) and (14) achieve different representation accuracies of the gradient and MV difference for different trade-offs between intermediate computational accuracy and bit width in the internal PROF derivation process. For example, if the input video contains a lot of noise, the derived gradient may not be reliable in representing the true local horizontal / vertical gradient values ​​at each sample. In such cases, it makes more sense to use more bits to represent the MV difference than the gradient. On the other hand, when the input video exhibits stable motion, the MV difference derived by the affine model should be very small. In such cases, using a high-accuracy MV difference does not provide the additional benefit of improving the accuracy of the derived PROF refinement. In other words, in such cases, it is more beneficial to use more bits to represent the gradient value. Based on the above considerations, in one or more embodiments of the present disclosure, one general method for computed the gradient and MV difference for PROF is proposed below. Specifically, the horizontal and vertical gradients at each sample position are n a This is calculated by applying a right shift of 1 to the difference of adjacent predicted samples, i.e.,

number

number

number

[0082] In another embodiment, a different PROF bit depth control method is proposed as follows: In this method, the horizontal and vertical gradients at each sample position are still shifted to the right by n. a The bits are calculated similarly to (13) by applying them to the difference values ​​of adjacent predicted samples. The corresponding PROF MV difference Δv(x,y) at the sample position is: It should be calculated as JPEG2026048758000027.jpg14159.

[0083] In addition, to maintain the overall PROF derivation at an appropriate internal bit depth, clipping is applied to the derived MV difference as follows: JPEG2026048758000028.jpg14159 Here, the limit is The threshold is equal to JPEG2026048758000029.jpg10131, and clip3(min,max,x) is a function that clips a given value x within the range [min,max]. In one example, n b The value is 2 max(5,bitdepth-7) It is set to be as follows. Finally, the PROF refinement of the sample is It will be calculated as JPEG2026048758000030.jpg9156.

[0084] In addition, one or more embodiments of the present disclosure propose a single PROF bit depth control solution. In this method, horizontal and vertical PROF motion refinement at each sample position (i,j) is performed. It is derived as JPEG2026048758000031.jpg16156.

[0085] Furthermore, the derived horizontal and vertical motion refinements are It will be clipped as JPEG2026048758000032.jpg17156.

[0086] Here, considering the motion refinement derived above, the final PROF sample refinement at location (i,j) is: It will be calculated as JPEG2026048758000033.jpg9156.

[0087] In another embodiment of the present disclosure, another PROF bit depth control solution is proposed. In the second method, the horizontal and vertical PROF motion refinement at the sample position (i,j) is performed. It is derived as JPEG2026048758000034.jpg16156.

[0088] Next, the derived motion refinement was, It will be clipped as JPEG2026048758000035.jpg16156.

[0089] Therefore, considering the motion refinement derived above, the final PROF sample refinement at location (i,j) is: It will be calculated as JPEG2026048758000036.jpg9156.

[0090] In one or more embodiments of this disclosure, it is proposed to combine a motion refinement accuracy control method in a solution with a PROF sample refinement derivation method in a second solution. Specifically, this method allows for horizontal and vertical PROF motion refinement at each sample position (i,j) to be performed. It is derived as JPEG2026048758000037.jpg16156.

[0091] Furthermore, the derived horizontal and vertical motion refinements are It will be clipped as JPEG2026048758000038.jpg18156.

[0092] Here, considering the motion refinement derived above, the final PROF sample refinement at location (i,j) is: It will be calculated as JPEG2026048758000039.jpg12164.

[0093] In one or more embodiments, the following PROF sample refinement derivation method is proposed.

[0094] Firstly, As shown in JPEG2026048758000040.jpg16164, the horizontal and vertical motion refinements of the PROF are calculated to an accuracy of 1 / 32 Pell by applying a fixed right shift.

[0095] Secondly, the calculated PROF motion refined values ​​are clipped to a single symmetrical range [-31, 31]. JPEG2026048758000041.jpg19164

[0096] Thirdly, the PROF refinement of the sample, It will be calculated as JPEG2026048758000042.jpg10164.

[0097] Figure 10 shows a method for representing the bit depth of PROF. This method can be applied, for example, to a decoder.

[0098] In step 1010, the decoder may obtain a reference picture I associated with a video block in the video signal.

[0099] In step 1012, the decoder may obtain a predicted sample I(i,j) of a video block from a reference block in the reference picture I, where i and j may represent the coordinates of a single sample containing a video block.

[0100] In step 1014, the decoder may control the internal PROF parameters of the PROF derivation process by applying a right shift to the internal PROF parameters based on a bit shift value to achieve a pre-set accuracy. The internal PROF parameters include the horizontal gradient value, vertical gradient value, horizontal motion difference value, and vertical motion difference value derived for the predicted sample I(i,j).

[0101] In step 1016, the decoder may obtain a predicted refined value for the sample in the video block, based on the fact that the PROF derivation process has been applied to the video block based on the predicted sample I(i,j).

[0102] In step 1018, the decoder may obtain predicted samples for the video block based on a combination of predicted samples and predicted refined values.

[0103] In addition, the same parameter derivation method is also v x =sGx2>0 ? Clip3(-31,31,-(sGxdI<<2)>>Floor(Log2(sGx2))):0 v y =sGy2>0 ? Clip3(-31,31,((sGydI<<2)-((v x *sGxGy m )<<12+v x *sGxGy s )>>1)>>Floor(Log2(sGy2))):0 As exemplified, this can be applied to the BDOF sample refinement process, where sGxdI, sGx2, sGxGy m sGxGy s , and sGy2 are intermediate BDOF derivation parameters.

[0104] Figure 11 shows a method for controlling the bit depth of a BDOF. This method can be applied, for example, to a decoder.

[0105] In Step 1110, the decoder receives the first reference picture I associated with the video block. (0) and second reference picture I (1) It is possible to obtain the first reference picture I in the display order. (0) It is before the current picture, and is the second reference picture I (1) It is located after the current picture.

[0106] In step 1112, the decoder controls the first reference picture I (0) First predicted sample I of the video block from the reference block within (0) (i,j) can be obtained, where i and j may represent the coordinates of one sample containing the current picture.

[0107] In step 1114, the decoder uses the second reference picture I (1) Second predicted sample I of the video block from the reference block within (1) (i,j) can be obtained.

[0108] In step 1116, the decoder may control the internal BDOF parameters of the BDOF derivation process by applying a shift to the internal BDOF parameters. The internal BDOF parameters are used for the first predicted sample I (0) (i,j), Second prediction sample I (1) (i,j), First prediction sample I (0) (i,j) and the second predicted sample I (1)This includes the sample difference between (i,j) and the horizontal and vertical gradient values ​​derived based on the intermediate BDOF derivation parameters. The intermediate BDOF derivation parameters include the sGxdI, sGydI, sGx2, sGxGy, and sGy2 parameters. sGxdI and sGydI include the cross-correlation values ​​between the horizontal gradient value and the sample difference value, and between the vertical gradient value and the sample difference value. sGx2 and sGy2 include the autocorrelation values ​​of the horizontal and vertical gradient values. sGxGy includes the cross-correlation value between the horizontal and vertical gradient values.

[0109] In step 1118, the decoder receives the first prediction sample I (0) (i,j) and the second predicted sample I (1) Based on (i,j), BDOF is applied to the video block, and motion refinement can be obtained for the samples within the video block.

[0110] In step 1120, the decoder may obtain two predicted samples of the video block based on motion refinement.

[0111] Harmonized workflow of BDOF and PROF for biprediction As previously discussed, when a single affine coded block is bipredicted, the current PROF is applied unilaterally. More specifically, PROF sample refinements are derived separately and applied to the prediction samples in lists L0 and L1. The refined prediction signals from lists L0 and L1, respectively, are then averaged to produce the final biprediction signal for the block. This is in contrast to the BDOF design, where sample refinement is derived and applied to the biprediction signal. Such differences between the biprediction workflows of BDOF and PROF can be undesirable for practical codec pipeline design.

[0112] To facilitate hardware pipeline design, one simplification method relating to this disclosure is to modify the PROF's dual prediction process so that the workflows of two prediction refinement methods are harmonized. Specifically, instead of applying refinement separately for each prediction direction, the proposed PROF method derives prediction refinement at once based on the control point MV of lists L0 and L1, and the derived prediction refinement is then applied to the combined L0 and L1 prediction signals to improve quality. Specifically, based on the MV difference derived in equation (14), the final dual prediction sample of one affine coding block is obtained by the proposed method,

number

[0113] Figure 13 shows the corresponding PROF process when the proposed bipredictive PROF method is applied. PROF process 1300 includes L0 motion compensation 1310, L1 motion compensation 1320, and bipredictive PROF 1330. L0 motion compensation 1310 may be, for example, a list of motion compensation samples from a previous reference picture. The previous reference picture is the reference picture before the current picture in the video block. L1 motion compensation 1320 may be, for example, a list of motion compensation samples from a next reference picture. The next reference picture is the reference picture after the current picture in the video block. Bipredictive PROF 1330 takes motion compensation samples from L1 motion compensation 1310 and L1 motion compensation 1320 as described above and outputs bipredictive samples.

[0114] To demonstrate the potential benefits of the proposed method for hardware pipeline design, Figure 14 (described below) shows one example illustrating the pipeline stages when both BDOF and the proposed PROF are applied. In Figure 14, the decoding process for one interblock mainly involves three things.

[0115] First, the MV of the encoded block is analyzed / decoded, and the reference sample is fetched.

[0116] Secondly, the L0 and / or L1 prediction signals of the coded block are generated.

[0117] Thirdly, sample-level refinement of the generated biprediction samples is performed based on BDOF when the coded block is predicted by a single non-affine mode, or on PROF when the coded block is predicted by an affine mode.

[0118] Figure 14 illustrates an exemplary pipeline stage when both BDOF and the proposed PROF are applied as relating to this disclosure. Figure 14 demonstrates the potential benefits of the proposed method for hardware pipeline design. Pipeline stage 1400 includes 1410 for parsing / decoding the MV and fetching a reference sample, 1420 for motion compensation, and 1430 for BDOF / PROF. Pipeline stage 1400 encodes video blocks BLK0, BKL1, BKL2, BKL3, and BLK4. Each video block starts in 1410 for parsing / decoding the MV and fetching a reference sample, then moves sequentially to motion compensation 1420, and then to motion compensation 1420, BDOF / PROF 1430. This means that BLK0 does not start processing in pipeline stage 1400 until BLK0 moves to motion compensation 1420. The same applies to all stages and video blocks as time progresses from T0 to T1, T2, T3, and T4.

[0119] As shown in FIG. 14, after the proposed harmonization method is applied, both BDOF and PROF are directly applied to the dual prediction samples. Considering that BDOF and PROF are applied to different types of encoding blocks (i.e., BDOF is applied to non-affine blocks and PROF is applied to affine blocks), the two encoding tools cannot be called simultaneously. Therefore, their corresponding decoding processes can be implemented by sharing the same pipeline stage. This is more efficient than existing PROF designs where it is difficult to allocate the same pipeline stage to both BDOF and PROF due to the different workflows of dual prediction.

[0120] In the above discussion, the proposed method only considers the harmonization of the BDOF and PROF workflows. However, according to existing designs, the basic operation units of the two encoding tools are also implemented in different sizes. For example, in the case of BDOF, one encoding block is split into multiple sub-blocks of size W s ×H s , where W s = min(W, 16) and H s=min(H,16), where W and H are the width and height of the coded block. BDOF operations such as gradient calculation and sample refinement derivation are performed independently for each subblock. On the other hand, as previously described, an affine coded block is divided into 4x4 subblocks, and each subblock is assigned one individual MV derived based on either a 4-parameter affine model or a 6-parameter affine model. Since PROF is applied only to affine blocks, its basic unit of operation is a 4x4 subblock. Similar to the bipredictive workflow problem, using a different basic unit of operation size for PROF than for BDOF is undesirable for hardware implementations and makes it difficult for BDOF and PROF to share the same pipeline stage of the overall decoding process. To solve such problems, in one or more embodiments, it is proposed to adjust the subblock size of the affine mode to be the same as the subblock size of the BDOF. For example, according to the proposed method, if one coded block is coded by affine mode, then one coded block is W s ×H s It is split into subblocks having the size of W, however, s =min(W,16) and H s=min(H,16), where W and H are the width and height of the coding block. Each subblock is assigned one individual MV and is considered an independent PROF operating unit. It is worth mentioning that independent PROF operating units ensure that PROF operations on them are performed without referencing information from adjacent PROF operating units. For example, the PROF MV difference at one sample location is calculated as the difference between the MV at the sample location and the MV at the center of the PROF operating unit where the sample resides, and the gradient used by the PROF derivation is calculated by padding the samples along each PROF operating unit. The asserted benefits of the proposed method include, primarily, the following aspects: 1) a simplified pipeline architecture with an integrated basic operating unit size for both motion compensation and BDOF / PROF refinement; 2) reduced memory bandwidth usage due to an enlarged subblock size for affine motion compensation; and 3) reduced per-sample computation complexity for fractional sample interpolation.

[0121] It should also be noted that the reduced computational complexity of the proposed method (i.e., item 3)) may eliminate the existing 6-tap interpolation filter constraint for affine-coded blocks. Instead, the default 8-tap interpolation for non-affine-coded blocks is also used for affine-coded blocks. The overall computational complexity in this case is still comparable to the existing PROF design (based on 4x4 subblocks with 6-tap interpolation filters).

[0122] Harmonizing gradient derivations for BDOF and PROF As previously explained, both BDOF and PROF calculate the gradient of each sample within the current encoded block, which accesses one additional row / column of predicted samples on each side of the block. To avoid additional interpolation complexity, the required predicted samples within the extended region around the block boundary are copied directly from the integer reference samples. However, as noted in the "Statement of the Problem" section, integer samples at different locations are used to calculate the gradient values ​​for BDOF and PROF.

[0123] To achieve a single, more unified design, two methods are proposed below for integrating the gradient derivation methods used by BDOF and PROF. The first method proposes adjusting the gradient derivation method of PROF to be the same as that of BDOF. For example, by the first method, the integer positions used to generate predictive samples within the extended region are determined by flooring down the fractional sample positions, i.e., the selected integer sample positions are to the left of the fractional sample positions (for horizontal gradients) and above the fractional sample positions (for vertical gradients).

[0124] The second method proposes adjusting the gradient derivation method for BDOF to be the same as that for PROF. More specifically, when the second method is applied, the integer reference sample closest to the predicted sample is used for gradient calculation.

[0125] Figure 15 shows an example of using the BDOF gradient derivation method relating to this disclosure. In Figure 15, the blank circle 1510 represents a reference sample at an integer position, the triangle 1530 represents a fractional prediction sample for the current block, and the black circle 1520 represents an integer reference sample used to fill the extension region of the current block.

[0126] Figure 16 shows an example of using the PROF gradient derivation method relating to this disclosure. In Figure 16, the blank circle 1610 represents a reference sample at an integer position, the triangle 1630 represents a fractional prediction sample for the current block, and the black circle 1620 represents an integer reference sample used to fill the extended region of the current block.

[0127] Figures 15 and 16 illustrate the corresponding integer sample locations used to derive the gradients for BDOF and PROF when the first method (Figure 15) and the second method (Figure 16) are applied, respectively. In Figures 15 and 16, blank circles represent reference samples at integer locations, triangles represent fractional predicted samples for the current block, and patterned circles represent integer reference samples used to fill the extended region of the current block for gradient deriving.

[0128] In addition, according to existing BDOF and PROF designs, predictive sample padding is performed at different coding levels. For example, in the case of BDOF, padding is applied along the boundaries of sbWidth × sbHeight subblocks, where sbWidth = min(CUWidth, 16) and sbHeight = min(CUHeight, 16), where CUWidth and CUHeight are the width and height of one CU. On the other hand, padding for PROF is always applied at the 4 × 4 subblock level. In the above discussion, only the padding method is unified between BDOF and PROF, but the padding subblock sizes still differ. This is also undesirable for actual hardware implementations, given that different modules must be implemented for the padding process of BDOF and PROF. To achieve one more integrated design, it is proposed to unify the subblock padding sizes of BDOF and PROF. In one or more embodiments of this disclosure, it is proposed to apply predictive sample padding for BDOF at the 4 × 4 level. For example, using this method, the CU is first divided into multiple 4x4 subblocks, and after motion compensation for each 4x4 subblock, the extended samples along the top / bottom and left / right boundaries are padded by copying the corresponding integer sample positions.

[0129] Figures 19A, 19B, 19C, and 19D illustrate one example of the proposed padding method being applied to a single 16×16 BDOF CU, where the dashed lines represent 4×4 subblock boundaries and the blue bands represent the padded samples in each 4×4 subblock.

[0130] Figure 19A shows a proposed padding method applicable to a 16×16 BDOF CU according to the present disclosure, where the dashed line represents the upper left 4×4 subblock boundary 1920.

[0131] Figure 19B shows a proposed padding method applicable to a 16×16 BDOF CU according to the present disclosure, where the dashed line represents the upper right 4×4 subblock boundary 1940.

[0132] Figure 19C shows a proposed padding method applicable to a 16×16 BDOF CU according to the present disclosure, where the dashed line represents the 4×4 subblock boundary 1960 in the lower left.

[0133] Figure 19D shows the proposed padding method applicable to a 16×16 BDOF CU according to the present disclosure, where the dashed line represents the 4×4 subblock boundary 1980 in the lower right.

[0134] High-level signaling syntax for enabling / disabling BDOF, PROF, and DMVR In existing BDOF and PROF designs, two different flags in the Sequence Parameter Set (SPS) signal the enabling / disabling of the two encoding tools separately. However, due to the similarity between BDOF and PROF, it is more desirable to enable and / or disable BDOF and PROF from a high level with a single, identical control flag. Based on such considerations, a new flag called sps_bdof_prof_enabled_flag is introduced into the SPS, as shown in Table 1. As shown in Table 1, enabling and disabling BDOF depends solely on sps_bdof_prof_enabled_flag. When the flag is equal to 1, BDOF is enabled to encode video content in the sequence. Otherwise, when sps_bdof_prof_enabled_flag is equal to 0, BDOF is not applied. On the other hand, in addition to sps_bdof_prof_enabled_flag, the SPS level affine control flag, namely sps_affine_enabled_flag, is also used to conditionally enable and disable PROF. When both flags sps_bdof_prof_enabled_flag and sps_affine_enabled_flag are equal to 1, PROF is enabled for all encoded blocks encoded in affine mode. When flag sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0, PROF is disabled. [Table 1]

[0135] The sps_bdof_prof_enabled_flag specifies whether bidirectional optical flow and predictive refinement by optical flow are enabled. When sps_bdof_prof_enabled_flag is equal to 0, both bidirectional optical flow and predictive refinement by optical flow are disabled. When sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 1, both bidirectional optical flow and predictive refinement by optical flow are enabled. Otherwise (sps_bdof_prof_enabled_flag is equal to 1 and sps_affine_enabled_flag is equal to 0), bidirectional optical flow is enabled and predictive refinement by optical flow is disabled.

[0136] The sps_bdof_prof_dmvr_slice_preset_flag specifies when the flag slice_disable_bdof_prof_dmvr_flag is signaled at the slice level. When the flag is equal to 1, the syntax slice_disable_bdof_prof_dmvr_flag is signaled for each slice that references the current sequence parameter set. Otherwise (when sps_bdof_prof_dmvr_slice_present_flag is equal to 0), the syntax slice_disabled_bdof_prof_dmvr_flag is not signaled at the slice level. When this flag is not signaled, it is inferred to be 0.

[0137] Furthermore, when the proposed SPS level BDOF and PROF control flags are used, the corresponding control flag no_bdof_constraint_flag in the general constraint information syntax should also be modified according to the table below. [Table 2]

[0138] A no_bdof_prof_constraint_flag equal to 1 specifies that sps_bdof_prof_enabled_flag is equal to 0. A no_bdof_constraint_flag equal to 0 means no constraints are imposed.

[0139] In addition to the SPS BDOF / PROF syntax described above, it is proposed to introduce another control flag at the slice level, namely slice_disable_bdof_prof_dmvr_flag, to disable BDOF, PROF, and DMVR. The SPS flag sps_bdof_prof_dmvr_slice_present_flag, which is signaled in the SPS when either the sps-level control flag for DMVR or BDOF / PROF is true, is used to indicate the presence of slice_disable_bdof_prof_dmvr_flag. If present, slice_disable_bdof_dmvr_flag is signaled. Table 2 illustrates the modified slice header syntax table after the proposed syntax has been applied. In another embodiment, it is proposed to still use two control flags in the slice header to control the enabling / disabling of BDOF and DMVR, as well as the enabling / disabling of PROF, separately. For example, in this way, two flags are used in the slice header. One flag, slice_disable_bdof_dmvr_slice_flag, is used to control the on / off state of BDOF and DMVR, while the other flag, disable_prof_slice_flag, is used to control the on / off state of PROF independently. [Table 3]

[0140] In another embodiment, it is proposed to control BDOF and PROF separately using two different SPS flags. For example, two separate SPS flags, sps_bdof_enable_flag and sps_prof_enable_flag, are introduced to enable / disable the two tools separately. In addition, a high-level control flag, no_prof_constraint_flag, needs to be added to the general_constrain_info() syntax table to forcibly disable the PROF tool. [Table 4]

[0141] The sps_bdof_enabled_flag specifies whether bidirectional optical flow is enabled or disabled. When sps_bdof_enabled_flag is equal to 0, bidirectional optical flow is disabled. When sps_bdof_enabled_flag is equal to 1, bidirectional optical flow is enabled.

[0142] The sps_prof_enabled_flag flag specifies whether or not optical flow-based predictive refinement is enabled. When sps_prof_enabled_flag is equal to 0, optical flow-based predictive refinement is disabled. When sps_prof_enabled_flag is equal to 1, optical flow-based predictive refinement is enabled. [Table 5]

[0143] A no_prof_constraint_flag equal to 1 specifies that sps_prof_enabled_flag is equal to 0. A no_prof_constraint_flag equal to 0 means no constraints are imposed.

[0144] At the slice level, in one or more embodiments of the present disclosure, it is proposed to introduce an additional control flag at the slice level, namely slice_disable_bdof_prof_dmvr_flag, to disable BDOF, PROF, and DMVR together. In another embodiment, it is proposed to add two separate flags at the slice level, namely slice_disable_bdof_dmvr_flag and slice_disable_prof_flag. The first flag (i.e., slice_disable_bdof_dmvr_flag) is used to adaptively switch BDOF and DMVR on / off for a single slice, and the second flag (i.e., slice_disable_prof_flag) is used to control the enabling and disabling of the PROF tool at the slice level. In addition, when the second method is applied, the flag slice_disable_bdof_dmvr_flag should only be signaled when either the SPS BDOF or SPS DMVR flag is enabled, and the flag should only be signaled when the SPS PROF flag is enabled.

[0145] At the 16th JVET meeting, picture headers were adopted for the VVC draft. Picture headers are signaled once per slice, and syntax elements are present in the slice header as before and do not change from slice to slice.

[0146] Based on the adopted picture header, one or more embodiments of the present disclosure propose controlling BDOF, DMVR, and PROF control flags from the current slice header to the picture header. For example, in the proposed method, three different control flags sps_dmvr_picture_header_present_flag, sps_bdof_picture_header_present_flag, and sps_prof_picture_header_present_flag are signaled within the SPS. When one of the three flags is signaled as true, one additional control flag is signaled within the picture header to indicate that the corresponding tools (i.e., DMVR, BDOF, and PROF) are enabled or disabled for slices referencing the picture header. The proposed syntax elements are specified as follows: [Table 6]

[0147] The sps_dmvr_picture_header_preset_flag specifies whether the flag picture_disable_dmvr_flag is signaled in the picture header. When the flag is equal to 1, the syntax picture_disable_dmvr_flag is signaled for each picture that references the current sequence parameter set. Otherwise, the syntax picture_disable_dmvr_flag is not signaled in the picture header. When this flag is not signaled, it is inferred to be 0.

[0148] The sps_bdof_picture_header_preset_flag specifies whether the flag picture_disable_bdof_flag is signaled in the picture header. When the flag is equal to 1, the syntax picture_disable_bdof_flag is signaled for each picture that references the current sequence parameter set. Otherwise, the syntax picture_disable_bdof_flag is not signaled in the picture header. When this flag is not signaled, it is inferred to be 0.

[0149] The sps_prof_picture_header_preset_flag specifies whether the flag picture_disable_prof_flag is signaled in the picture header. When the flag is equal to 1, the syntax picture_disable_prof_flag is signaled for each picture that references the current sequence parameter set. Otherwise, the syntax picture_disable_prof_flag is not signaled in the picture header. When this flag is not signaled, it is inferred to be 0. [Table 7]

[0150] The `picture_disable_dmvr_flag` flag specifies whether the dmvr tool is enabled for slices that reference the current picture header. When the flag is equal to 1, the dmvr tool is enabled for slices that reference the current picture header. Otherwise, the dmvr tool is disabled for slices that reference the current picture header. If this flag is not present, it is inferred to be 0.

[0151] The `picture_disable_bdof_flag` flag specifies whether the bdof tool is enabled for slices that reference the current picture header. When the flag is equal to 1, the bdof tool is enabled for slices that reference the current picture header. Otherwise, the bdof tool is disabled for slices that reference the current picture header.

[0152] The `picture_disable_prof_flag` flag specifies whether the prof tool is enabled for slices that reference the current picture header. When the flag is equal to 1, the prof tool is enabled for slices that reference the current picture header. Otherwise, the prof tool is disabled for slices that reference the current picture header.

[0153] Figure 12 shows the BDOF, PROF, and DMVR methods. These methods can be applied, for example, to decoders.

[0154] In step 1210, the decoder may receive three control flags in the Sequence Parameter Set (SPS). The first control flag indicates whether BDOF is enabled to decode the video blocks in the current video sequence. The second control flag indicates whether PROF is enabled to decode the video blocks in the current video sequence. The third control flag indicates whether DMVR is enabled to decode the video blocks in the current video sequence.

[0155] In step 1212, the decoder may receive the first presence flag in the SPS when the first control flag is true, the second presence flag in the SPS when the second control flag is true, and the third presence flag in the SPS when the third control flag is true.

[0156] In step 1214, the decoder may receive a first picture control flag in the picture header of each picture when the first presence flag in the SPS indicates that BDOF is disabled for the video block in the picture.

[0157] In step 1216, the decoder may receive a second picture control flag in the picture header of each picture when the second presence flag in the SPS indicates that PROF is disabled for the video block in the picture.

[0158] In step 1218, the decoder may receive a third picture control flag in the picture header of each picture when the third presence flag in the SPS indicates that DMVR is disabled for the video block in the picture.

[0159] Early termination of PROF based on control point MV difference According to the current PROF design, PROF is always invoked when a coding block is predicted by an affine mode. However, as shown in equations (6) and (7), the subblock MV of an affine block is derived from the control point MV. Therefore, when the difference between control point MVs is relatively small, the MV at each sample position should be consistent. In such cases, the benefit of applying PROF can be very limited. Thus, to further reduce the average computational complexity of PROF, it is proposed to adaptively skip PROF-based sample refinement based on the maximum MV difference between the sample-level MV and the subblock-level MV within a single 4x4 subblock. Since the PROF MV difference values ​​for samples within a single 4x4 subblock are symmetric with respect to the subblock center, the maximum horizontal and vertical PROF MV differences are based on equation (10).

number

[0160] According to this disclosure, different metrics may be used when determining whether the MV difference is small enough to skip the PROF process.

[0161] In one example, based on equation (19), the sum of the absolute maximum horizontal MV difference and the absolute maximum vertical MV difference is less than one predefined threshold, i.e.,

number

[0162] In another example, If the maximum value of JPEG2026048758000053.jpg28153 is below the threshold, the PROF process may be skipped.

number

[0163] In addition to the two examples above, the intent of this disclosure also applies when other metrics are used to determine whether the MV difference is small enough to skip the PROF process.

[0164] In the method described above, PROF is skipped based on the magnitude of the MV difference. On the other hand, in addition to the MV difference, PROF sample refinement is also calculated based on local gradient information at each sample location within a single motion-compensated block. For prediction blocks containing details that are not very high frequency (e.g., flat areas), the gradient values ​​tend to be small, and as a result, the derived sample refinement value should be small. Taking this into account, according to another embodiment, it is proposed to apply PROF only to prediction samples in blocks that contain sufficiently high frequency information.

[0165] When determining whether a block contains sufficient high-frequency information such that it is worthwhile to call the PROF process for that block, different metrics can be used. In one example, the determination is made based on the magnitude (i.e., absolute value) of the average of the gradients of the samples within the prediction block. If the magnitude of the average is less than a threshold, the prediction block is classified as a flat area and the PROF should not be applied. Otherwise, the prediction block is considered to contain sufficient high-frequency details for which the PROF may still be applicable. In another example, the maximum magnitude of the gradients of the samples within the prediction block can be used. If the maximum magnitude is less than a threshold, the PROF should be skipped for the block. In yet another example, the difference I max -I min between the maximum sample value and the minimum sample value of the prediction block can be used to determine whether the PROF should be applied to the block. If such a difference value is less than the threshold, the PROF should be skipped for the block. It is worth noting that the spirit of the present disclosure is applicable also when some other metric is used in determining whether a given block contains sufficient high-frequency information.

[0166] Handling the interaction between PROF and LIC for the affine mode Since the adjacent reconstructed samples (i.e., templates) of the current block are used by the LIC to derive linear model parameters, decoding a single LIC-coded block depends on the complete reconstruction of its adjacent samples. Due to such interdependence, in a real hardware implementation, the LIC must be performed at a reconstruction stage where adjacent reconstructed samples become available for LIC parameter derivation. Since block reconstruction must be performed sequentially (i.e., one by one), throughput (i.e., the amount of work that can be performed in parallel per unit time) is one important issue to consider when applying other coding methods together with LIC-coded blocks. In this section, two methods are proposed for handling the interaction when both PROF and LIC are enabled for affine modes.

[0167] In a first embodiment of this disclosure, it is proposed to apply the PROF mode and the LIC mode exclusively to a single affine coding block. As previously discussed, existing designs implicitly apply PROF to all affine blocks without signaling, while a single LIC flag is signaled or inherited at the coding block level to indicate whether or not the LIC mode applies to a single affine block. According to the method in this disclosure, it is proposed to conditionally apply PROF based on the value of the LIC flag for a single affine block. When the flag is equal to 1, only LIC is applied by refining the predicted samples for the entire coding block based on the LIC weights and offsets. Otherwise (i.e., the LIC flag is equal to 0), PROF is applied to the affine coding block to refine the predicted samples for each subblock based on an optical flow model.

[0168] Figure 18A illustrates one exemplary flowchart of the decoding process based on the proposed method in which PROF and LIC are not permitted to be applied simultaneously.

[0169] Figure 18A illustrates an example of a decryption process based on the proposed method in which PROF and LIC are not permitted, as relating to this disclosure. The decryption process 1820 includes a decision step 1822, an LIC 1824, and a PROF 1826. The decision step 1822 is to determine whether the LIC flag is on, and according to that decision, the following is taken: LIC 1824 is the application of LIC if the LIC flag is set. PROF 1826 is the application of PROF if the LIC flag is not set.

[0170] In a second embodiment of the present disclosure, it is proposed to apply LIC after PROF to generate a predictive sample of one affine block. For example, after subblock-based affine motion compensation is performed, the predictive sample is refined based on PROF sample refinement, and then LIC is applied.

number

[0171] Figure 18B illustrates an example of a decoding process to which PROF and LIC are applied according to the present disclosure. The decoding process 1860 includes affine motion compensation 1862, LIC parameter derivation 1864, PROF 1866, and LIC sample adjustment 1868. Affine motion compensation 1862 applies affine motion and is an input to LIC parameter derivation 1864 and PROF 1866. LIC parameter derivation 1864 is applied to derive the LIC parameters. PROF 1866 is the PROF being applied. LIC sample adjustment 1868 is the LIC weight parameters and offset parameters combined with the PROF.

[0172] Figure 18B illustrates an exemplary decoding workflow when the second method is applied. As shown in Figure 18B, since LIC uses templates (i.e., adjacent reconstructed samples) to compute the LIC linear model, the LIC parameters can be derived as soon as adjacent reconstructed samples become available. This means that PROF refinement and LIC parameter derivation can be performed simultaneously.

[0173] LIC weights and offsets (i.e., α and β) and PROF refinement (i.e., ΔI[x]) are generally floating-point. In preferred hardware implementations, these floating-point operations are typically implemented as multiplication with a single integer value, followed by a right shift operation by several bits. Existing LIC and PROF designs have the two tools designed separately, so each is N LIC Bits and N PROF Two different bitwise right shifts are applied in two stages.

[0174] According to the third embodiment, in order to improve the coding gain when PROF and LIC are applied together to an affine coded block, it is proposed to apply LIC-based sample adjustment and PROF-based sample adjustment with high precision. This is done by combining those two right shift operations into one and applying it last to derive the final predicted sample (as shown in ((12)) of the current block).

[0175] Handling the multiplication overflow problem when combining PROF by weighted prediction and bi-prediction (BCW) by CU-level weights According to the current PROF design in the VVC working draft, PROF can be applied together with weighted prediction (WP). For example, when they are combined, the prediction signal of one affine CU is generated by the following procedure.

[0176] First, for each sample at position (x, y), calculate the L0 prediction refinement ΔI0(x, y) based on PROF and add the refinement to the original L0 prediction sample I0(x, y), that is,

Equation

[0177] Second, for each sample at position (x, y), calculate the L1 prediction refinement ΔI1(x, y) based on PROF and add the refinement to the original L1 prediction sample I1(x, y), that is,

Equation

[0178] Thirdly, by combining refined L0 and L1 prediction samples, that is,

number

[0179] As can be seen from the equation above, sample-by-sample refinement, i.e., ΔI0(x,y) and ΔI1(x,y), results in the predicted samples after PROF (i.e., I0'(x,y) and I1'(x,y)) having one more dynamic range than the original predicted samples (i.e., I0(x,y) and I1(x,y)). Considering that the refined predicted samples are multiplied by WP and BCW weight coefficients, this increases the length of the required multipliers. For example, based on the current design, when the internal coding bit depth is in the range of 8 to 12 bits, the dynamic range of the predicted signals I0(x,y) and I1(x,y) is 16 bits. However, after PROF, the dynamic range of the predicted signals I0'(x,y) and I1'(x,y) becomes 17 bits. Therefore, when PROF is applied, a 16-bit multiplication overflow problem may occur in some cases. Several methods are proposed below to correct such overflow issues.

[0180] Firstly, in the first method, it is proposed that WP and BCW be disabled when PROF is applied to one affine CU.

[0181] Secondly, in the second method, it is proposed to apply a clipping operation to the derived sample refinement before adding it to the original predicted samples such that the dynamic range of the refined predicted samples I0'(x,y) and I1'(x,y)) has the same dynamic bit depth as that of the original predicted samples I0(x,y) and I1(x,y). For example, by such a method, the sample refinements ΔI0(x,y) and ΔI1(x,y) in (23) and (24) are

number

[0182] Firstly, in the third method, it is proposed to directly clip the refined predicted sample instead of clipping the sample refinement so that the refined sample has the same dynamic range as the original predicted sample. For example, by the third method, the refined L0 and L1 samples are:

number

[0183] Secondly, in the fourth method, it is proposed to apply a specific right shift to the refined L0 and L1 prediction samples before WP and BCW, and then adjust the final prediction samples to the original accuracy by an additional left shift. For example, the final prediction samples are:

number

[0184] Thirdly, in the fifth method, it is proposed to split each multiplication of the L0 / L1 predicted samples into two multiplications by the corresponding WP / BCW weights in (25), where both multiplications are,

number

[0185] Figure 20 shows the computing environment 2010 coupled with the user interface 2060. The computing environment 2010 may be part of a data processing server. The computing environment 2010 includes a processor 2020, memory 2040, and an I / O interface 2050.

[0186] Processor 2020 typically controls the overall operation of the computing environment 2010, including operations associated with display, data acquisition, data communication, and image processing. Processor 2020 may include one or more processors for executing instructions to perform all or part of the tasks described above. Furthermore, Processor 2020 may include one or more modules that facilitate interconnection between Processor 2020 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.

[0187] Memory 2040 is configured to store various types of data to support the operation of computing environment 2010. Memory 2040 may include predetermined software 2042. Examples of such data include instructions for any application or method operating on computing environment 2010, a video data set, image data, and the like. Memory 2040 may be implemented by using any type of volatile or non-volatile memory device such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk, or a combination thereof.

[0188] I / O interface 2050 provides an interface between processor 2020 and peripheral interface modules such as a keyboard, click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. I / O interface 2050 may be coupled with an encoder and a decoder.

[0189] In some embodiments, a non-transitory computer-readable storage medium including a plurality of programs, such as in memory 2040, executable by a processor 2020 within computing environment 2010 for implementing the methods described above is also provided. For example, the non-transitory computer-readable storage medium may be a ROM, a RAM, a CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0190] A non-temporary computer-readable storage medium stores multiple programs for execution by a computing device having one or more processors, wherein the multiple programs, when executed by one or more processors, cause the computing device to perform the methods described above for motion prediction.

[0191] In some embodiments, the computing environment 2010 may be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphical processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to carry out the above methods.

[0192] The descriptions in this disclosure are provided for illustrative purposes only and are not intended to be exhaustive or limitful to this disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art who have an interest in the teachings presented in the foregoing description and the associated drawings.

[0193] Examples have been selected and described to illustrate the principles of this disclosure and to enable those skilled in the art to understand this disclosure in terms of various implementations and to best utilize the underlying principles and various implementations with various modifications suitable for a particular intended use. Therefore, it should be understood that the scope of this disclosure is not limited to the specific examples of the disclosed implementations, and that modifications and other implementations are intended to be included within the scope of this disclosure.

Claims

1. A bit depth representation method for predictive refinement (PROF) by optical flow for decoding video signals, In the decoder, the reference picture I associated with the video block in the video signal is obtained, In the decoder, a predicted sample I(i,j) of the video block is obtained from the reference block in the reference picture I, wherein i and j represent the coordinates of one sample having the video block. The decoder controls the internal PROF parameter of the PROF derivation process by applying a right shift to the internal PROF parameter based on a bit shift value to achieve a pre-set accuracy, wherein the internal PROF parameter includes the horizontal gradient value, vertical gradient value, horizontal motion difference value, and vertical motion difference value derived for the prediction sample I(i,j). In the decoder, based on the application of the PROF derivation process to the video block based on the predicted sample I(i,j), a predicted refined value is obtained for the sample in the video block. A method comprising obtaining a predicted sample of the video block based on a combination of the predicted sample and the predicted refined value in the decoder.

2. The method according to claim 1, wherein the pre-set accuracy is equal to 1 / 32 Pell.

3. By applying a right shift to the aforementioned internal PROF parameter, the internal PROF parameter of the PROF derivation process can be controlled. In the decoder, the horizontal gradient value of the first predicted sample I(i,j) is obtained based on the difference between the first predicted sample I(i+1,j) and the first predicted sample I(i-1,j), In the decoder, the vertical gradient value of the first predicted sample I(i,j) is obtained based on the difference between the first predicted sample I(i,j+1) and the first predicted sample I(i,j-1), The decoder obtains the control point motion vector (MV) of the first predicted sample I(i,j), wherein the control point MV includes the MVs of the upper left corner, upper right corner, and lower left corner blocks of one block containing the video block. In the decoder, the affine model parameters derived based on the control point MV are obtained, In the decoder, the horizontal MV difference Δv for the first predicted sample I(i,j) is determined based on the affine model parameters. x (i, j) and vertical MV difference Δv y To obtain (i, j), In the decoder, the horizontal MV difference Δv x (i, j) and the vertical MV difference Δv y The method according to claim 1, comprising right-shifting (i, j) by a bit shift value such that the bit shift value is equal to 8.

4. In the decoder, the horizontal MV difference Δv x Clipping (i, j) to the symmetrical range [-31, 31], In the decoder, the vertical MV difference Δv y Clipping (i, j) to the aforementioned symmetrical range [-31, 31], The method according to claim 3, further comprising:

5. To obtain the predicted refined value for the sample in the video block, In the decoder, the horizontal gradient value and the horizontal MV difference Δv x (i, j), the vertical gradient value, and the vertical MV difference Δv y The method according to claim 3, comprising obtaining the predicted refined value based on (i, j).

6. A method for controlling the bit depth of a bidirectional optical flow (BDOF) for decoding a video signal, In a decoder, obtaining a first reference picture I associated with a video block (0) and a second reference picture I (1) such that, in the display order, the first reference picture I (0) is before the current picture and the second reference picture I (1) is after the current picture In the decoder, the first reference picture I (0) From the reference block within, the first predicted sample I of the video block (0) Obtaining (i, j), where i and j represent the coordinates of one sample having the current picture, In the decoder, the second reference picture I (1) The second predicted sample I of the video block from the reference block within (1) To obtain (i, j), In the decoder, the internal BDOF parameter of the BDOF derivation process is controlled by applying a shift to the internal BDOF parameter, wherein the internal BDOF parameter is the first prediction sample I (0) (i, j), the second prediction sample I (1) (i, j), the first prediction sample I (0) (i, j) and the second prediction sample I (1) The intermediate BDOF derivation parameters include the sample difference between (i, j) and the horizontal and vertical gradient values ​​derived based on the intermediate BDOF derivation parameters, wherein the intermediate BDOF derivation parameters include the parameters sGxdI, sGydI, sGx2, sGxGy, and sGy2, where sGxdI and sGydI include the cross-correlation values ​​between the horizontal gradient value and the sample difference value and between the vertical gradient value and the sample difference value, sGx2 and sGy2 include the autocorrelation values ​​between the horizontal gradient value and the vertical gradient value, and sGxGy includes the cross-correlation value between the horizontal gradient value and the vertical gradient value. In the decoder, the first prediction sample I (0) (i, j) and the second prediction sample I (1) Based on the application of the BDOF to the video block based on (i, j), motion refinement is obtained for the samples in the video block, A method comprising obtaining a dual prediction sample of the video block based on the motion refinement in the decoder.

7. In the decoder, the internal bit depth of the BDOF derivation process is controlled for the accuracy of representing the internal BDOF derivation parameters. In the decoder, the first prediction sample I (0) (i, j), the second prediction sample I (1) Obtaining the intermediate BDOF derivation parameters based on (i, j), In the decoder, a horizontal motion refinement value is obtained based on the sGx2 and sGxdI parameters, In the decoder, a vertical motion refinement value is obtained based on the sGy2, sGydI, and sGxGy parameters, The method according to claim 6, further comprising clipping the horizontal motion refinement value and the vertical motion refinement value to a symmetric range of [-31, 31] in the decoder.

8. A method for bidirectional optical flow (BDOF), predictive refinement by optical flow (PROF), and decoder-side motion vector refinement (DMVR), The decoder receives three control flags in the sequence parameter set (SPS), wherein the first control flag indicates whether the BDOF is enabled to decode the video block in the current video sequence, the second control flag indicates whether the PROF is enabled to decode the video block in the current video sequence, and the third control flag indicates whether the DMVR is enabled to decode the video block in the current video sequence. The decoder receives the first presence flag in the SPS when the first control flag is true, the second presence flag in the SPS when the second control flag is true, and the third presence flag in the SPS when the third control flag is true. When the decoder receives the first picture control flag in the picture header of each picture when the first presence flag in the SPS indicates that the BDOF is disabled for the video block in the picture, When the decoder receives the second picture control flag in the picture header of each picture when the second presence flag in the SPS indicates that the PROF is disabled for the video block in the picture, A method comprising: receiving a third picture control flag in the picture header of each picture when the decoder indicates that the third presence flag in the SPS is disabled for the video block in the picture.

9. The decoder receives the three control flags in the SPS. The decoder receives the sps_bdof_enabled_flag flag, which indicates whether the BDOF is permitted to decode the video block in the sequence. The decoder receives the sps_prof_enabled_flag flag, which indicates whether the sps_prof_enabled_flag flag is permitted to decode the video block in the sequence. The method according to claim 8, comprising: the decoder receiving the sps_dmvr_enabled_flag flag, the sps_dmvr_enabled_flag flag indicating whether the DMVR is permitted to decode the video block in the sequence.

10. The decoder receives the sps_dmvr_picture_header_present_flag flag when the sps_dmvr_enabled_flag flag is true, and the sps_dmvr_picture_header_present_flag flag signals whether the picture_disable_dmvr_flag flag is signaled in the picture header of each picture that references the current SPS, The decoder receives the sps_bdof_picture_header_present_flag flag when the sps_bdof_enabled_flag flag is true, the sps_bdof_picture_header_present_flag flag signals whether the picture_disable_bdof_flag flag is signaled in the picture header of each picture that references the current SPS, The decoder receives the sps_prof_picture_header_present_flag flag when the sps_prof_enabled_flag flag is true, the sps_prof_picture_header_present_flag flag signals whether the picture_disable_prof_flag flag is signaled in the picture header of each picture that references the current SPS, The method according to claim 8, further comprising:

11. When the value of the sps_bdof_picture_header_present_flag flag is false, the decoder applies the BDOF to generate the predicted samples of the intercoded blocks that are not coded in affine mode, When the value of the sps_prof_picture_header_present_flag flag is false, the decoder applies the PROF to generate the predicted samples of the intercoded block encoded in affine mode, When the value of the sps_dmvr_picture_header_present_flag flag is false, the decoder applies the DMVR to generate the predicted samples of the intercoded block that is not encoded in affine mode, The method according to claim 10, further comprising:

12. The decoder further includes receiving a picture header control flag when the sps_dmvr_picture_header_present_flag flag is true, wherein the picture header control flag is the picture_disable_dmvr_flag flag, which signals that the DMVR is disabled for slices referencing the picture header. The method according to claim 10.

13. The decoder further includes receiving a picture header control flag when the sps_bdof_picture_header_present_flag flag is true, wherein the picture header control flag is the picture_disable_bdof_flag flag, which signals that the BDOF is disabled for slices referencing the picture header. The method according to claim 10.

14. The decoder further includes receiving a picture header control flag when the sps_prof_picture_header_present_flag flag is true, wherein the picture header control flag is the picture_disable_prof_flag flag, which signals that the PROF is disabled for slices that reference the picture header. The method according to claim 10.

15. A computing device, One or more processors, A non-temporary computer-readable storage medium that stores instructions executable by the one or more processors, To obtain a reference picture I associated with a video block in the aforementioned video signal, Obtaining a predicted sample I(i,j) of the video block from a reference block in the reference picture I, wherein i and j represent the coordinates of one sample having the video block, Controlling the internal predictive refinement (PROF) parameters by optical flow in the PROF derivation process by applying a right shift to the internal PROF parameters based on a bit shift value to achieve a pre-set accuracy, wherein the internal PROF parameters include the horizontal gradient value, vertical gradient value, horizontal motion difference value, and vertical motion difference value derived for the predictive sample I(i,j). Based on the application of the PROF derivation process to the video block based on the predicted sample I(i,j), a predicted refined value is obtained for the sample in the video block. A computing device configured to obtain a predicted sample of the video block based on a combination of the predicted sample and the predicted refined value.

16. The computing device according to claim 15, wherein the pre-set accuracy is equal to 1 / 32 Pel.

17. One or more processors configured to control the internal PROF parameter of the PROF derivation process by applying a right shift to the internal PROF parameter, Based on the difference between the first predicted sample I(i+1,j) and the first predicted sample I(i-1,j), the horizontal gradient value of the first predicted sample I(i,j) is obtained, Based on the difference between the first predicted sample I(i,j+1) and the first predicted sample I(i,j-1), the vertical gradient value of the first predicted sample I(i,j) is obtained, Obtaining the control point motion vector (MV) of the first prediction sample I(i,j), wherein the control point MV includes the MVs of the upper left corner, upper right corner, and lower left corner block of one block containing the video block, Obtaining affine model parameters derived based on the aforementioned control point MV, Based on the affine model parameters, the horizontal MV difference Δv for the first predicted sample I(i,j) x (i, j) and vertical MV difference Δv y To obtain (i, j), The horizontal MV difference Δv x (i, j) and the vertical MV difference Δv y The computing device according to claim 15, further configured to perform the following: shift (i, j) to the right by the bit shift value, wherein the bit shift value is equal to 8.

18. The one or more processors The horizontal MV difference Δv x Clipping (i, j) to the symmetrical range [-31, 31], The vertical MV difference Δv y The computing device according to claim 18, further configured to clip (i, j) to the symmetrical range [-31, 31].

19. One or more processors configured to obtain the predicted refined values ​​for the samples in the video block, The aforementioned horizontal gradient value, the aforementioned horizontal MV difference Δv x (i, j), the vertical gradient value, and the vertical MV difference Δv y The computing device according to claim 18, further configured to obtain the predicted refined value based on (i, j).

20. A computing device, One or more processors, A non-temporary computer-readable storage medium that stores instructions executable by the one or more processors, First reference picture I associated with the video block (0) and second reference picture I (1) This involves obtaining the first reference picture I in the display order. (0) It is in front of the current picture, and the second reference picture I (1) The fact that it is after the current picture, The first reference picture I (0) From the reference block within, the first predicted sample I of the video block (0) Obtaining (i, j), where i and j represent the coordinates of one sample having the current picture, The second reference picture I (1) The second predicted sample I of the video block from the reference block within (1) To obtain (i, j), The internal BDOF parameters of the BDOF derivation process are controlled by applying a shift to the internal BDOF parameters, wherein the internal BDOF parameters are the first predicted sample I (0) (i, j), the second prediction sample I (1) (i, j), the first prediction sample I (0) (i, j) and the second prediction sample I (1) The intermediate BDOF derivation parameters include the sample difference between (i, j) and the horizontal and vertical gradient values ​​derived based on the intermediate BDOF derivation parameters, wherein the intermediate BDOF derivation parameters include the parameters sGxdI, sGydI, sGx2, sGxGy, and sGy2, where sGxdI and sGydI include the cross-correlation values ​​between the horizontal gradient value and the sample difference value and between the vertical gradient value and the sample difference value, sGx2 and sGy2 include the autocorrelation values ​​between the horizontal gradient value and the vertical gradient value, and sGxGy includes the cross-correlation value between the horizontal gradient value and the vertical gradient value. The first prediction sample I (0) (i, j) and the second prediction sample I (1) Based on the application of the BDOF to the video block based on (i, j), motion refinement is obtained for the samples in the video block, A computing device configured to acquire a dual prediction sample of the video block based on the motion refinement described above.

21. One or more processors configured to control the internal bit depth of the BDOF derivation process for the accuracy of representing the internal BDOF derivation parameters, The first prediction sample I (0) (i, j), the second prediction sample I (1) Obtaining the intermediate BDOF derivation parameters based on (i, j), The horizontal motion refinement value is obtained based on the sGx2 and sGxdI parameters, The refined vertical motion values ​​are obtained based on the aforementioned sGy2, sGydI, and sGxGy parameters, The computing device according to claim 20, further configured to clip the horizontal motion refinement value and the vertical motion refinement value to a symmetric range of [-31, 31].

22. A non-temporary computer-readable storage medium for storing multiple programs to be executed by a computing device having one or more processors, wherein when the multiple programs are executed by the one or more processors, the computing device... The decoder receives three control flags in the Sequence Parameter Set (SPS), wherein the first control flag indicates whether bidirectional optical flow (BDOF) is enabled to decode the video block in the current video sequence, the second control flag indicates whether predictive refinement by optical flow (PROF) is enabled to decode the video block in the current video sequence, and the third control flag indicates whether decoder-side motion vector refinement (DMVR) is enabled to decode the video block in the current video sequence. The decoder receives the first presence flag in the SPS when the first control flag is true, the second presence flag in the SPS when the second control flag is true, and the third presence flag in the SPS when the third control flag is true. When the decoder receives the first picture control flag in the picture header of each picture when the first presence flag in the SPS indicates that the BDOF is disabled for the video block in the picture, When the decoder receives the second picture control flag in the picture header of each picture when the second presence flag in the SPS indicates that the PROF is disabled for the video block in the picture, A non-temporary computer-readable storage medium that causes the decoder to perform an action including receiving a third picture control flag in the picture header of each picture when the third presence flag in the SPS indicates that the DMVR is disabled for the video block in the picture.

23. The aforementioned multiple programs are installed on the computing device. The decoder receives the sps_dmvr_picture_header_present_flag flag when the sps_dmvr_enabled_flag flag is true, the sps_dmvr_picture_header_present_flag flag signals whether the picture_disable_dmvr_flag flag is signaled in the picture header of each picture that references the current SPS, The decoder receives the sps_bdoc_enable_flag flag when the sps_bdoc_enable_flag flag is true, the sps_bdoc_picture_header_present_flag flag signals whether the picture_disable_bdoc_flag flag is signaled in the picture header of each picture that references the current SPS, The non-temporary computer-readable storage medium according to claim 22, wherein the decoder is made to receive the sps_prof_picture_header_present_flag flag when the sps_prof_enable_flag flag is true, and the sps_prof_picture_header_present_flag flag signals whether the picture_disable_prof_flag flag is signaled in the picture header of each picture that references the current SPS.

24. The aforementioned multiple programs are installed on the computing device. When the value of the sps_bdof_picture_header_present_flag flag is false, the BDOF is applied in the decoder to generate the predicted samples of the intercoded blocks that are not encoded in affine mode, When the value of the sps_prof_picture_header_present_flag flag is false, the PROF is applied in the decoder to generate the predicted samples of the intercoded block encoded in affine mode, The non-temporary computer-readable storage medium according to claim 22, further comprising: applying the DMVR to the decoder to generate the predicted samples of the intercoded blocks that are not encoded in affine mode when the value of the sps_dmvr_picture_header_present_flag flag is false.

25. The aforementioned multiple programs are installed on the computing device. The non-temporary computer-readable storage medium according to claim 24, wherein the decoder further causes the decoder to receive a picture header control flag when the sps_dmvr_picture_header_present_flag flag is true, and the picture header control flag is the picture_disable_dmvr_flag flag which signals that the DMVR is disabled for slices that reference the picture header.

26. The aforementioned multiple programs are installed on the computing device. The non-temporary computer-readable storage medium according to claim 24, wherein the decoder further causes the decoder to receive a picture header control flag when the sps_bdof_picture_header_present_flag flag is true, and the picture header control flag is the picture_disable_bdof_flag flag which signals that the BDOF is disabled for slices that reference the picture header.

27. The aforementioned multiple programs are installed on the computing device. The non-temporary computer-readable storage medium according to claim 24, wherein the decoder further causes the decoder to receive a picture header control flag when the sps_prof_picture_header_present_flag flag is true, and the picture header control flag is the picture_disable_prof_flag flag which signals that the PROF is disabled for slices that reference the picture header.

28. The aforementioned multiple programs are installed on the computing device. The non-temporary computer-readable storage medium according to claim 24, wherein the decoder further causes the decoder to receive a picture header control flag when the sps_prof_picture_header_present_flag flag is true, and the picture header control flag is the picture_disable_prof_flag flag which signals that the PROF is disabled for slices that reference the picture header.