Calculating optical flow-based prediction refinement

Optical flow-based prediction refinements in video processing address bandwidth challenges by enhancing encoding and decoding efficiency, thereby reducing data requirements and improving video quality.

JP7741154B2Active Publication Date: 2025-09-17DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023207339
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-03-27
Filing Date
2023-12-08
Publication Date
2025-09-17
Estimated Expiration
2040-03-17

AI Technical Summary

Technical Problem

Digital video consumption continues to demand significant bandwidth due to inefficiencies in existing video compression technologies, necessitating improved methods for encoding and decoding to reduce data requirements.

Method used

Implementing optical flow-based prediction refinements in video processing, including techniques such as gradient components, motion displacements, and interweave prediction to enhance video block conversions and bitstream representations.

Benefits of technology

Enhances video encoding and decoding efficiency, reducing bandwidth demands and improving the quality of decompressed video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007741154000030
    Figure 0007741154000030
  • Figure 0007741154000031
    Figure 0007741154000031
  • Figure 0007741154000032
    Figure 0007741154000032
Patent Text Reader

Abstract

To provide a video processing method for optical flow-based prediction refinement.SOLUTION: A video processing method includes determining a first motion displacement Vx(x,y) at a position (x,y) and a second motion displacement Vx(x,y) at a position (x,y) in a video block encoded using an optical flow-based method, where x and y are fractions, determining Vx(x,y) and Vy(x,y) on the basis of at least the position (x,y) and a center position of the basic video block of the video block, and performing conversion between the video block and a bitstream representation of a current video block using the first motion displacement and the second motion displacement.SELECTED DRAWING: Figure 28D
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This is a divisional application of Patent Application No. 2021-555421, which is a national phase application of International Patent Application No. PCT / CN2020 / 079675 filed on March 17, 2020, which claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 078411 filed on March 17, 2019, International Patent Application No. PCT / CN2019 / 078501 filed on March 18, 2019, International Patent Application No. PCT / CN2019 / 078719 filed on March 19, 2019, and International Patent Application No. PCT / CN2019 / 079961 filed on March 27, 2019. The disclosures of the above applications are incorporated herein in their entireties.

[0002] This patent document relates to video encoding and decoding. [Background technology]

[0003] Despite advances in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video usage are expected to continue to increase. Summary of the Invention

[0004] Various techniques are provided that can be implemented by digital video encoders, transcoders, and decoders to use optical flow-based prediction refinements in processing video.

[0005] A first example of a video processing method includes: determining a refined prediction sample P'(x,y) at a position (x,y) within a video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers; and performing a conversion between the video block and a bitstream representation of the video block using reconstructed sample values ​​Rec(x,y) at the position (x,y) obtained based on the refined prediction sample P'(x,y) and residual sample values ​​Res(x,y).

[0006] A second example of a video processing method includes: determining a refined prediction sample P'(x,y) at a position (x,y) within a video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers; and encoding a bitstream representation of the video block to include residual sample values ​​Res(x,y) based on reconstructed sample values ​​Rec(x,y) at the position (x,y) that are based on at least the refined prediction sample P'(x,y).

[0007] A third example of a video processing method includes determining a first motion displacement Vx(x,y) at a position (x,y) and a second motion displacement Vy(x,y) at the position (x,y) within a video block to be encoded using an optical flow-based method, where x and y are fractional numbers, and Vx(x,y) and Vy(x,y) are determined based on at least the position (x,y) and a center position of a basic video block of the video block; and performing a conversion between the video block and a bitstream representation of a current video block using the first motion displacement and the second motion displacement.

[0008] A fourth example of a video processing method includes determining a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within a video block and a second gradient component Gy(x,y) in a second direction estimated at the position (x,y) within the video block, where the first gradient component and the second gradient component are based on a final predicted sample value of a predicted sample P(x,y) at the position (x,y), and x and y are integers; and performing a conversion between the video block and a bitstream representation of a current video block using a reconstructed sample value Rec(x,y) at the position (x,y) obtained based on a residual sample value Res(x,y) plus the final predicted sample value of a predicted sample P(x,y), which has been refined using the gradients Gx(x,y), Gy(x,y).

[0009] A fifth example of a video processing method includes determining a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within a video block and a second gradient component Gy(x,y) in a second direction estimated at the position (x,y) within the video block, where the first gradient component and the second gradient component are based on a final predicted sample value of a predicted sample P(x,y) at the position (x,y), and where x and y are integers; and encoding a bitstream representation of the video block to include residual sample values ​​Res(x,y) that are based on reconstructed sample values ​​Rec(x,y) at the position (x,y), where the reconstructed sample values ​​Rec(x,y) are based on the residual sample values ​​Res(x,y) plus the final predicted sample value of a predicted sample P(x,y) refined using the gradients Gx(x,y), Gy(x,y).

[0010] A sixth example of a video processing method includes determining a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within a video block and a second gradient component Gy(x,y) in a second direction estimated at the position (x,y) within the video block, where the first gradient component and the second gradient component are based on intermediate predicted sample values ​​of predicted samples P(x,y) at the position (x,y), where final predicted sample values ​​of the predicted samples P(x,y) are based on the intermediate predicted sample values, and where x and y are integers; and performing a conversion between the video block and a bitstream representation of a current video block using reconstructed sample values ​​Rec(x,y) at the position (x,y) obtained based on the final predicted sample values ​​of predicted samples P(x,y) and residual sample values ​​Res(x,y).

[0011] A seventh example of a video processing method includes determining a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within a video block and a second gradient component Gy(x,y) in a second direction estimated at the position (x,y) within the video block, where the first gradient component and the second gradient component are based on intermediate predicted sample values ​​of predicted samples P(x,y) at the position (x,y), where final predicted sample values ​​of predicted samples P(x,y) are based on the intermediate predicted sample values, and where x and y are integers; and encoding a bitstream representation of the video block to include residual sample values ​​Res(x,y) based on reconstructed sample values ​​Rec(x,y) at the position (x,y), where the reconstructed sample values ​​Rec(x,y) are based on the final predicted sample values ​​of predicted samples P(x,y) and the residual sample values ​​Res(x,y).

[0012] An eighth example of a video processing method includes the steps of: determining a refined prediction sample P'(x,y) at a position (x,y) in an affine-coded video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where the first direction is orthogonal to the second direction, and x and y are integers; and determining a reconstructed sample value Rec(x,y) at the location (x,y) based on a refined prediction sample P'(x,y) and a residual sample value Res(x,y); determining a refined reconstructed sample value Rec'(x,y) at the location (x,y) within the affine-coded video block, where Rec'(x,y) = Rec(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y); and using the refined reconstructed sample value Rec'(x,y) to convert between the affine-coded video block and a bitstream representation of the affine-coded video block.

[0013] A ninth example of a video processing method includes determining a refined prediction sample P'(x,y) at a position (x,y) in an affine-coded video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), wherein the first direction is orthogonal to the second direction, and x and y are integers. determining a reconstructed sample value Rec(x,y) at the location (x,y) based on the refined prediction sample P'(x,y) and residual sample values ​​Res(x,y); determining a refined reconstructed sample value Rec'(x,y) at the location (x,y) within the affine-coded video block, where Rec'(x,y)=Rec(x,y)+Gx(x,y)×Vx(x,y)+Gy(x,y)×Vy(x,y); and encoding a bitstream representation of the affine-coded video block to include the residual sample values ​​Res(x,y).

[0014] A tenth example of a video processing method includes the steps of: determining a motion vector with 1 / N pixel accuracy for an affine mode video block; determining an estimated motion displacement vector (Vx(x,y), Vy(x,y)) for a position (x,y) within the video block, the motion displacement vector being derived with 1 / M pixel accuracy, N and M being positive integers, and x and y being integers; and using the motion vector and the motion displacement vector to perform a conversion between the video block and a bitstream representation of the video block.

[0015] An eleventh example of a video processing method includes a step of determining two sets of motion vectors for a video block or a sub-block of the video block, each set of the two sets of motion vectors having a different motion vector pixel precision, and the two sets of motion vectors being determined using a temporal motion vector prediction (TMVP) technique or a sub-block-based temporal motion vector prediction (SbTMVP) technique, and a step of performing a conversion between the video block and a bitstream representation of the video block based on the two sets of motion vectors.

[0016] A twelfth example of a video processing method includes the steps of: performing an interweave prediction technique on a video block to be coded using an affine coding mode by dividing the video block into a plurality of partitions using K different sub-block patterns, where K is an integer greater than 1; generating a predicted sample for the video block by performing motion compensation using a first of the K different sub-block patterns, where a predicted sample at a location (x, y) is denoted as P(x, y), where x and y are integers; and generating an Lth pattern of the K different sub-block patterns, where K is an integer greater than 1. For at least one of the remaining sub-block patterns, determining an offset value OL(x,y) at the location (x,y) based on a predicted sample derived with a first sub-block pattern and a difference between a motion vector derived using the first of the K sub-block patterns and a motion vector derived using the L patterns; determining a final predicted sample for the location (x,y) as a function of OL(x,y) and P(x,y); and using the final predicted sample to perform a conversion between a bitstream representation of the video block and the video block.

[0017] A thirteenth example of a video processing method includes a step of performing a conversion between a bitstream representation of a video block and the video block using a final prediction sample, which is derived from a refined intermediate prediction sample by (a) performing an interweave prediction technique and a subsequent optical flow-based prediction refinement technique based on a rule, or (b) performing a motion compensation technique.

[0018] A fourteenth example of a video processing method includes, when bi-prediction is applied, performing a conversion between a bitstream representation of a video block and the video block using a final prediction sample, the final prediction sample being derived from a refined intermediate prediction sample by (a) disabling an interweave prediction technique and performing an optical flow-based prediction refinement technique, or (b) performing a motion compensation technique.

[0019] A fifteenth example of a video processing method includes a step of performing a conversion between a bitstream representation of a video block and the video block using prediction samples, the prediction samples being derived from refined intermediate prediction samples by performing an optical flow-based prediction refinement technique, wherein the performing of the optical flow-based prediction refinement technique depends on only one of a first set of motion displacements Vx(x,y) estimated in a first direction for the video block or a second set of motion displacements Vy(x,y) estimated in a second direction for the video block, where x and y are integers and the first direction is orthogonal to the second direction.

[0020] A sixteenth example of a video processing method includes a step of obtaining a refined motion vector for a video block by refining the motion vector of the video block, where the motion vector is refined before performing a motion compensation technique, the refined motion vector having 1 / N pixel accuracy and the motion vector having 1 / M pixel accuracy; a step of obtaining a final prediction sample by performing an optical flow-based prediction refinement technique on the video block, where the optical flow-based prediction refinement technique is applied to the difference between the refined motion vector and the motion vector; and a step of performing a conversion between a bitstream representation of the video block and the video block using the final prediction sample.

[0021] A seventeenth example of a video processing method includes a step of determining a final motion vector for a video block using a multi-step decoder-side motion vector refinement process, the final motion vector having 1 / N pixel accuracy, and a step of performing a conversion between the current block and a bitstream representation using the final motion vector.

[0022] An eighteenth example of a video processing method includes the steps of obtaining refined intermediate prediction samples of a video block by performing an interweave prediction technique and an optical flow-based prediction refinement technique on the intermediate prediction samples of the video block, deriving final prediction samples from the refined intermediate prediction samples, and using the final prediction samples to perform a conversion between a bitstream representation of the video block and the video block.

[0023] A 19th example of a video processing method includes the steps of obtaining refined intermediate prediction samples of a video block by performing an interweave prediction technique and a phase variational affine sub-block motion compensation (PAMC) technique on the intermediate prediction samples of the video block, deriving final prediction samples from the refined intermediate prediction samples, and using the final prediction samples to perform a conversion between a bitstream representation of the video block and the video block.

[0024] A twentieth example of a video processing method includes the steps of obtaining refined intermediate prediction samples of a video block by performing an optical flow-based prediction refinement technique and a phase variational affine sub-block motion compensation (PAMC) technique on the intermediate prediction samples of the video block, deriving final prediction samples from the refined intermediate prediction samples, and using the final prediction samples to perform a conversion between a bitstream representation of the video block and the video block.

[0025] A 21st example of a video processing method includes, in converting between a video block and a bitstream representation of the video block, determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) as a function of a gradient in a first direction and / or a second direction estimated at the position (x,y) and a first motion displacement and / or a second motion displacement estimated for the position (x,y), and performing the conversion using a reconstructed sample value Rec(x,y) from the refined prediction sample P'(x,y).

[0026] A 22nd example of a video processing method includes a step of determining a first displacement vector Vx(x,y) and a second displacement vector Vy(x,y) at a position (x,y) within the video block corresponding to an optical flow-based method of encoding the video block based on information from adjacent blocks or basic blocks, and a step of performing a conversion between the video block and a bitstream representation of the current video block using the first displacement vector and the second displacement vector.

[0027] A 23rd example of a video processing method includes, in converting between a video block and a bitstream representation of the video block, a step of determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying a prediction sample P(x,y) at the position (x,y), wherein a gradient in a first direction and a gradient in a second direction at the position (x,y) are determined based on the refined prediction sample P'(x,y) and a final prediction value determined from a residual sample value at the position (x,y), and a step of performing the conversion using the gradient in the first direction and the gradient in the second direction.

[0028] A 24th example of a video processing method includes the steps of determining a reconstructed sample Rec(x,y) at a position (x,y) within a video block to be affine-coded, refining Rec(x,y) using first and second displacement vectors and first and second gradients at the position (x,y) to obtain a refined reconstructed sample Rec'(x,y), and using the refined reconstructed sample to perform a conversion between the video block and a bitstream representation of a current video block.

[0029] A twenty-fifth example of a video processing method includes, in converting between a video block coded using an affine coding mode and a bitstream representation of the video block, performing interweave prediction of the video block by dividing the video block into a plurality of partitions using K sub-block patterns, where K is an integer greater than 1; performing motion compensation using a first of the K sub-block patterns to generate a predicted sample for the video block, wherein the predicted sample at a location (x, y) is denoted as P(x, y); for at least one of the remaining K sub-block patterns, denoted an L-th pattern, determining an offset value OL(x, y) at the location (x, y) based on P(x, y) and a difference between a motion vector derived using the first of the K sub-block patterns and a motion vector derived using the L-th pattern; determining a final predicted sample for the location (x, y) as a function of OL(x, y) and P(x, y); and performing the conversion using the final predicted sample.

[0030] In yet another exemplary embodiment, a video encoder device configured to implement one of the methods described in this patent document is disclosed.

[0031] In yet another exemplary embodiment, a video decoder device configured to implement one of the methods described in this patent document is disclosed.

[0032] In yet another aspect, a computer readable medium is disclosed having processor executable code stored thereon for implementing one of the methods described in this patent document, thus being a non-transitory computer readable medium having code for implementing the method described above and any of the methods described in this patent document.

[0033] These and other aspects are described in detail herein. [Brief explanation of the drawings]

[0034] [Figure 1] 1 illustrates an example of a derivation process for building a merge candidate list. [Figure 2] 10 shows examples of spatial merge candidate locations. [Figure 3] 10 shows examples of candidate pairs that are considered for redundancy checking of spatial merge candidates. [Figure 4] 10 shows examples of the location of the second PU for Nx2N and 2NxN partitions. [Figure 5] 10 shows an illustration of motion vector scaling for temporal merge candidates. [Figure 6] 10 shows example candidate positions for temporal merge candidates C0 and C1. [Figure 7] 1 illustrates an example of combined bi-predictive merge candidates. [Figure 8] This summarizes the derivation process for motion vector prediction candidates. [Figure 9] 1 shows an illustrative diagram of motion vector scaling for spatial motion vector candidates; [Figure 10] 10 shows an example of advanced temporal motion vector predictor ATMVP motion prediction for a coding unit CU. [Figure 11] An example of one CU having four sub-blocks (AD) and their adjacent blocks (ad) is shown. [Figure 12] 1 is an example of an explanatory diagram of a sub-block to which OBMC is applied. [Figure 13] 1 shows an example of adjacent samples used to derive IC parameters. [Figure 14] 1 shows a simplified affine motion model. [Figure 15] 1 shows an example of affine MVF for each sub-block. [Figure 16]An example of an MVP for AF_INTER is shown below. [Figure 17A] Figures 17A-17B show the candidates for AF_MERGE. [Figure 17B] Figures 17A-17B show the candidates for AF_MERGE. [Figure 18] 1 illustrates an example of bilateral matching. [Figure 19] 1 shows an example of template matching. [Figure 20] 1 shows an example of unilateral motion estimation ME in frame rate up-conversion FRUC. [Figure 21] 1 shows an example of an optical flow trajectory. [Figure 22A] 22A-22B show an example of an access location outside a block and how padding is used to avoid extra memory accesses and computations. [Figure 22B] 22A-22B show an example of an access location outside a block and how padding is used to avoid extra memory accesses and computations. [Figure 23] An example of an interpolated sample used in BIO is shown. [Figure 24] 1 shows an example of DMVR based on bilateral template matching. [Figure 25] An example of a sub-block MV VSB and a pixel Δv(i,j) (shown as an arrow) is shown. [Figure 26] 1 shows an example of how to derive Vx(x,y) and / or Vy(x,y). [Figure 27] 1 illustrates an example of a video processing device. [Figure 28A] 28A-28U are example flowcharts of methods for video processing. [Figure 28B] 28A-28U are example flowcharts of methods for video processing. [Figure 28C]28A-28U are example flowcharts of methods for video processing. [Figure 28D] 28A-28U are example flowcharts of methods for video processing. [Figure 28E] 28A-28U are example flowcharts of methods for video processing. [Figure 28F] 28A-28U are example flowcharts of methods for video processing. [Figure 28G] 28A-28U are example flowcharts of methods for video processing. [Figure 28H] 28A-28U are example flowcharts of methods for video processing. [Figure 28I] 28A-28U are example flowcharts of methods for video processing. [Figure 28J] 28A-28U are example flowcharts of methods for video processing. [Figure 28K] 28A-28U are example flowcharts of methods for video processing. [Figure 28L] 28A-28U are example flowcharts of methods for video processing. [Figure 28M] 28A-28U are example flowcharts of methods for video processing. [Figure 28N] 28A-28U are example flowcharts of methods for video processing. [Figure 28O] 28A-28U are example flowcharts of methods for video processing. [Figure 28P] 28A-28U are example flowcharts of methods for video processing. [Figure 28Q] 28A-28U are example flowcharts of methods for video processing. [Figure 28R] 28A-28U are example flowcharts of methods for video processing. [Figure 28S] 28A-28U are example flowcharts of methods for video processing. [Figure 28T]28A-28U are example flowcharts of methods for video processing. [Figure 28U] 28A-28U are example flowcharts of methods for video processing. [Figure 29] 1 shows an example of a split pattern in interweave prediction. [Figure 30] 1 shows an example of phase variation horizontal filtering. [Figure 31] An example of applying a single 8-tap horizontal filtering is shown. [Figure 32] 1 shows an example of non-uniform phase vertical filtering. [Figure 33] FIG. 1 is a block diagram illustrating an example of a video processing system in which various techniques disclosed herein may be implemented. [Figure 34] FIG. 1 is a block diagram illustrating a video encoding system in accordance with some embodiments of the present disclosure. [Figure 35] FIG. 2 is a block diagram illustrating an encoder according to some embodiments of the present disclosure. [Figure 36] FIG. 2 is a block diagram illustrating a decoder according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0035] This document provides various techniques that can be used by decoders of image or video bitstreams to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" is used herein to include both a series of pictures (traditionally called a video) and individual images. Video encoders may also implement these techniques in the encoding process to reconstruct decoded frames that are used for further encoding.

[0036] Section headings are used in this document for ease of understanding, but they do not limit the embodiments and techniques to the corresponding section, and therefore, embodiments from one section can be combined with embodiments from other sections.

[0037] 1. Overview The technology described in this patent document relates to video coding technology. Specifically, the described technology relates to motion compensation in video coding. It may be applied to existing video coding standards such as HEVC, or to emerging standards (Versatile Video Coding). It may also be applicable to future video coding standards or video codecs.

[0038] 2. Background Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on hybrid video coding architectures that utilize transform coding in addition to temporal prediction. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, numerous new methods have been adopted by the JVET and incorporated into reference software named the Joint Exploration Model (JEM). In April 2018, the Joint Video Expert Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, which aims to reduce the bitrate by 50% compared to HEVC.

[0039] The latest version of the VVC draft, Versatile Video Coding (Draft 2), is: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 11_Ljubljana / wg11 / JVET-K1001-v7.zip can be found at.

[0040] The latest reference software for VVC, called VTM, is: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-2.1 can be found at.

[0041] 2.1 Inter Prediction in HEVC / H.265 Each PU with inter prediction has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly coded as a delta to the predictor.

[0042] When a CU is coded in skip mode, one PU is associated with the CU, there are no significant residual coefficients, and there are no coded motion vector deltas or reference picture indices. A merge mode is defined, in which motion parameters for the current PU are obtained from neighboring PUs, including spatial and temporal candidates. The merge mode can be applied to any inter-predicted PU, not just to skip mode. An alternative to the merge mode is explicit signaling of motion parameters, in which motion vectors (more precisely, motion vector difference (MVD) compared to a motion vector predictor), corresponding reference picture indices for each reference picture list, and reference picture list usage are explicitly signaled for each PU. Such a mode is referred to as advanced motion vector prediction (AMVP) in this disclosure.

[0043] When signaling indicates that one of two reference picture lists is used, the PU is generated from samples of one block. This is called "uni-prediction." Uni-prediction is available for both P slices and B slices.

[0044] When signaling indicates that both reference picture lists are used, the PU is generated from samples of two blocks. This is called "bi-prediction." Bi-prediction is only available for B slices.

[0045] The following text provides details about the inter prediction modes specified in HEVC, starting with merge mode.

[0046] Merge Mode 2.1.1.1. Deriving Candidates for Merge Modes When a PU is predicted using merge mode, an index pointing to an entry in the merge candidate list is parsed from the bitstream and used to extract the motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps: Step 1: Derive initial candidates - Step 1.1: Spatial candidate derivation - Step 1.2: Redundancy check on spatial candidates - Step 1.3: Derive time candidates Step 2: Insert additional candidates - Step 2.1: Creating bi-prediction candidates - Step 2.2: Insertion of zero motion candidates

[0047] Figure 1 also shows a schematic of these steps. For spatial merge candidate derivation, up to four merge candidates are selected from candidates at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Because a fixed number of candidates is assumed for each PU at the decoder, additional candidates are generated if the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Because the number of candidates is fixed, the index of the best merge candidate is coded using truncated unary binarization (TU). When the size of a CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of a 2N × 2N prediction unit.

[0048] The following provides a detailed description of the processes associated with the above steps.

[0049] 2.1.1.2. Spatial candidate derivation In the derivation of spatial merge candidates, up to four merge candidates are selected from the candidates located at the positions shown in FIG. 2. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs located at positions A1, B1, B0, and A0 is unavailable (e.g., because it belongs to another slice or tile) or is intra-coded. After the candidate located at position A1 is added, the remaining candidates are subjected to a redundancy check to ensure that candidates with the same motion information are removed from the list, thereby improving coding efficiency. To reduce computational complexity, the aforementioned redundancy check does not consider all possible candidate pairs. Instead, only pairs connected by arrows in FIG. 3 are considered, and a candidate is added to the list only if the corresponding candidate used in the redundancy check does not have the same motion information. Another source of overlapping motion information is a "second PU" associated with a partition different from 2N×2N. As an example, FIG. 4 shows the second PUs for N×2N and 2N×N cases, respectively. When the current PU is divided into Nx2N, the candidate at A1 is not considered for list construction. In fact, adding this candidate would lead to two prediction units with the same motion information, which is redundant for having only one PU in the coding unit. Similarly, when the current PU is divided into 2NxN, the position B1 is not considered.

[0050] 2.1.1.3. Time candidate derivation In this step, only one candidate is added to the list. In particular, in this derivation of the temporal merge candidate, a scaled motion vector is derived based on the co-located PU belonging to the picture with the smallest POC difference from the current picture in a given reference picture list. The reference picture list used to derive the co-located PU is explicitly signaled in the slice header. The scaled motion vector for the temporal merge candidate is obtained as shown by the dotted line in FIG. 5 , which is scaled from the motion vector of the co-located PU (col_PU) using POC distances tb and td, where tb is defined as the POC difference between the reference picture (curr_ref) of the current picture (curr_pic) and the current picture, and td is defined as the POC difference between the reference picture (col_ref) of the co-located picture (col_pic) and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. The actual implementation of the scaling process is defined in the HEVC specification. For a B slice, two motion vectors are obtained, one with respect to reference picture list 0 and the other with respect to reference picture list 1, which are combined to form a bi-predictive merge candidate.

[0051] FIG. 5 shows an illustration of motion vector scaling for temporal merge candidates.

[0052] For a co-located PU(Y) belonging to the reference frame, a position for a temporal candidate is selected between candidates C0 and C1, as shown in Figure 6. If the PU at position C0 is unavailable, or is intra-coded, or is outside the current CTU row, position C1 is used. Otherwise, position C0 is used to derive the temporal merge candidate.

[0053] FIG. 6 shows example candidate positions for temporal merge candidates C0 and C1.

[0054] 2.1.1.4. Inserting additional candidates In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidates and zero merge candidates. Combined bi-predictive merge candidates are generated by utilizing spatial and temporal merge candidates. Combined bi-predictive merge candidates are used only for B slices. Combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of one original candidate with the second reference picture list motion parameters of another. If these two tuples provide different motion hypotheses, they form a new bi-predictive candidate. As an example, Figure 7 shows the case where two candidates in the original list (left side), with mvL0 and refIdxL0, or mvL1 and refIdxL1, are used to create a combined bi-predictive merge candidate that is added to the final list (right side). There are many rules regarding the combinations considered to generate these additional merge candidates.

[0055] Zero motion candidates are inserted to fill the remaining entries in the merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have a spatial displacement of zero and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. The number of reference frames used by these candidates is 1 and 2 for unidirectional and bidirectional prediction, respectively. Finally, no redundancy check is performed on these candidates.

[0056] 2.1.1.5 Motion Estimation Regions for Parallel Processing To speed up the encoding process, motion estimation can be performed in parallel, whereby motion vectors for all prediction units in a given region are derived simultaneously. Deriving merge candidates from spatial neighbors can hinder parallel processing, as one prediction unit cannot derive motion parameters from neighboring PUs until its associated motion estimation is complete. To mitigate the tradeoff between coding efficiency and processing latency, HEVC defines a motion estimation region (MER), whose size is signaled in the picture parameter set using the “log2_parallel_merge_level_minus2” syntax element. When an MER is specified, merge candidates that fall in the same region are marked as unavailable and therefore not considered in list creation.

[0057] 2.1.2 AMVP AMVP exploits the spatial-temporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the upper-left temporally neighboring PU position, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best predictor from the candidate list and transmit the corresponding index pointing to the selected candidate. Similarly, in merge index signaling, the index of the best motion vector candidate is coded using a truncated unary. The maximum value coded in this case is 2 (see Figure 8). The following section provides details of the motion vector prediction candidate derivation process.

[0058] 2.1.2.1 Derivation of AMVP Candidates In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. In spatial motion vector candidate derivation, two motion vector candidates are ultimately derived based on the motion vectors of each PU at five different positions as shown in Figure 2.

[0059] In temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates that are derived based on two different co-located positions. After a first list of spatio-temporal candidates is created, duplicate motion vector candidates in the list are removed. If the number of possible candidates is greater than two, motion vector candidates whose reference picture index in the associated reference picture list is greater than one are removed from the list. If the number of spatio-temporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.

[0060] 2.1.2.2. Spatial Motion Vector Candidates In deriving spatial motion vector candidates, up to two candidates out of five possible candidates are considered, which are derived from PUs located as shown in Figure 2, and their positions are the same as the positions of the motion merge. The derivation order for the left side of the current PU is defined as A0, A1, scaled A0, scaled A1. The derivation order for the top side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two cases do not need to use spatial scaling, and two cases do use spatial scaling. The four different cases are summarized as follows: No spatial scaling - (1) Same reference picture list and same reference picture index (same POC) - (2) Different reference picture lists, but the same reference picture (same POC) Spatial scaling - (3) Same reference picture list, but different reference pictures (different POC) - (4) Different reference picture lists and different reference pictures (different POCs)

[0061] The case without spatial scaling is checked first, followed by spatial scaling. Spatial scaling is considered when the POC differs between the reference picture of the neighboring PU and the reference picture of the current PU, regardless of the reference picture list. If all PUs of the left candidate are unavailable or intra-coded, scaling is enabled for the upper motion vector to facilitate parallel derivation of left and upper MV candidates. Otherwise, spatial scaling is not enabled for the upper motion vector.

[0062] FIG. 9 shows an explanatory diagram of motion vector scaling for spatial motion vector candidates.

[0063] In the spatial scaling process, the motion vectors of neighboring PUs are scaled in a similar manner as for temporal scaling, as shown in Figure 9. The main difference is that the reference picture list and the index of the current PU are given as input, and the actual scaling process is the same as that of temporal scaling.

[0064] 2.1.2.3. Temporal Motion Vector Candidates Apart from the derivation of the reference picture index, all the processes for the derivation of temporal merge candidates are the same as for the derivation of spatial motion vector candidates (see Figure 6). The reference picture index is signaled to the decoder.

[0065] 2.2. New Inter-Prediction Method in JEM 2.2.1 Sub-CU-based motion vector prediction In JEM with QTBT, each CU can have at most one set of motion parameters for each prediction direction. Two sub-CU level motion vector prediction methods are considered in the encoder by dividing a large CU into multiple sub-CUs and deriving motion information for all sub-CUs of the large CU. The alternative temporal motion vector prediction (ATMVP) method allows each CU to fetch multiple sets of motion information from multiple blocks smaller than the current CU in a co-located reference picture. In the spatial-temporal motion vector prediction (STMVP) method, the motion vector of a sub-CU is recursively derived by using a temporal motion vector predictor and spatial neighboring motion vectors.

[0066] To preserve more accurate motion fields for sub-CU motion estimation, motion compression on reference frames is currently disabled.

[0067] 2.2.1.1. Alternative Temporal Motion Vector Prediction In the alternative temporal motion vector prediction (ATMVP) method, the temporal motion vector prediction (TMVP) of the motion vector is modified by fetching multiple sets of motion information (including motion vectors and reference indexes) from multiple blocks smaller than the current CU. As shown in Figure 10, a sub-CU is a square NxN block (N is set to 4 by default).

[0068] ATMVP predicts motion vectors for multiple sub-CUs within a CU in two steps. The first step is to identify corresponding blocks in a reference picture 1050 using a so-called temporal vector. The reference picture 1050 is also called a motion source picture. The second step is to divide the current CU 1000 into sub-CUs 1001, as shown in Figure 10, and obtain the motion vectors and reference indexes of each sub-CU from the blocks corresponding to each sub-CU.

[0069] In the first step, the reference picture and corresponding block are determined by the motion information of the spatial neighboring blocks of the current CU. To avoid repeated scanning of neighboring blocks, the first merge candidate in the merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set to be the temporal vector and index into the motion source picture. Thus, in ATMVP, corresponding blocks can be identified more accurately compared to TMVP, and corresponding blocks (sometimes called co-located blocks) are always located in the bottom-right or center position relative to the current CU.

[0070] In the second step, the corresponding block of the sub-CU is identified by the time vector in the motion source picture by adding the time vector to the coordinates of the current CU. For each sub-CU, the motion information of the corresponding block (the smallest motion grid covering the center sample) is used to derive the motion information for that sub-CU. After the motion information of the corresponding NxN block is identified, it is converted into the motion vector and reference index of the current sub-CU, similar to TMVP in HEVC, where motion scaling and other procedures are applied. For example, the decoder checks whether the low-delay condition (i.e., the POC of all reference pictures of the current picture is smaller than the POC of the current picture) is met, and then uses the motion vector MVx (the motion vector corresponding to reference picture list X) to derive the motion vector MV for each sub-CU. y (where X is equal to 0 or 1 and Y is equal to 1-X).

[0071] 2.2.1.2. Spatial-Temporal Motion Vector Prediction In this method, motion vectors for sub-CUs are derived recursively according to the raster scan order. Figure 11 illustrates this concept. Consider an 8x8 CU containing four 4x4 sub-CUs A, B, C, and D. Label the adjacent 4x4 blocks in the current frame as a, b, c, and d.

[0072] Motion derivation for sub-CU A begins by identifying its two spatial neighbors. The first neighbor is an N×N block (block c) above sub-CU A. If block c is unavailable or intra-coded, other N×N blocks above sub-CU A are examined (starting from block c and working from left to right). The second neighbor is a block to the left of sub-CU A (block b). If block b is unavailable or intra-coded, other blocks to the left of sub-CU A are examined (starting from block b and working from top to bottom). The motion information obtained from these neighboring blocks for each list is scaled relative to the first reference frame for a given list. Next, a temporal motion vector predictor (TMVP) for sub-block A is derived by following the same TMVP derivation procedure as specified in HEVC. The motion information of the co-located block at position D is fetched and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to three) are averaged separately for each reference list. The averaged motion vector is assigned as the motion vector of the current sub-CU.

[0073] 2.2.1.3. Sub-CU Motion Prediction Mode Signaling Sub-CU modes are enabled as additional merge candidates, and no additional syntax elements are required to signal these modes. Two additional merge candidates are added to the merge candidate list of each CU to represent ATMVP and STMVP modes. Up to seven merge candidates are used when the sequence parameter set indicates that ATMVP and STMVP are enabled. The encoding logic for these additional merge candidates is the same as for merge candidates in HM, which means that for each CU in a P slice or B slice, two more RD checks are required for the two additional merge candidates.

[0074] In JEM, all bins of the merge index are context coded by CABAC, whereas in HEVC, only the first bin is context coded and the remaining bins are context-bypass coded.

[0075] 2.2.2. Adaptive Motion Vector Difference Resolution In HEVC, when use_integer_mv_flag is equal to 0 in the slice header, the motion vector difference (MVD) (between a motion vector and a PU's predicted motion vector) is signaled in units of 1 / 4 luma sample. In JEM, locally adaptive motion vector resolution (LAMVR) is introduced. In JEM, MVD can be coded in units of 1 / 4 luma sample, integer luma sample, or 4 luma sample. This MVD resolution is controlled at the coding unit (CU) level, and an MVD resolution flag is conditionally signaled for each CU that has at least one non-zero MVD component.

[0076] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in that CU. If the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or 4 luma sample MV precision is used.

[0077] If the first MVD resolution flag of a CU is zero or not coded for the CU (meaning all MVDs in the CU are zero), 1 / 4 luma sample MV resolution is used for that CU. If the CU uses integer luma sample MV precision or 4 luma sample MV precision, the MVPs in the AMVP candidate list for that CU are rounded to the corresponding precision.

[0078] At the encoder, a CU-level RD check is used to determine which MVD resolution should be used for a CU, i.e., three CU-level RD checks are performed, one for each MVD resolution. To accelerate the encoder speed, the following coding scheme is applied in JEM: In the RD test of a CU at regular 1 / 4 luma sample MVD resolution, the motion information (integer luma sample accuracy) of the current CU is stored. The stored motion information (after rounding) is used as the starting point for a smaller range of motion vector refinement in the RD test of the same CU at integer luma sample and 4 luma sample MVD resolution, thereby avoiding the time-consuming motion estimation process being repeated three times.

[0079] RD check for CUs at 4 luma sample MVD resolution is conditionally invoked. If the RD cost for a CU at integer luma sample MVD resolution is much larger than that at 1 / 4 luma sample MVD resolution, RD check for that CU at 4 luma sample MVD resolution is skipped.

[0080] 2.2.3 Higher Motion Vector Storage Accuracy In HEVC, motion vector precision is 1 / 4 pel (1 / 4 luma sample and 1 / 8 chroma sample for 4:2:0 video). In JEM, the precision of the internal motion vector storage and merge candidates is increased to 1 / 16 pel. This higher motion vector precision (1 / 16 pel) is used in motion compensated inter prediction for CUs coded in skip / merge mode. For CUs coded in regular AMVP mode, either integer-pel or 1 / 4-pel motion is used, as described in Section 0.

[0081] An SHVC upsampling interpolation filter, which has the same filter length and normalization factor as the HEVC motion compensated interpolation filter, is used as the motion compensated interpolation filter for the additional fractional pel positions. The chroma component motion vector precision is 1 / 32 sample in JEM, and the additional interpolation filter for the 1 / 32-pel fractional position is obtained by using the average of the filters for two adjacent 1 / 16-pel fractional positions.

[0082] 2.2.4 Overlapped Block Motion Compensation Previously, overlapped block motion compensation (OBMC) has been used in H.263. Unlike in H.263, JEM allows OBMC to be switched on and off using CU-level syntax. When OBMC is used in JEM, it is performed on all motion compensation (MC) block boundaries except the right and bottom boundaries of a CU. Furthermore, this applies to both luma and chroma components. In JEM, an MC block corresponds to a coding block. When a CU is coded in a sub-CU mode (including sub-CU merge, affine, and FRUC modes), each sub-block of that CU is an MC block. To handle CU boundaries in a uniform manner, OBMC is performed at the sub-block level for all MC block boundaries by setting the sub-block size equal to 4×4, as shown in FIG. 12.

[0083] When OBMC is applied to a current sub-block, in addition to the current motion vector, the motion vectors of the four adjacent neighboring sub-blocks, if available and not identical to the current motion vector, are also used to derive a prediction block for the current sub-block, and these multiple prediction blocks based on multiple motion vectors are combined to generate a final prediction signal for the current sub-block.

[0084] The predicted block based on the motion vector of the neighboring sub-block is P, where N indicates the indexes for the neighboring upper, lower, left and right sub-blocks. N and the predicted block based on the motion vector of the current sub-block is P C It is written as P N If P is based on the motion information of neighboring sub-blocks that contain the same motion information as the current sub-block, then OBMC is N otherwise, P N All samples of P C are summed to the same sample in P N The four rows / columns of P C The weighting coefficients {1 / 4, 1 / 8, 1 / 16, 1 / 32} are added to P N The weighting factors {3 / 4,7 / 8,15 / 16,31 / 32} are used for P C The exception is small MC blocks (i.e., when the height or width of the coding block is equal to 4, or when the CU is coded in sub-CU mode), in which case P N Only two rows / columns of P are added to Pc. In this case, the weighting coefficients {1 / 4, 1 / 8} are added to Pc. N The weighting factors {3 / 4,7 / 8} are used for P C is used for P, which is generated based on the motion vectors of vertically (horizontally) adjacent sub-blocks. N So, P N The samples in the same row (column) of P C are added together.

[0085] In JEM, for CUs with a size of 256 luma samples or less, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU. For CUs with a size larger than 256 luma samples or CUs not coded in AMVP mode, OBMC is applied by default. When OBMC is applied to a CU in the encoder, its impact is taken into account during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the upper and left neighboring blocks is used to compensate the upper and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.

[0086] 2.2.5 Local illumination compensation Local Illumination Compensation (LIC) is based on a linear model of illumination changes with a scaling factor a and an offset b, and is adaptively enabled or disabled for each inter-mode coded coding unit (CU).

[0087] FIG. 13 shows an example of adjacent samples used to derive IC parameters.

[0088] When LIC is applied to a CU, a least square error method is adopted to derive parameters a and b by using neighboring samples of the current CU and their corresponding reference samples. More specifically, as illustrated in Figure 13, subsampled (2:1 subsampled) neighboring samples of the CU and corresponding samples in the reference picture (identified by the motion information of the current CU or current sub-CU) are used. IC parameters are derived and applied separately for each prediction direction.

[0089] If the CU is coded in merge mode, the LIC flag is copied from the neighboring block in a manner similar to the motion information copying in merge mode; otherwise, the LIC flag is signaled to indicate whether LIC is applied or not for the CU.

[0090] When LIC is enabled for a picture, an additional CU-level RD check is required to determine whether LIC applies to the CU. When LIC is enabled for a CU, the mean-removed sum of absolute difference (MR-SAD) and the mean-removed sum of absolute Hadamard-transformed difference (MR-SATD) are used instead of SAD and SATD for integer-pixel motion search and fractional-pixel motion search, respectively.

[0091] To reduce the encoding complexity, an encoding scheme is applied in JEM: LIC is disabled for the entire picture when there is no obvious illumination change between the current picture and its reference pictures. To identify this situation, the encoder calculates the histogram of the current picture and the histograms of all of its reference pictures. If the histogram difference between the current picture and all of its reference pictures is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.

[0092] 2.2.6 Affine Motion Compensated Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). Meanwhile, in the real world, many types of motion exist, such as zoom in / out, rotation, perspective motion, and other irregular motions. In JEM, a simplified affine transformation motion compensation prediction is applied. As shown in Figure 14, the affine motion field of a block is described by two control point motion vectors.

[0093] The motion vector field (MVF) of a block is given by:

number

[0094] where (v 0x ,v 0y ) is the motion vector of the control point in the upper left corner, and (v 1x ,v 1y ) is the motion vector of the control point in the upper right corner.

[0095] To further simplify the motion compensation prediction, sub-block-based affine transformation prediction is applied. The sub-block size M×N is expressed as Equation 2:

number

[0096] After being derived by Equation 2, M and N should be adjusted downwards if necessary to make them divisors of w and h, respectively.

[0097] To derive a motion vector for each M×N sub-block, the motion vector of the center sample of each sub-block is calculated according to Equation 1 and rounded to 1 / 16 fractional precision, as shown in Figure 15. Then, the motion compensated interpolation filter described in Section 0 is applied to generate a prediction for each sub-block with the derived motion vector.

[0098] After the MCP, the high-precision motion vectors for each sub-block are rounded and stored with the same precision as the regular motion vectors.

[0099] In JEM, there are two affine motion modes: AF_INTER mode and AF_MERGE mode. For CUs whose width and height are both greater than 8, the AF_INTER mode can be applied. An affine flag at the CU level is signaled in the bitstream to indicate whether the AF_INTER mode is used. In this mode, the motion vector pair {(v0,v1)|v0={v A ,v B ,v C},v1={v D ,v E A candidate list with {}} is constructed using neighboring blocks. As shown in Figure 16, v0 is selected from the motion vectors of blocks A, B, or C. The motion vectors from neighboring blocks are scaled according to the reference list and the relationship between the POC of the reference for the neighboring block, the POC of the reference for the current CU, and the POC of the current CU. The approach for selecting v1 from neighboring blocks D and E is similar. If the number of candidate lists is less than two, the list is padded with motion vector pairs constructed by duplicating each of the AMVP candidates. If the candidate list is greater than two, the candidates are first sorted according to the consistency of adjacent motion vectors (the similarity of the two motion vectors in a pair of candidates), and only the first two candidates are retained. An RD cost test is used to determine which motion vector pair candidate is selected as the control point motion vector prediction (CPMVP) for the current CU. An index indicating the position of the CPMVP in the candidate list is then signaled in the bitstream. After the CPMVP of the current affine CU is determined, affine motion estimation is applied to find the control point motion vector (CPMV), and the difference between the CPMV and the CPMVP is signaled in the bitstream.

[0100] When a CU is applied in AF_MERGE mode, it obtains the first block coded in affine mode from the valid neighboring reconstructed blocks. The selection order of candidate blocks is from left to right, top, top right, bottom left, top left, as shown in Figure 17A. If the neighboring bottom-left block is coded in affine mode, as shown in Figure 17B, motion vectors v2, v3, and v4 are derived for the top-left, top-right, and bottom-left corners of the CU containing block A. Then, the top-left motion vector v0 of the current CU is calculated according to v2, v3, and v4. Next, the top-right motion vector v1 of the current CU is calculated.

[0101] After the CPMVs v0 and v1 of the current CU are calculated according to Equation 1 of the simplified affine motion model, the MVF of the current CU is generated. To specify whether the current CU is coded in AF_MERGE mode, an affine flag is signaled in the bitstream when at least one neighboring block is coded in affine mode.

[0102] 2.2.7 Pattern Matching Motion Vector Derivation The pattern matched motion vector derivation (PMMVD) mode is a special merge mode based on the Frame-Rate Up Conversion (FRUC) technique, in which the motion information of blocks is derived at the decoder side without being signaled.

[0103] If the merge flag is true, the FRUC flag is signaled for the CU. If the FRUC flag is false, the merge index is signaled and normal merge mode is used. If the FRUC flag is true, an additional FRUC mode flag is signaled to indicate which method (bilateral matching or template matching) is used to derive motion information for the block.

[0104] At the encoder side, the decision to use the FRUC merge mode for a CU is based on RD cost selection as is done for regular merge candidates. That is, for a CU, two matching modes (bilateral matching and template matching) are both examined using RD cost selection. The one that leads to the smallest cost is then compared with the other CU modes. If the FRUC matching mode is the most efficient one, the FRUC flag is set to true for that CU and the associated matching mode is used.

[0105] The motion derivation process in FRUC merge mode has two steps. First, a CU-level motion search is performed, followed by sub-CU-level motion refinement. At the CU level, an initial motion vector is derived for the entire CU based on bilateral matching or template matching. First, a list of MV candidates is generated, and the candidate that leads to the minimum matching cost is selected as the starting point for further CU-level refinement. Then, a local search based on bilateral matching or template matching is performed around the starting point, and the MV that leads to the minimum matching cost is taken as the MV for the entire CU. Subsequently, using the derived CU motion vector as a starting point, motion information is further refined at the sub-CU level.

[0106] For example, for deriving W×H CU motion information, the following derivation process is performed: In the first step, the MV for the entire W×H CU is derived; In the second step, the CU is further divided into M×M sub-CUs. The value of M is calculated as in equation (16), and D is a predetermined division depth, which is set to 3 by default in JEM.

number

[0107] As shown in Figure 18, bilateral matching is used to derive the motion information of a current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. Under the assumption of continuous motion trajectories, the motion vectors MV0 and MV1 pointing to two reference blocks are proportional to the temporal distances between the current picture and the two reference pictures, i.e., TD0 and TD1. As a special case, when the current picture is temporally between two reference pictures and the temporal distances from the current picture to the two reference pictures are the same, bilateral matching becomes a mirror-based bidirectional MV.

[0108] As shown in Figure 19, template matching is used to derive motion information for the current CU by finding the closest match between a template in the current picture (the neighboring block above and / or to the left of the current CU) and a block in the reference picture (the same size as the template). Except for the FRUC merge mode mentioned above, template matching is also applied to AMVP mode. In JEM, as in HEVC, AMVP has two candidates. A new candidate is derived using the template matching method. If the newly derived candidate by template matching is different from the first existing AMVP candidate, it is inserted into the first part of the AMVP candidate list, and the list size is set to 2 (meaning the second existing AMVP candidate is deleted). When applied to AMVP mode, only CU-level search is applied.

[0109] 2.2.7.1 CU-level MV candidate set MV candidates set at the CU level are: (i) If the CU is currently in AMVP mode, the original AMVP candidate (ii) All merge candidates (iii) Some MVs in the interpolated MV field introduced in Section 0 (iv) Upper and left adjacent motion vectors It consists of:

[0110] When bilateral matching is used, each valid MV of a merge candidate is used as an input to generate an MV pair assuming bilateral matching. For example, one valid MV of a merge candidate is (MVa, refa) in reference list A. Then, the reference picture refb of the paired bilateral MV is found in the other reference list B, such that refa and refb are on different temporal sides of the current picture. If such refb is not available in reference list B, refb is determined as a reference in list B that is different from refa and has the smallest temporal distance to the current picture. After refb is determined, MVb is derived by scaling MVa based on the temporal distance between the current picture and refa and refb.

[0111] Four MVs from the interpolated MV field are also added to the CU-level candidate list. More specifically, the interpolated MVs at positions (0,0), (W / 2,0), (0,H / 2), and (W / 2,H / 2) of the current CU are added.

[0112] When FRUC is applied in AMVP mode, the original AMVP candidates are also added to the CU-level MV candidate set.

[0113] At the CU level, up to 15 MVs for an AMVP CU and up to 13 MVs for a merged CU are added to the candidate list.

[0114] 2.2.7.2 Sub-CU level MV candidate set MV candidates configured at the sub-CU level are: (i) MV determined from CU-level exploration (ii) Top, left, top-left, and top-right adjacent MVs (iii) Scaled co-located MVs from the reference picture (iv) Up to four ATMVP candidates (v) Up to four STMVP candidates It consists of:

[0115] The scaled MV from the reference picture is derived as follows: All reference pictures in both lists are considered. The MV at the co-located position of the sub-CU in the reference picture is scaled to the reference of the source CU-level MV.

[0116] ATMVP and STMVP candidates are limited to the first four candidates.

[0117] At the sub-CU level, up to 17 MVs are added to the candidate list.

[0118] 2.2.7.3 Generating Interpolated MV Fields Before encoding a frame, an interpolated motion field is generated for the whole picture based on unilateral ME, and the motion field can later be used as a CU-level or sub-CU-level MV candidate.

[0119] First, the motion field of each reference picture in both reference lists is considered at the 4x4 block level. For each 4x4 block, if the motion associated with that block through the 4x4 block in the current picture (as shown in Figure 20) does not have any interpolated motion assigned, the motion of the reference block is scaled to the current picture according to the temporal distances TD0 and TD1 (the same way as MV scaling of TMVP in HEVC), and the scaled motion is assigned to the block in the current frame. If a 4x4 block does not have a scaled MV assigned, the motion of that block is marked as unavailable in the interpolated motion field.

[0120] 2.2.7.4 Interpolation and Matching Costs When motion vectors point to fractional sample positions, motion compensated interpolation is required. To reduce complexity, bilinear interpolation is used instead of the usual 8-tap HEVC interpolation for both bilateral and template matching.

[0121] The calculation of the matching cost is slightly different in different steps. When selecting a candidate from the candidate set at the CU level, the matching cost is the sum of absolute differences (SAD) of bilateral matching or template matching. After the starting MV is determined, the matching cost C of bilateral matching in the sub-CU level search is calculated as follows:

number

[0122] In FRUC mode, MV is derived using only luma samples. The derived motion will be used for both luma and chroma for MC inter prediction. After the MV is determined, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.

[0123] 2.2.7.5 MV Refinement MV refinement is a pattern-based MV search using the bilateral matching cost or template matching cost criterion. JEM supports two search patterns: unrestricted center-biased diamond search (UCBDS) and adaptive cross search for MV refinement at the CU and sub-CU levels. In both CU and sub-CU level MV refinement, MVs are directly searched with 1 / 4 luma sample MV precision, followed by 1 / 8 luma sample MV refinement. The search range for MV refinement in the CU and sub-CU steps is set equal to 8 luma samples.

[0124] 2.2.7.6 Prediction Direction Selection in Template Matching FRUC Merge Mode In bilateral matching merge mode, bi-prediction is always applied, since the motion information of a CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. Template matching merge mode has no such restriction. In template matching merge mode, the encoder can choose between uni-prediction from list0, uni-prediction from list1, or bi-prediction for the CU. The choice is based on the cost of template matching as follows: If costBi<=factor*min(cost0,cost1), Bi-prediction is used; Otherwise, if cost0<=cost1, One prediction from list0 is used; In other cases, A single prediction from list1 is used; where cost0 is the SAD of list0 template matching, cost1 is the SAD of list1 template matching, and costBi is the SAD of bi-predictive template matching. The value of factor is equal to 1.25, which means that this selection process is biased towards bi-prediction.

[0125] The inter prediction direction selection is only applied to the CU-level template matching process.

[0126] 2.2.8 Generalized Biprediction Improvement The Generalized Bi Prediction Improvement (GBi) proposed in JVET-L0646 has been adopted in VTM--3.0.

[0127] GBi was proposed in JVET-C0047. JVET-K0248 improved the gain-complexity tradeoff in GBi and adopted it in BMS2.1. BMS2.1 GBi applies unequal weights to the predictor from L0 and the predictor from L1 in bi-prediction mode. In inter-prediction mode, multiple weight pairs, including the equal weight pair (1 / 2, 1 / 2), are evaluated based on rate-distortion optimization (RDO), and the GBi index of the selected weight pair is signaled to the decoder. In merge mode, the GBi index is inherited from the neighboring CU. In BMS2.1 GBi, predictor generation in bi-prediction mode is shown in Equation (1).

[0128] P GBi =(w0*P L0 +w1*P L1 +RoundingOffset GBi )>>shiftNum GBi where P GBi is the final predictor of GBi, and w0 and w1 are the selected GBi weight pairs applied to the predictors in list 0 (L0) and list 1 (L1), respectively. GBi and shiftNum GBi is used to normalize the final predictor in GBi. The supported w1 weight set is {-1 / 4, 3 / 8, 1 / 2, 5 / 8, 5 / 4}, where these five weights correspond to one equal weight pair and four unequal weight pairs. The blend gain, i.e., the sum of w1 and w0, is fixed to 1.0. Therefore, the corresponding w0 weight set is {5 / 4, 5 / 8, 1 / 2, 3 / 8, -1 / 4}. This weight pair selection is at the CU level.

[0129] For non-low latency pictures, the weight set size is reduced from 5 to 3, with the w1 weight set being {3 / 8, 1 / 2, 5 / 8} and the w0 weight set being {5 / 8, 1 / 2, 3 / 8}. The weight set size reduction for non-low latency pictures applies to BMS2.1 GBi and all GBi tests in this contribution.

[0130] In this JVET-L0646, a combined solution based on JVET-L0197 and JVET-L0296 is proposed to further improve GBi performance. Specifically, the following changes are applied on top of the existing GBi design of BMS2.1:

[0131] 2.2.8.1 GBi encoder bug fix To reduce GBi encoding time, in the current encoder design, the encoder stores unidirectionally predicted motion vectors estimated from a GBi weight equal to 4 / 8 and reuses them for unidirectionally predicted searches of other GBi weights. This fast encoding method applies to both translational and affine motion models. In VTM2.0, a 6-parameter affine model was adopted along with a 4-parameter affine model. The BMS2.1 encoder does not distinguish between the 4-parameter affine model and the 6-parameter affine model when storing unidirectionally predicted affine MVs when the GBi weight is equal to 4 / 8. Therefore, after encoding with a GBi weight of 4 / 8, the 4-parameter affine MV may be overwritten by the 6-parameter affine MV. The stored 6-parameter affine MV can be used for 4-parameter affine ME for other GBi weights, or the stored 4-parameter affine MV can be used for 6-parameter affine ME. The proposed GBi encoder bug fix separates the 4-parameter affine MV storage from the 6-parameter affine MV storage. The encoder stores the affine MVs based on the affine model type if the GBi weights are equal to 4 / 8, and for other GBi weights, reuses the corresponding affine MVs based on the affine model type.

[0132] 2.2.8.2 GBi encoder speedup Five encoder speed-up methods are proposed to reduce encoding time when GBi is enabled.

[0133] (1) Conditionally skip affine motion estimation for some GBi weights.

[0134] In BMS2.1, affine ME, including 4-parameter and 6-parameter affine ME, is performed for all GBi weights. We propose to conditionally skip affine ME for unequal GBi weights (weights not equal to 4 / 8). Specifically, affine ME is performed for other GBi weights only if affine mode is selected as the current best mode after evaluating 4 / 8 GBi weights and it is not affine merge mode. If the current picture is a non-low latency picture, when affine ME is performed, bi-predictive ME of the translational model is skipped for unequal GBi weights. If affine mode is not selected as the current best mode or affine merge is selected as the current best mode, affine ME is skipped for all other GBi weights.

[0135] (2) Reduce the number of weights for RD cost check for low-delay pictures in 1-pel and 4-pel MVD precision encoding.

[0136] For low-latency pictures, there are five weights for RD cost checking at all MVD precisions, including 1 / 4-pel, 1-pel, and 4-pel. The encoder will first check the RD cost at 1 / 4-pel MVD precision. We propose to skip some of the GBi weights for the RD cost checking at 1-pel and 4-pel MVD precision. We order the unequal weights according to their RD costs at 1 / 4-pel MVD precision. Only the first two weights with the smallest RD costs, along with the GBi weight 4 / 8, will be evaluated for encoding at 1-pel and 4-pel MVD precision. Therefore, for low-latency pictures, at most three weights will be evaluated for 1-pel and 4-pel MVD precision.

[0137] (3) Conditionally skip bi-predictive search if the L0 reference picture and the L1 reference picture are the same.

[0138] For some pictures in RA, the same picture may occur in both reference picture lists (list-0 and list-1). For example, in a random access coding structure in CTC, the reference picture structure for the first group of pictures (GOP) is listed as follows: POC:16, TL:0, [L0: 0 ][L1: 0 ] POC:8, TL:1, [L0: 0 16 ] [L1: 16 0 ] POC:4, TL:2, [L0:0 8 ] [L1: 8 16] POC:2,TL:3,[L0:0 4 ] [L1: 4 8] POC:1, TL:4, [L0:0 2 ] [L1: 2 4] POC:3, TL:4, [L0:2 0] [L1:4 8] POC:6,TL:3,[L0:4 0] [L1:8 16] POC:5, TL:4, [L0:4 0] [L1:68] POC:7, TL:4, [L0:6 4] [L1:8 16] POC:12, TL:2, [L0: 8 0] [L1:16 8 ] POC:10, TL:3, [L0:8 0] [L1:12 16] POC:9, TL:4, [L0:8 0] [L1:10 12] POC:11, TL:4, [L0:10 8] [L1:12 16] POC:14, TL:3, [L0: 12 8] [L1: 12 16] POC:13, TL:4, [L0:12 8] [L1:14 16] POC:15, TL:4, [L0: 14 12] [L1:1614 ]

[0139] It can be seen that pictures 16, 8, 4, 2, 1, 12, 14, and 15 have the same reference picture in both lists. In bi-prediction for these pictures, it is possible that the reference picture is the same in L0 and L1. We propose that the encoder skip bi-predictive ME for unequal GBi weights if 1) the two reference pictures in bi-prediction are the same, and 2) the temporal layer is greater than 1, and 3) the MVD precision is 1 / 4 pel. In affine bi-predictive ME, this fast skip method only applies to 4-parameter affine ME.

[0140] (4) Skip RD cost check for unequal GBi weights based on the temporal layer and the POC distance between the reference picture and the current picture.

[0141] We propose to skip the RD cost evaluation for unequal GBi weights when the temporal stratum is equal to 4 (the highest temporal stratum in RA) or the POC distance between the reference picture (either list-0 or list-1) and the current picture is equal to 1 and the coding QP is greater than 32.

[0142] (5) Change floating-point calculations to fixed-point calculations for unequal GBi weights in ME.

[0143] In existing bi-predictive search, the encoder fixes the MV of one list and refines the MV of the other list. To reduce computational complexity, the target is modified before ME. For example, if the MV of list-1 is fixed and the encoder refines the MV of list-0, the target of MV refinement of list-0 is modified by Equation (5). O is the original signal, P1 is the predicted signal of list-1, and w is the GBi weight for list-1.

number

[0144] Here, the term (1 / (8-w)) is stored in floating-point precision, which increases the computational complexity. We convert equation (5) to equation (6):

number

number

[0145] 2.2.8.3 CU Size Constraints for GBi In this method, GBi is disabled for small CUs. In inter prediction mode, if bi-prediction is used and the CU region is smaller than 128 luma samples, GBi is disabled without any signaling.

[0146] 2.2.9 Bidirectional Optical Flow 2.2.9.1 Theoretical analysis In BIO, motion compensation is first performed to generate a first prediction (in each prediction direction) of the current block. The first prediction is used to derive spatial gradients, temporal gradients, and optical flow for each sub-block / pixel within the block, which are then used to generate a second prediction, i.e., the final prediction for the sub-block / pixel. The details are as follows:

[0147] Bidirectional optical flow (BIO) is a sample-by-sample motion refinement performed on top of block-by-block motion compensation for bidirectional prediction. This sample-level motion refinement does not use signaling.

[0148] FIG. 21 shows an example of an optical flow trajectory.

[0149] The luma value from reference k (k=0,1) after block motion compensation is I (k) Let ∂I (k) / ∂x, ∂I (k) / ∂y are I (k) are the horizontal and vertical components of the gradient. Assuming that the optical flow is valid, the motion vector field (v x ,v y ) is the expression:

number

[0150] Combining this optical flow equation with Hermite interpolation for the motion trajectory of each sample yields the function value I (k) and derivative ∂I (k) / ∂x, ∂I (k) We obtain a unique cubic polynomial that fits both ∂y and ∂y. The value of this polynomial at t=0 is the BIO prediction:

number

[0151] where τ0 and τ1 represent the distance to the reference frame as shown in FIG. 21. The functions τ0 and τ1 are calculated based on the POCs for Ref0 and Ref1 as τ0 = POC(current) - POC(Ref0) and τ1 = POC(Ref1) - POC(current). If both predictions are from the same temporal direction (either both from the past or both from the future), the signs are different (i.e., τ0 · τ1 < 0). In this case, BIO is only applied if the predictions are not from the same point in time (i.e., τ0 ≠ τ1), both reference regions have non-zero motion (MVx0, MVy0, MVx1, MVy1 ≠ 0), and the block motion vectors are proportional to the temporal distance (MVx0 / MVx1 = MVy0 / MVy1 = -τ0 / τ1).

[0152] Motion vector field (v x ,v y ) is determined by minimizing the difference Δ in values ​​between point A (which is the intersection of the motion trajectory and the reference frame plane in Figure 9) and point B. The model is a local Taylor expansion with respect to Δ:

number

[0153] All values ​​in equation (9) depend on the sample position (i',j'), which we have omitted from the notation so far. Assuming that the motion is consistent within the local surrounding area, we minimize Δ within a (2M+1) × (2M+1) square window Ω centered at the current prediction point (i,j), where M is equal to 2:

number

[0154] For this optimization problem, JEM uses a simplified approach that minimizes first vertically and then horizontally:

number

number

[0155] To avoid division by zero or very small values, regularization parameters r and m are introduced in equations (11) and (12).

number

[0156] where d is the bit depth of the video samples.

[0157] To keep the memory access for BIO the same as for regular bi-predictive motion compensation, all prediction and gradient values ​​I (k) , ∂I (k) / ∂x, ∂I (k)In equation (13), a (2M+1) × (2M+1) square window Ω centered on the current prediction point on the boundary of the prediction block needs to access positions outside the block (as shown in Figure 22A). In JEM, I outside the block (k) , ∂I (k) / ∂x, ∂I (k) The value of / ∂y is set equal to the nearest available value inside the block. For example, this can be implemented as padding, as shown in Figure 22B.

[0158] Using BIO, it is possible to refine the motion field for each sample. To reduce the computational complexity, JEM uses a block-based design of BIO. The motion refinement is calculated based on a 4x4 block. In block-based BIO, s in equation (13) for all samples in a 4x4 block is used. n The values ​​of are aggregated, and s n The aggregate value of is used to derive the BIO motion vector offset for that 4x4 block. More specifically, for block-based BIO derivation, the following equation:

number

[0159] In some cases, the MV refinement of BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the magnitude of the MV refinement is clipped to a threshold thBIO. This threshold is determined based on whether the reference pictures of the current picture are all from one direction. If all the reference pictures of the current picture are from one direction, the value of the threshold is 12×2.14-d otherwise it is set to 12 x 2 13-d is set to

[0160] The gradients for BIOs are calculated simultaneously with the motion-compensated interpolation, using arithmetic consistent with the HEVC motion compensation process (2D separable FIR). The inputs of this 2D separable FIR are the same reference frame samples as the motion compensation process and the fractional position (fracX, fracY) according to the fractional part of the block motion vector. For horizontal gradients ∂I / ∂x, the signal is first vertically interpolated using BIOfilterS corresponding to fractional position fracY with a descaling shift of d-8, and then BIOfilterG corresponding to fractional position fracX with a descaling shift of 18-d is applied horizontally. For vertical gradients ∂I / ∂y, the gradient filter is first vertically applied using BIOfilterG corresponding to fractional position fracY with a descaling shift of d-8, and then the signal is displaced horizontally using BIOfilterS corresponding to fractional position fracX with a descaling shift of 18-d. The lengths of the interpolation filters for gradient calculation BIOfilterG and signal displacement BIOfilterF are made shorter (6 taps) to maintain reasonable complexity. Table 1 shows the filters used for gradient calculation for different fractional positions of block motion vectors in BIO.

[0161] Table 2 shows the interpolation filters used to generate the predicted signal in BIO. [Table 1] [Table 2]

[0162] In JEM, BIO is applied to all bi-predictive blocks when the two predictions are from different reference pictures. If LIC is enabled for a CU, BIO is disabled.

[0163] In JEM, OBMC is applied to a block after the normal MC process. To reduce the computational complexity, BIO is not applied during the OBMC process. This means that BIO is only applied in the MC process for a block when it uses its own MV, but not in the MC process when the MV of a neighboring block is used in the OBMC process.

[0164] 2.2.9.2 BIO in VTM-3.0 proposed in JVET-L0256 Step 1: Determine if BIO is applicable (W and H are the width and height of the current block).

[0165] BIO is not applicable if: Affine coded ATMVP encoded · (iPOC-iPOC0)*(iPOC-iPOC1)>=0 H==4 or (W==4 and H==8) Use weighted prediction · GBi weight is not (1,1).

[0166] BIO is not used if: - The total SAD between two reference blocks (denoted as R0 and R1) is less than a threshold.

number

[0167] Step 2: Data Preparation For a WxH block, (W+2)x(H+2) samples are interpolated.

[0168] The inner WxH samples are interpolated with an 8-tap interpolation filter as in normal motion compensation.

[0169] The four outer lines of the sample (black circles in Figure 23) are interpolated using a bilinear filter.

[0170] For each position, the gradient is calculated on two reference blocks (denoted R0 and R1): Gx0(x,y)=(R0(x+1,y)-R0(x-1,y))>>4 Gy0(x,y)=(R0(x,y+1)-R0(x,y-1))>>4 Gx1(x,y)=(R1(x+1,y)-R1(x-1,y))>>4 Gy1(x,y)=(R1(x,y+1)-R1(x,y-1))>>4.

[0171] For each position, the internal value is T1=(R0(x,y)>>6)-(R1(x,y)>>6),T2=(Gx0(x,y)+Gx1(x,y))>>3,T3=(Gy0(x,y)+Gy1(x,y))>>3 B1(x,y)=T2*T2,B2(x,y)=T2*T3,B3(x,y)=-T1*T2,B5(x,y)=T3*T3,B6(x,y)=-T1*T3 It is calculated as:

[0172] Step 3: Calculate predictions for each block If the SAD between two 4x4 reference blocks is less than a threshold, the BIO is skipped for the 4x4 block.

[0173] Calculate Vx and Vy.

[0174] Calculate the final prediction for each location in the 4x4 block: b(x,y)=(Vx(Gx 0 (x,y)-Gx 1 (x,y))+Vy(Gy 0 (x,y)-Gy 1 (x,y))+1)>>1 P(x,y)=(R 0 (x,y)+R 1 (x,y)+b(x,y)+offset)>>shift.

[0175] b(x,y) is known as the correction term.

[0176] 2.2.9.3 BIO in VTM-3.0 The section numbers below refer to sections in the latest version of the VVC standard document.

[0177] 8.3.4 Inter-Block Decoding Process - If predFlagL0 and predFlagL1 are equal to 1, DiffPicOrderCnt(currentPic, refPicList0[refIdx0])*DiffPicOrderCnt(currPic, refPicList1[refIdx1])<0, MotionModelIdc[xCb][yCb] is equal to 0, and MergeModeList[merge_idx[xCb]][yCb]] is not equal to SbCol, set the value of bioAvailableg to TRUE. - Otherwise, set the value of bioAvailableFlag to FALSE. ...(original specification text continues) - If bioAvailableFlag is equal to TRUE, the following applies: - The variable Shift is set equal to Max(2,14-bitDepth). The variables cuLevelAbsDiffThres and subCuLevelAbsDiffThres are set equal to (1<<(bitDepth-8+shift))*cbWidth*cbHeight and 1<<(bitDepth-3+shift). The variable cuLevelSumAbsoluteDiff is set to 0. - for xSbIdx=0..(cbWidth>>2)-1 and ySbIdx=0..(cbHeight>>2)-1, the variables subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx] and the bidirectional optical flow utilization flag of the current sub-block, bioUtilizationFlag[xSbIdx][ySbIdx], are derived as follows: i,j=0..3, subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx]=Σ i Σ j Abs(predSamplesL0L[(xSbIdx<<2)+1+i][(ySbIdx<<2)+1+j]- predSamplesL1L[(xSbIdx<<2)+1+i][(ySbIdx<<2)+1+j]) bioUtilizationFlag[xSbIdx][ySbIdx]= subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx]>=subCuLevelAbsDiffThres cuLevelSumAbsoluteDiff+=subCuLevelSumAbsoluteDiff[xSbIdx][ySbIdx] - If cuLevelSumAbsoluteDiff is less than cuLevelAbsDiffThres, set bioAvailableFlag to FALSE. - if bioAvailableFlag is equal to TRUE, the prediction samples within the current luma coding sub-block, predSamplesL[xL+xSb][yL+ySb], with xL=0..sbWidth-1 and yL=0..sbHeight-1, are derived by invoking the bidirectional optical flow sample prediction process specified in clause 8.3.4.5 using the luma coding sub-block width sbWidth, luma coding sub-block height sbHeight, sample arrays preSamplesL0L and preSamplesL1L, and variables predFlagL0, predFlagL1, refIdxL0, refIdxL1.

[0178] 8.3.4.3 Fractional Sample Interpolation Process 8.3.4.3.1 General The inputs to this process are: - a luma position (xSb, ySb) that defines the top left sample of the currently coded sub-block relative to the top left luma sample of the current picture, - a variable sbWidth that specifies the width of the currently coded sub-block in luma samples, - a variable sbHeight that specifies the height of the currently coded sub-block in luma samples, - luma motion vector mvLX given in 1 / 16 luma sample units, - chroma motion vector mvCLX given in 1 / 32 chroma sample units, - the selected reference picture sample array refPicLXL and the arrays refPicLXCb and refPicLXCr, - Bidirectional optical flow enable flag bioAvailableFlag.

[0179] The output of this process is: - predSamplesLXL, a (sbWidth) x (sbHeight) array of predicted luma sample values ​​if bioAvailableFlag is FALSE, or predSamplesLXL, a (sbWidth+2) x (sbHeight+2) array of predicted luma sample values ​​if bioAvailableFlag is TRUE; - Two (sbWidth / 2) x (sbHeight / 2) arrays of predicted chroma sample values: preSamplesLXCb and preSamplesLXCr.

[0180] Let (xIntL, yIntL) be the luma position given in full samples, and (xFracL, yFracL) be the offset given in 1 / 16 samples. These variables are only used in this section to specify the position of fractional samples inside the reference sample arrays refPicLXL, refPicLXCb, and refPicLXCr.

[0181] If bioAvailableFlag is equal to TRUE, then for each luma sample position (xL=-1..sbWidth, yL=-1..sbHeight) inside the predicted luma sample array preSamplesLXL, the corresponding predicted luma sample value preSamplesLXL[xL][yL] is derived as follows: The variables xIntL, yIntL, xFracL and yFracL are derived as follows: xIntL=xSb-1+(mvLX[0]>>4)+xL yIntL=ySb-1+(mvLX[1]>>4)+yL xFracL=mvLX[0]&15 yFracL=mvLX[1]&15 - The value of bilinearFiltEnabledFlag is derived as follows: - If xL is equal to -1 or sbWidth, or if yL is equal to -1 or sbHeight, set the value of bilinearFiltEnabledFlag to TRUE. - Otherwise, set the value of bilinearFiltEnabledFlag to FALSE. The predicted luma sample values ​​predSamplesLXL[xL][yL] are derived by invoking the process specified in Section 8.3.4.3.2 using (xIntL, yIntL), (xFracL, yFracL), refPicLXL, and bilinearFiltEnabledFlag as input.

[0182] If bioAvailableFlag is equal to FALSE, then for each luma sample position (xL=0..sbWidth-1, yL=0..sbHeight-1) inside the predicted luma sample array preSamplesLXL, the corresponding predicted luma sample value preSamplesLXL[xL][yL] is derived as follows: The variables xIntL, yIntL, xFracL and yFracL are derived as follows: xIntL=xSb+(mvLX[0]>>4)+xL yIntL=ySb+(mvLX[1]>>4)+yL xFracL=mvLX[0]&15 yFracL=mvLX[1]&15 - The variable bilinearFiltEnabledFlag is set to FALSE. The predicted luma sample values ​​predSamplesLXL[xL][yL] are derived by invoking the process specified in Section 8.3.4.3.2 using (xIntL, yIntL), (xFracL, yFracL), refPicLXL, and bilinearFiltEnabledFlag as input. ...(original specification text continues) 8.3.4.5 Bidirectional Optical Flow Prediction Process The inputs to this process are: - two variables nCbW and nCbH that define the width and height of the current coding block, - two (nCbW+2) x (nCbH+2) luma prediction sample arrays predSamplesL0 and predSamplesL1, - prediction list usage flags predFlagL0 and predFlagL1, - reference indices refIdxL0 and refIdxL1, - Bidirectional optical flow utilization flag bioUtilizationFlag[xSbIdx][ySbIdx], where xSbIdx=0..(nCbW>2)-1, ySbIdx=0...(nCbH>2)-1.

[0183] The output of this process is a (nCbW) x (nCbH) array of luma prediction sample values, pbSamples.

[0184] The variable bitDepth is set equal to BitDepthY.

[0185] The variable shift2 is set equal to Max(3,15-bitDepth) and the variable offset2 is set equal to 1<<(shift2-1).

[0186] The variable mvRefinesThres is set equal to 1<<(13-bitDepth).

[0187] For xSbIdx=0..(nCbW>>2)-1 and ySbIdx=0..(nCbH>>2)-1, - If bioUtilizationFlag[xSbIdx][ySbIdx] is FALSE, for x=xSb..xSb+3, y=ySb..ySb+3, the predicted sample values ​​of the current prediction unit are derived as follows: pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1, (predSamplesL0[x][y]+predSamplesL1[x][y]+offset2)>>shift2) - Otherwise, the predicted sample values ​​for the current prediction unit are derived as follows: The position (xSb, ySb) defining the top left sample of the current sub-block relative to the top left samples of the prediction sample arrays predSamplesL0 and predSamplesL1 is derived as follows: xSb=(xSbIdx<<2)+1 ySb=(ySbIdx<<2)+1 - For x=xSb-1..xSb+4, y=ySb-1..ySb+4, the following applies: The position (hx,vy) for each corresponding sample (x,y) inside the prediction sample array is derived as follows: hx=Clip3(1,nCbW,x) vy=Clip3(1,nCbH,y) The variables gradientHL0[x][y], gradientVL0[x][y], gradientHL1[x][y] and gradientVL1[x][y] are derived as follows: gradientHL0[x][y]=(predSamplesL0[hx+1][vy]-predSampleL0[hx-1][vy])>>4 gradientVL0[x][y]=(predSampleL0[hx][vy+1]-predSampleL0[hx][vy-1])>>4 gradientHL1[x][y]=(predSamplesL1[hx+1][vy]-predSampleL1[hx-1][vy])>>4 gradientVL1[x][y]=(predSampleL1[hx][vy+1]-predSampleL1[hx][vy1])>>4 - The temporary temp, tempX, tempY, and other characters are as follows: temp[x][y]=(predSamplesL0[hx][vy]>>6)-(predSamplesL1[hx][vy]>>6) tempX[x][y]=(gradientHL0[x][y]+gradientHL1[x][y])>>3 tempY[x][y]=(gradientVL0[x][y]+gradientVL1[x][y])>>3 sGx2, sGy2, sGxGy, sGxdI, and sGydI. x,y=-1..4 and sGx2=Σ x Σ y (tempX[xSb+x][ySb+y]*tempX[xSb+x][ySb+y]) x,y=-1..4 and sGy2=Σ x Σ y (tempY[xSb+x][ySb+y]*tempY[xSb+x][ySb+y]) x,y=-1..4 and sGxGy=Σ x Σ(tempX[xSb+x][ySb+y]*tempY[xSb+x][ySb+y]) x,y=-1..4、sGxdIΣ x Σ(-tempX[xSb+x][ySb+y]*temp[xSb+x][ySb+y]) x,y=-1..4, sGydI=Σ x Σ y (-tempY[xSb+x][ySb+y]*temp[xSb+x][ySb+y]) - The horizontal and vertical motion refinement of the current sub-block is derived as follows: vx=sGx2>0?Clip3(-mvRefineThres,mvRefineThres,-(sGxdI<<3)>>Floor(Log2(sGx2))):0 vy=sGy2>0?Clip3(-mvRefineThres,mvRefineThres,((sGydI<<3)-((vx*sGxGym)<<12+vx*sGxGys)>>1)>>Floor(Log2(sGy2))):0 sGxGym=sGxGy>>12; sGxGys=sGxGy&((1<<12)-1)

[0188] For x=xSb-1..xSb+2, y=ySb-1..ySb+2, the following applies: sampleEnh=Round((vx*(gradientHL1[x+1][y+1]-gradientHL0[x+1][y+1]))>>1) +Round((vy*(gradientVL1[x+1][y+1]-gradientVL0[x+1][y+1]))>>1) pbSamples[x][y]=Clip3(0,(1< <bitDepth)-1,(predSamplesL0[x+1][y+1] +predSamplesL1[x+1][y+1]+sampleEnh+offset2)>>shift2)

[0189] 2.2.10 Decoder-Side Motion Vector Refinement (DMVR) DMVR is a type of decoder-side motion vector derivation (DMVD).

[0190] In bi-prediction operation, for prediction of one block region, two prediction blocks formed using the motion vector (MV) of list0 and the MV of list1 are combined to form a single prediction signal. In decoder-side motion vector refinement (DMVR) method, the two bi-predictive motion vectors are further refined by a bilateral template matching process. To obtain the refined MV without transmitting additional motion information, bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed sample in the reference picture.

[0191] In DMVR, a bilateral template is generated as a weighted combination (i.e., average) of two prediction blocks from the original MV0 in list0 and MV1 in list1, as shown in Figure 24. The template matching operation consists of calculating a cost metric between the generated template and a sample region (near the original prediction block) in the reference picture. For each of the two reference pictures, the MV that produces the smallest template cost is considered as the updated MV of that list to replace the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and eight surrounding MVs that are offset by one luma sample horizontally or vertically or both relative to the original MV. Finally, these two new MVs, i.e., MV0' and MV1' shown in Figure 24, are used to generate the final bi-prediction result. The sum of absolute differences (SAD) is used as the cost metric. It should be noted that when calculating the cost of a predicted block generated by one surrounding MV, the rounded MV (to integer pels) is actually used instead of the actual MV to obtain the predicted block.

[0192] DMVR is applied to bi-predictive merge modes that use one MV from a past reference picture and another MV from a future reference picture without transmitting additional syntax elements. In JEM, DMVR is not applied when LIC, affine motion, FRUC, or sub-CU merge candidates are enabled for a CU.

[0193] 2.2.11 JVET-N0236 This contribution proposes a method to refine subblock-based affine motion compensation prediction using optical flow. After subblock-based affine motion compensation is performed, the predicted samples are refined by adding the difference derived by the optical flow equation, which is referred to as prediction refinement with optical flow (PROF). The proposed method can achieve inter prediction at pixel-level granularity without increasing memory access bandwidth.

[0194] To achieve finer-granularity motion compensation, this contribution proposes a method to refine sub-block-based affine motion compensation prediction using optical flow. After sub-block-based affine motion compensation is performed, the luma prediction samples are refined by adding the difference derived by the optical flow equation. The proposed prediction refinement with optical flow (PROF) can be described as the following four steps:

[0195] Step 1) Perform sub-block based affine motion compensation to generate the sub-block prediction I(i,j).

[0196] Step 2) At each sample position, we use a 3-tap filter [-1,0,1] to calculate the spatial gradient of the subblock prediction g x (i,j) and g y Calculate (i,j):

number

[0197] The subblock prediction is extended by one pixel on each side for gradient calculation. To reduce memory bandwidth and complexity, the pixels on the extended boundary are copied from the nearest integer pixel position in the reference picture. Thus, additional interpolation for the padding area is avoided.

[0198] Step 3) Luma prediction refinement using optical flow equations:

number

[0199] Since the affine model parameters and pixel position relative to the sub-block center do not change between sub-blocks, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks within the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the center of the sub-block, respectively, and Δv(x,y) can be calculated as follows:

number

[0200] In the four-parameter affine model,

number

[0201] In the six-parameter affine model,

number

[0202] Step 4) Finally, add the luma prediction refinement to the sub-block prediction I(i,j). The final prediction I' is given by:

number

[0203] 2.2.12 JVET-N0510 Phase-variant affine subblock motion compensation (PAMC) To better approximate the affine motion model in affine sub-blocks, phase variational MC is applied to the sub-blocks. In the proposed method, the affine coding block is also divided into 4x4 sub-blocks, and a sub-block MV is derived for each sub-block as done in VTM4.0. The MC for each sub-block is divided into two stages. The first stage filters a (4+L-1)x(4+L-1) reference block window with (4+L-1) rows of horizontal filtering, where L is the filter tap length of the interpolation filter. However, unlike translational MC, in the proposed phase variational affine sub-block MC, the filter phase for each sample row is different. For each sample row, the MVx is derived as follows: MVx=(subblockMVx<<7+dMvVerX×(rowIdx-L / 2-2))>>7 (Formula 1)

[0204] The filter phase for each sample row is derived from MVx. subblockMVx is the x-component of MV for subblock MV, derived as in VTM4.0. rowIdx is the sample row index. dMvVerX is (cuBottomLeftCPMVx-cuTopLeftCPMVx)<<(7-log2LumaCbHeight), where cuBottomLeftCPMVx is the x-component of the CU bottom-left control point MV, cuTopLeftCPMVx is the x-component of the CU top-left control point MV, and LumaCbHeight is the log2 of the height of the luma coding block (CB).

[0205] After horizontal filtering, 4×(4+L-1) horizontally filtered samples are generated. Figure 1 shows the proposed horizontal filtering concept. The gray dots are the samples of the reference block window, and the orange dots represent the horizontally filtered samples. The blue tubular 8×1 samples represent applying one 8-tap horizontal filtering, as shown in Figures 30 and 31. Each sample row requires four horizontal filterings. The filter phase of one sample row is the same. However, the filter phases of different rows are different. 4×11 skewed samples are generated.

[0206] In the second stage, the 4 × (4 + L - 1) horizontally filtered samples (orange samples in Figure 1) are further vertically filtered. For each sample column, MVy is derived as follows: MVy=(subblockMVy<<7+dMvHorY×(columnIdx-2))>>7 (Formula 2)

[0207] The filter phase for each sample row is derived from MVy. subblockMVy is the y-component of MV for subblock MV, derived as in VTM4.0. columnIdx is the sample column index. dMvHorY is (cuTopRightCPMVy-cuTopLeftCPMVy)<<(7-log2LumaCbWidth), where cuTopRightCPMVy is the y-component of the CU top-right control point MV, cuTopLeftCPMVy is the y-component of the CU top-left control point MV, and log2LumaCbWidth is the log2 of the luma CB width.

[0208] After vertical filtering, 4x4 affine sub-block prediction samples are generated. Figure 32 shows the proposed vertical filtering concept. The light orange dots are horizontally filtered samples from the first stage. The red dots are vertically filtered samples as the final prediction samples.

[0209] In this proposal, the interpolation filter set used is the same as that in VTM4.0. The only difference is that the horizontal filter phase for one sample row is different and the vertical filter phase for one sample column is different. The number of filtering operations for each affine sub-block in the proposed method is the same as that in VTM4.0.

[0210] 3. Examples of problems solved by the disclosed technical solutions 1. BIO only considers bi-prediction 2. The derivation of vx and vy does not take into account the motion information of neighboring blocks. 3. JVET-N0236 uses optical flow for affine prediction, which is far from optimal.

[0211] 4. Examples of embodiments and techniques To address these issues, we propose a different configuration for deriving refinement prediction samples using optical flow. In addition, we propose using information (e.g., reconstructed samples or motion information) of neighboring (e.g., adjacent or non-adjacent) blocks and / or gradients of one sub-block / block and its predicted block to obtain the final predicted block of the current sub-block / block.

[0212] The techniques and embodiments listed below should be considered as examples to illustrate the general concepts. These embodiments should not be construed in a narrow sense. Furthermore, these techniques can be combined in any suitable manner.

[0213] The reference pictures of the current picture from list 0 and list 1 are denoted by Ref0 and Ref1, respectively, and τ0 = POC(current) - POC(Ref0) and τ1 = POC(Ref1) - POC(current). Also, the reference blocks of the current block from Ref0 and Ref1 are denoted by refblk0 and Refblk1, respectively. For a sub-block in the current block, the MV of the corresponding reference sub-block of refblk0 that points to refblk1 is (v x ,v y ) The MVs of the current sub-block that refer to Ref0 and Ref1 are expressed as (mvL0 x ,mvL0 y ) and (mvL1 x ,mvL1 y ) is written as

[0214] In the following description, SatShift(x,n) is

number

[0215] In one example, offset0 and / or offset1 are (1<<n)> >1 or (1<<(n-1). In another example, offset0 and / or offset1 are set to 0.

[0216] In another example, offset0=offset1=((1<<n)> >1)-1 or ((1<<(n-1)))-1.

[0217] Clip3(min,max,x)

number

[0218] In the following description, an operation between two motion vectors means that the operation is applied to both components of the motion vector. For example, MV3 = MV1 + MV2 means that MV3 x =MV1 x +MV2 x and MV3 y =MV1 y +MV2 y Alternatively, the operation may be applied to only the horizontal or vertical components of the two motion vectors.

[0219] In the following description, the left adjacent block, the lower left adjacent block, the upper adjacent block, the upper right adjacent block, and the upper left adjacent block are represented as blocks A1, A0, B1, B0, and B2 shown in FIG. 1. It is proposed that the predicted sample P(x,y) at position (x,y) within a block can be refined as P'(x,y) = P(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y). P'(x,y) is used together with the residual sample Res(x,y) to generate the reconstructed Rec(x,y). (Gx(x,y), Gy(x,y)) represent the gradient at position (x,y), e.g., along the horizontal and vertical directions, respectively. (Vx(x,y), Vy(x,y)) represent the motion displacement at position (x,y), which can be derived on the fly; Alternatively, a weighting function can be applied to the prediction samples, gradients, and motion displacements, for example, P'(x,y) = α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), where α(x,y), β(x,y), and γ(x,y) are weight values ​​at position (x,y), which can be integers or real numbers; i. For example, P'(x,y) = (α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y) + offsetP) / (α(x,y) + β(x,y) + γ(x,y)). In one example, offsetP is set to 0. Alternatively, the division may be replaced by a shift; ii. For example, P'(x,y)=P(x,y)-Gx(x,y)×Vx(x,y)+Gy(x,y)×Vy(x,y); iii. For example, P'(x,y)=P(x,y)-Gx(x,y)×Vx(x,y)-Gy(x,y)×Vy(x,y); iv. For example, P'(x,y)=P(x,y)+Gx(x,y)×Vx(x,y)-Gy(x,y)×Vy(x,y); v. For example, P'(x,y)=0.5×P(x,y)+0.25×Gx(x,y)×Vx(x,y)+0.25×Gy(x,y)×Vy(x,y); vi. For example, P'(x,y)=0.5×P(x,y)+0.5×Gx(x,y)×Vx(x,y)+0.5×Gy(x,y)×Vy(x,y); vii. For example, P'(x,y)=P(x,y)+0.5×Gx(x,y)×Vx(x,y)+0.5×Gy(x,y)×Vy(x,y); b. Alternatively, P'(x,y) = Shift(α(x,y) × P(x,y),n1) + Shift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + Shift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where α(x,y), β(x,y), and γ(x,y) are weight values ​​at position (x,y) and are integers. n1, n2, and n3 are non-negative integers, e.g., 1; c. Alternatively, P'(x,y) = SatShift(α(x,y) × P(x,y),n1) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + SatShift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where α(x,y), β(x,y), and γ(x,y) are weight values ​​at position (x,y) and are integers. n1, n2, and n3 are non-negative integers, e.g., 1; d. Alternatively, P'(x,y) = Shift(α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n1), where α(x,y), β(x,y), and γ(x,y) are weight values ​​at position (x,y) and are integers. n1 is a non-negative integer, e.g., 1; e. Alternatively, P'(x,y) = SatShift(α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n1), where α(x,y), β(x,y) and γ(x,y) are weight values ​​at position (x,y) and are integers. n1 is a non-negative integer, e.g., 1; f. Alternatively, P'(x,y) = α(x,y) × P(x,y) + Shift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + Shift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where α(x,y), β(x,y), and γ(x,y) are integer weights at position (x,y). n2, n3 are non-negative integers, e.g., 1; g. Alternatively, P'(x,y) = α(x,y) × P(x,y) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + SatShift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where α(x,y), β(x,y), and γ(x,y) are integer weights at position (x,y). n2 and n3 are non-negative integers, e.g., 1; h. Alternatively, P'(x,y) = α(x,y) × P(x,y) + Shift(β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n3), where α(x,y), β(x,y), and γ(x,y) are integer weights at position (x,y). n3 is a non-negative integer, e.g., 1; i. Alternatively, P'(x,y) = α(x,y) × P(x,y) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n3), where α(x,y), β(x,y), and γ(x,y) are integer weights at position (x,y). n3 is a non-negative integer, e.g., 1; j. Or, P'(x,y)=f0(P(x,y))+f1(Gx(x,y)×Vx(x,y))+f2(Gy(x,y)×Vy(x,y)), where f0, f1 and f2 are three functions; k. In one example, (Gx(x,y), Gy(x,y)) is calculated by P(x1,y1), where x1 is in the range [x-Bx0,x+Bx1], y1 is in the range [y-By0,y+By1], and Bx0, Bx1, By0, By1 are integers; i. For example, Gx(x,y)=P(x+1,y)-P(x-1,y), Gy(x,y)=P(x,y+1)-P(x,y-1); (i) Alternatively, Gx(x,y) = Shift(P(x+1,y) - P(x-1,y),n1), Gy(x,y) = Shift(P(x,y+1) - P(x,y-1),n2), for example, n1 = n2 = 1; (ii) Alternatively, Gx(x,y) = SatShift(P(x+1,y) - P(x-1,y), n1), Gy(x,y) = SatShift(P(x,y+1) - P(x,y-1), n2), e.g., n1 = n2 = 1; l. P(x,y) can be the prediction value of one-sided prediction (inter prediction with one MV); m. P(x,y) can be the final predicted value after bi-prediction (inter-prediction with two MVs); i. For example, Vx(x,y) and Vy(x,y) may be derived according to the method specified in BIO (also known as bidirectional optical flow, BDOF); n. P(x,y) may be a multi-hypothesis inter-prediction (inter-prediction with three or more MVs); o. P(x,y) may be an affine prediction; p. P(x,y) may be an intra prediction; q. P(x,y) may be an intra-block copy (IBC) prediction; r. P(x,y) may be generated by triangular prediction mode (TPM) or geographic prediction mode (GPM) technique; s. P(x,y) may be an inter-intra joint prediction; t. P(x,y) may be a global inter prediction, where the regions share the same motion model and parameters; u. P(x,y) may be generated by palette coding mode; v. P(x,y) may be inter-view prediction in multiview or 3D video coding; w. P(x,y) may be inter-layer prediction in scalable video coding; x. P(x,y) may be filtered before being refined; y. P(x,y) may be the final prediction to be summed with the residual sample values ​​to obtain the reconstructed sample values. In some embodiments, P(x,y) may be the final prediction when no refinement process is applied. In some embodiments, P'(x,y) may be the final prediction when a refinement process is applied; i. In one example, for a block (or sub-block) where bi-prediction or multi-hypothesis prediction is applied, the above function is applied once to the final predicted value; ii. In one example, in a block (or sub-block) to which bi-prediction or multi-hypothesis prediction is applied, for each prediction block according to one prediction direction or reference picture or motion vector, the above function is applied multiple times, so that the above process is called to update the prediction block. Then, the updated prediction block can be used to generate a final prediction block; iii. Alternatively, P(x,y) may be an intermediate prediction that will be used to derive a final prediction; (i) For example, P(x,y) may be a prediction from one reference picture list when the current block is bi-predictive and inter-predicted; (ii) For example, P(x,y) may be a prediction from one reference picture list when the current block is inter-predicted with TPM or GPM techniques; (iii) For example, P(x,y) may be a prediction from one reference picture when the current block is inter-predicted with multiple hypotheses; (iv) For example, P(x,y) may be inter-predicted if the current block is jointly inter-intra predicted; (v) For example, P(x,y) may be the inter prediction before local illumination compensation (LIC) is applied if the current block uses LIC; (vi) For example, P(x,y) may be an inter prediction before DMVR (or other type of DMVD) is applied if the current block uses DMVR (or other type of DMVD); (vii) For example, P(x,y) may be the inter prediction before being multiplied by the weighting factor if the current block uses weighted prediction or generalized bi-prediction (GBi); z. A gradient, denoted G(x,y), e.g., Gx(x,y) or / and Gy(x,y), may be derived on the final prediction to be summed with the residual sample values ​​to obtain the reconstructed sample values. In some embodiments, if no refinement process is applied, the final predicted sample values ​​are added to the residual sample values ​​to obtain the reconstructed sample values; i. For example, G(x,y) can be derived on P(x,y); ii. Alternatively, G(x,y) can be derived on intermediate predictions that will be used to derive the final prediction; (i) For example, G(x,y) may be derived on prediction from one reference picture list when the current block is bi-predictively inter-predicted; (ii) For example, G(x,y) can be derived based on prediction from one reference picture list when the current block is inter-predicted with TPM or GPM techniques; (iii) For example, G(x,y) can be derived on prediction from one reference picture when the current block is inter-predicted with multiple hypotheses; (iv) For example, G(x,y) can be derived on inter prediction when the current block is inter-intra jointly predicted; (v) For example, if the current block uses local illumination compensation (LIC), G(x,y) may be derived on the inter prediction before LIC is applied; (vi) For example, if the current block uses DMVR (or other type of DMVD), G(x,y) may be derived on inter prediction before DMVR (or other type of DMVD) is applied; (vii) For example, G(x,y) may be derived on the inter prediction before being multiplied by the weighting factor if the current block uses weighted prediction or generalized bi-prediction (GBi); Alternatively, P'(x,y) may be further processed by other methods to obtain the final predicted sample; bb. Alternatively, a reconstructed sample Rec(x,y) at position (x,y) within a block can be refined as Rec'(x,y) = Rec(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y). Rec'(x,y) will be used to replace the reconstructed Rec(x,y). (Gx(x,y), Gy(x,y)) represent the gradient at position (x,y), e.g., along the horizontal and vertical directions, respectively. (Vx(x,y), Vy(x,y)) represent the motion displacement at position (x,y), which can be derived on the fly; i. In one example, (Gx(x,y), Gy(x,y)) are derived on the reconstruction samples. 2. It is proposed that Vx(x,y) and / or Vy(x,y) used in optical flow based methods (such as in bullet 1) may depend on spatial or temporal neighboring blocks; Alternatively, Vx(x,y) and / or Vy(x,y) in the process of BIO (also known as BDOF) may depend on spatial or temporal neighboring blocks; b. In one example, the "dependence" on spatial or temporal neighboring blocks may include dependence on motion information (e.g., MV), coding mode (e.g., inter-coding or intra-coding), neighboring CU size, neighboring CU position, etc.; c. In one example, (Vx(x,y),Vy(x,y)) is the MV Mix can be equal to; i. In one example, MV Mix is Wc(x,y)×MVc+W N1 (x,y)×MV N1 +W N2 (x,y)×MV N2 +...+W Nk (x,y)×MV Nk where MVc is the MV of the current block, and MV N1 ...MV Nk are the MVs of k spatially or temporally adjacent blocks: N1...Nk. Wc, W N1 ...W Nk is the weight value, which can be an integer or a real number; ii. Or MV Mix is Shift(Wc(x,y)×MVc+W N1 (x,y)×MV N1 +W N2 (x,y)×MV N2 +...+W Nk (x,y)×MV Nk , n1), where MVc is the MV of the current block, and MV N1 ...MV Nk are the MVs of k spatially or temporally adjacent blocks: N1...Nk. Wc, W N1 ...W Nk is a weight value that is an integer; iii. Or MV Mix is SatShift(Wc(x,y)×MVc+W N1 (x,y)×MV N1 +W N2 (x,y)×MV N2 +...+W Nk (x,y)×MV Nk , n1), where MVc is the MV of the current block, and MV N1 ...MV Nk are the MVs of k spatially or temporally adjacent blocks: N1...Nk. Wc, W N1 ...W Nk is a weight value that is an integer; iv. In one example, Wc(x,y)=0; v. In one example, k=1 and N1 is a spatially adjacent block; (i) In one example, N1 is the closest spatial neighboring block to location (x, y); (ii) In one example, W N1 (x,y) is larger the closer the position (x,y) is to N1; vi. In one example, k=1, and N1 is a time-adjacent block; (i) For example, N1 is a collocated block in a collocated picture for position (x, y); In one example, different locations may use different spatial or temporal neighboring blocks; vii. Figure 26 shows an example of how to derive Vx(x,y) and / or Vy(x,y). In the figure, each block represents a basic block (e.g., a 4x4 block). The current block is marked with a bold line; (i) Prediction samples within shaded basic blocks, e.g., those that are not at the top or left boundary, are not refined in the optical flow; (ii) The predicted samples in the basic blocks at the upper boundary, e.g., C00, C10, C20 and C30, will be refined with optical flow; a. For example, the MV for the predicted samples in the basic block at the upper boundary Mix is derived depending on the adjacent upper neighboring blocks. For example, MV for the predicted samples in C10 Mix is derived depending on the upper neighboring block T1; (iii) The predicted samples in the basic blocks at the left boundary, e.g., C00, C01, C02 and C03, will be refined with optical flow; a. For example, the MV for the predicted samples in the basic block at the left boundary Mix is derived depending on the adjacent left neighboring block. For example, MV for the predicted sample in C01 Mix will be derived depending on the left neighboring block L1; viii. In one example, MVc and MV N1 ...MV Nk can be scaled to the same reference picture; (i) In one example, they are scaled to the reference picture that MVc refers to; ix. In one example, MV is generated using spatial or temporal neighboring block Ns only if it is not intra-coded. Mix can be derived; x. In one example, MV is generated using spatial or temporal neighboring block Ns only if it is not IBC coded. Mix can be derived; xi. In one example, MVc is generated using spatial or temporal neighboring block Ns only if MVs references the same reference picture as MVc. Mix can be derived; d. In one example, (Vx(x,y),Vy(x,y)) is f(MV Mix ,MVc), where f is a function and MVc is the MV of the current block; i. For example, (Vx(x,y),Vy(x,y)) is MV Mix -can be equal to MVc; ii. For example, (Vx(x,y),Vy(x,y)) is MVc-MV Mix can be equal to; iii. For example, (Vx(x,y),Vy(x,y)) is p×MV Mix +q×MVc, where p and q are real numbers. Some examples of p and q are p=q=0.5, or p=1, q=−0.5, or q=1, p=−0.5, etc.; (i) Or, (Vx(x,y), Vy(x,y)) is Shift(p×MV Mix +q×MVc,n) or SatShift(p×MV Mix +p×MVc,n), where p, q, and n are integers. Some examples of n, p, and q are n=1, p=2, q=-1, or n=1, p=q=1, or n=1, p=-1, q=2, etc.; e. In one example, the current block is uni-predictive and inter-predicted, and MVc may refer to reference picture list 0; f. In one example, the current block is uni-predictive and inter-predicted, and MVc may refer to reference picture list 1; g. In one example, the current block is bi-predictive and inter-predicted, and MVc may refer to reference picture list 0 or reference picture list 1; i. In one example, the final prediction is refined by optical flow: (Vx(x,y), Vy(x,y)) may be derived with MVc referring to one of the reference picture lists, e.g., reference picture list 0 or reference picture list 1; ii. In one example, the prediction from reference list 0 is refined by optical flow: (Vx(x,y),Vy(x,y)) can be derived with MVc referencing reference list 0; iii. In one example, the prediction from reference list 1 is refined by optical flow: (Vx(x,y),Vy(x,y)) can be derived with MVc referring to reference list 1; iv. The prediction from reference list 0 after being refined by optical flow and the prediction from reference list 1 after being independently refined by optical flow can be combined (e.g., averaged or weighted averaged) to obtain the final prediction; h. In one example, the current block is bi-predictively inter-predicted, BDOF is applied, and Vx(x,y), Vy(x,y) are modified according to spatial or temporal neighboring blocks; i. For example, if (Vx(x,y), Vy(x,y)) derived in the BDOF process is expressed as V'(x,y) = (V'x(x,y), V'y(x,y)), and (Vx(x,y)), Vy(x,y)) derived in the disclosed method are expressed as V"(x,y) = (V"x(x,y), V"y(x,y)), the final V(x,y) = (Vx(x,y), Vy(x,y)) can be derived as follows: (i) For example, V(x,y) = V'(x,y) × W'(x,y) + V"(x,y) × W"(x,y), where W'(x,y) and W"(x,y) are integers or real numbers. For example, W'(x,y) = 0, W"(x,y) = 1, or W'(x,y) = 1, W"(x,y) = 0, or W'(x,y) = 0.5, W"(x,y) = 0.5; (ii) For example, V(x,y) = Shift(V'(x,y) × W'(x,y) + V"(x,y) × W"(x,y), n1), where W'(x,y) and W"(x,y) are integers. n1 is a non-negative integer, for example, 1; (iii) For example, V(x,y) = SatShift(V'(x,y) × W'(x,y) + V"(x,y) × W"(x,y), n1), where W'(x,y) and W"(x,y) are integers. n1 is a non-negative integer, for example, 1; ii. For example, whether to modify (Vx(x,y), Vy(x,y)) in BDOF according to spatial or temporal neighboring blocks may depend on the position (x,y); (i) For example, (Vx(x,y), Vy(x,y)) in the shaded basic blocks in Figure 26, which are not at the top or left boundary, are not modified according to spatial or temporal neighboring blocks; i. Vx(x,y) and / or Vy(x,y) may be clipped; j. Alternatively, the spatially or temporally adjacent blocks in the above method may be replaced by non-adjacent blocks of the current block; k. Alternatively, the spatially or temporally adjacent block in the above method may be replaced by a non-adjacent sub-block of the current sub-block; l. Alternatively, the spatially or temporally adjacent blocks in the above-mentioned methods may be replaced by non-adjacent sub-blocks of the current block / current CTU / current VPDU / current region covering the current sub-block; m. Alternatively, the spatial or temporal neighboring blocks in the above method may be replaced by entries in a history-based motion vector. 3. It is proposed that Vx(x,y) and / or Vy(x,y) in the refinement on affine prediction using optical flow disclosed in JVET-N0236 may be derived as follows: a. Vx(x,y)=a×(x-xc)+b×(y-yc), Vy(x,y)=c×(x-xc)+d×(y-yc), where (x,y) is the position under consideration, (xc,yc) is the center position of a basic block of dimensions w×h (e.g. 4×4 or 8×8) that covers the position (x,y), and a, b, c and d are affine parameters; i. Or, Vx(x,y)=Shift(a×(x-xc)+b×(y-yc),n1), Vy(x,y)=Shift(c×(x-xc)+d×(y-yc),n1), where n1 is an integer; ii. Or, Vx(x,y)=SatShift(a×(x-xc)+b×(y-yc),n1), Vy(x,y)=SatShift(c×(x-xc)+d×(y-yc),n1), where n1 is an integer; iii. For example, if the top left position of the basic block (e.g., 4x4 or 8x8) covering the position (x,y) is (x0,y0), then (xc,yc)=(x0+(w / 2,y0+(h / 2)); (i) Or, (xc, yc) = (x0 + (w / 2) - 1, y0 + (h / 2) - 1); (ii) Or, (xc, yc) = (x0 + (w / 2), y0 + (h / 2) - 1); (iii) Or, (xc, yc) = (x0 + (w / 2) - 1, y0 + (h / 2)); iv. In one example, if the current block applies a 4-parameter affine mode, then c=-b and d=a; v. In one example, a, b, c, and d, along with the width (W) and height (H) of the current block, may be derived from the CPMV. For example,

number

[0220] FIG. 27 is a block diagram of a video processing device 2700. The device 2700 may be used to implement one or more of the methods described herein. The device 2700 may be embodied in a smartphone, tablet, computer, or Internet of Things (IoT) receiver. The device 2700 may include one or more processors 2702, one or more memories 2704, and video processing hardware 2706. The processor(s) 2702 may be configured to execute one or more of the methods described herein. The memory(s) 2704 may be used to store data and code used to execute the methods and techniques described herein. The video processing hardware 2706 may be used to implement some of the techniques described herein in hardware circuitry and may be partially or fully part of the processor 2702 (e.g., a graphics processor core GPU or other signal processing circuitry).

[0221] 28A is a flowchart of an example video processing method. Method 2800A includes determining (2802) a refined prediction sample P′(x,y) at a position (x,y) in the video block by modifying (2802) the prediction sample P(x,y) at the position (x,y) as a function of a gradient in a first direction and / or a second direction estimated at the position (x,y) and a first motion displacement and / or a second motion displacement estimated for the position (x,y), and performing (2804) the conversion using reconstructed sample values ​​Rec(x,y) from the refined prediction sample P′(x,y).

[0222] FIG. 28B is a flowchart of an example video processing method. Method 2800B includes determining (2812) a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers, and performing (2814) a conversion between the video block and a bitstream representation of the video block using reconstructed sample values ​​Rec(x,y) at the position (x,y) obtained based on the refined prediction sample P'(x,y) and residual sample values ​​Res(x,y).

[0223] 28C is a flowchart of an example video processing method. The method 2800C includes determining (2822) a refined prediction sample P′(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers, and encoding (2824) a bitstream representation of the video block to include residual sample values ​​Res(x,y) based on reconstructed sample values ​​Rec(x,y) at the position (x,y) that are based on at least the refined prediction sample P′(x,y).

[0224] In some embodiments of method 2800B and / or 2800C, the first direction and the second direction are orthogonal to each other. In some embodiments of method 2800B and / or 2800C, the first motion displacement represents a direction parallel to the first direction and the second motion displacement represents a direction parallel to the second direction. In some embodiments of method 2800B and / or 2800C, P'(x,y) = P(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y). In some embodiments of methods 2800B and / or 2800C, P'(x,y) = α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), where α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), and α(x,y), β(x,y), and γ(x,y) are integers or real numbers. In some embodiments of method 2800B and / or 2800C, P'(x,y) = (α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y) + offsetP) / (α(x,y) + β(x,y) + γ(x,y)), where α(x,y), β(x,y), and γ(x,y) are weighting values ​​at position (x,y), and α(x,y), β(x,y), and γ(x,y) are integers or real numbers. In some embodiments of method 2800B and / or 2800C, offsetP is equal to 0. In some embodiments of method 2800B and / or 2800C, P'(x,y) is obtained using a binary shift operation.

[0225] In some embodiments of method 2800B and / or 2800C, P'(x,y) = P(x,y) - Gx(x,y) x Vx(x,y) + Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = P(x,y) - Gx(x,y) x Vx(x,y) - Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = P(x,y) + Gx(x,y) x Vx(x,y) - Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = 0.5 x P(x,y) + 0.25 x Gx(x,y) x Vx(x,y) + 0.25 x Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = 0.5 x P(x,y) + 0.5 x Gx(x,y) x Vx(x,y) + 0.5 x Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = P(x,y) + 0.5 x Gx(x,y) x Vx(x,y) + 0.5 x Gy(x,y) x Vy(x,y). In some embodiments of method 2800B and / or 2800C, P'(x,y) = Shift(α(x,y) × P(x,y),n1) + Shift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + Shift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where the Shift() function refers to a binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n1, n2, n3 are non-negative integers. In some embodiments of method 2800B and / or 2800C, n1, n2, n3 are equal to 1.

[0226] In some embodiments of method 2800B and / or 2800C, P'(x,y) = SatShift(α(x,y) × P(x,y),n1) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + SatShift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where the SatShift() function refers to a saturated binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n1, n2, n3 are non-negative integers. In some embodiments of method 2800B and / or 2800C, n1, n2, n3 are equal to 1. In some embodiments of method 2800B and / or 2800C, P'(x,y) = Shift(α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n1), where the Shift() function refers to a binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n1 is a non-negative integer. In some embodiments of method 2800B and / or 2800C, n1 is equal to 1.

[0227] In some embodiments of method 2800B and / or 2800C, P'(x,y) = SatShift(α(x,y) × P(x,y) + β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n1), where the SatShift() function refers to a saturated binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n1 is a non-negative integer. In some embodiments of method 2800B and / or 2800C, n1 is equal to 1. In some embodiments of method 2800B and / or 2800C, P'(x,y) = α(x,y) × P(x,y) + Shift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + Shift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where the Shift() function refers to a binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n2 and n3 are non-negative integers. In some embodiments of method 2800B and / or 2800C, n2 and n3 are equal to 1.

[0228] In some embodiments of method 2800B and / or 2800C, P'(x,y) = α(x,y) × P(x,y) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y),n2) + SatShift(γ(x,y) × Gy(x,y) × Vy(x,y),n3), where the SatShift() function refers to a saturated binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n2 and n3 are non-negative integers. In some embodiments of method 2800B and / or 2800C, n2 and n3 are equal to 1. In some embodiments of method 2800B and / or 2800C, P'(x,y) = α(x,y) × P(x,y) + Shift(β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y),n3), where the Shift() function refers to a binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n3 is a non-negative integer. In some embodiments of method 2800B and / or 2800C, n3 is equal to 1.

[0229] In some embodiments of method 2800B and / or 2800C, P'(x,y) = α(x,y) × P(x,y) + SatShift(β(x,y) × Gx(x,y) × Vx(x,y) + γ(x,y) × Gy(x,y) × Vy(x,y), n3), where the SatShift() function refers to a saturated binary shift operation, α(x,y), β(x,y), and γ(x,y) are weighting values ​​at location (x,y), α(x,y), β(x,y), and γ(x,y) are integers, and n3 is a non-negative integer. In some embodiments of method 2800B and / or 2800C, n3 is equal to 1. In some embodiments of method 2800B and / or 2800C, 1-4, P'(x,y)=f0(P(x,y))+f1(Gx(x,y)×Vx(x,y))+f2(Gy(x,y)×Vy(x,y)), where f0(), f1(), and f2() are three functions. In some embodiments of method 2800B and / or 2800C, the first gradient component Gx(x,y) and the second gradient component Gy(x,y) are calculated using the second predicted sample P(x1,y1), where x1 is in a first range of [x-Bx0,x+Bx1], y1 is in a second range of [y-By0,y+By1], Bx0 and By0 are integers, and Bx1 and By1 are integers.

[0230] In some embodiments of methods 2800B and / or 2800C, the first gradient component Gx(x,y) = P(x+1,y) - P(x-1,y) and the second gradient component Gy(x,y) = P(x,y+1) - P(x,y-1). In some embodiments of methods 2800B and / or 2800C, the first gradient component Gx(x,y) = Shift(P(x+1,y) - P(x-1,y),n1) and the second gradient component Gy(x,y) = Shift(P(x,y+1) - P(x,y-1),n2), where the Shift() function refers to a binary shift operation. In some embodiments of method 2800B and / or 2800C, the first gradient component Gx(x,y) = SatShift(P(x+1,y) - P(x-1,y), n1), and the second gradient component Gy(x,y) = SatShift(P(x,y+1) - P(x,y-1), n2), where the SatShift() function refers to a saturated binary shift operation. In some embodiments of method 2800B and / or 2800C, n1 and n2 are equal to 1. In some embodiments of method 2800B and / or 2800C, the predicted sample P(x,y) is a uni-predicted sample at position (x,y). In some embodiments of method 2800B and / or 2800C, the predicted sample P(x,y) is the final result of bi-prediction.

[0231] In some embodiments of methods 2800B and / or 2800C, Vx(x,y) and Vy(x,y) are derived using bidirectional optical flow (BIO) techniques. In some embodiments of method 2800B and / or 2800C, the prediction sample P(x,y) satisfies any one of the following: the result of a multiple hypothesis inter prediction technique; the result of an affine prediction technique; the result of an intra prediction technique; the result of an intra block copy (IBC) prediction technique; generated by a triangular prediction mode (TPM) technique; generated by a geographic prediction mode (GPM) technique; the result of a joint inter-intra prediction technique; the result of a global inter prediction technique, which includes regions sharing the same motion model and parameters; the result of a palette coding mode; the result of inter-view prediction in multiview or 3D video coding; the result of inter-layer prediction in scalable video coding; or the result of a filtering operation before determining the refined prediction sample P'(x,y).

[0232] In some embodiments of methods 2800B and / or 2800C, the prediction samples P(x,y) are the final prediction sample values ​​when no refinement process is applied, and the reconstructed sample values ​​Rec(x,y) are obtained by adding the prediction samples P(x,y) with the residual sample values ​​Res(x,y). In some embodiments of methods 2800B and / or 2800C, the refined prediction samples P′(x,y) refined from the prediction samples P(x,y) are the final prediction sample values ​​when the refinement process is applied, and the reconstructed sample values ​​Rec(x,y) are obtained by adding the refined prediction samples P′(x,y) with the residual sample values ​​Res(x,y). In some embodiments of methods 2800B and / or 2800C, a bi-prediction technique or a multi-hypothesis prediction technique is applied to the video block or to a sub-block of the video block, and the first gradient component, the second gradient component, the first motion displacement, and the second motion displacement are applied once to the final predicted sample value.

[0233] In some embodiments of methods 2800B and / or 2800C, a bi-prediction technique or a multiple hypothesis prediction technique is applied to the video block or to a sub-block of the video block, and a first gradient component, a second gradient component, a first motion displacement, and a second motion displacement are applied multiple times to a prediction block of the video block to obtain multiple sets of the first gradient component, the second gradient component, the first motion displacement, and the second motion displacement, and an updated prediction block is obtained by updating each prediction block based on the refined prediction sample P'(x,y), and the updated prediction block is used to generate a final prediction block for the video block. In some embodiments of methods 2800B and / or 2800C, the first set applies to a first predicted block at a location at least one of a first gradient component, a second gradient component, a first motion displacement, or a second motion displacement that is different from the corresponding at least one of a first gradient component, a second gradient component, a first motion displacement, or a second motion displacement applied to a second predicted block at the same location in the second set.

[0234] In some embodiments of method 2800B and / or 2800C, prediction sample P(x,y) is an intermediate prediction sample value from which a final prediction sample value is derived. In some embodiments of method 2800B and / or 2800C, in response to the video block being inter-predicted using bi-prediction techniques, prediction sample P(x,y) is a prediction sample from one reference picture list. In some embodiments of method 2800B and / or 2800C, in response to the video block being inter-predicted using triangular prediction mode (TPM) techniques, prediction sample P(x,y) is a prediction sample from one reference picture list. In some embodiments of method 2800B and / or 2800C, in response to the video block being inter-predicted using geometric partition mode (GPM) techniques, prediction sample P(x,y) is a prediction sample from one reference picture list. In some embodiments of methods 2800B and / or 2800C, in response to the video block being inter-predicted using a multiple hypothesis prediction technique, the prediction sample P(x,y) is a prediction sample from one reference picture list.

[0235] In some embodiments of method 2800B and / or 2800C, prediction sample P(x,y) is an inter-predicted sample for a video block that is inter-inter jointly predicted. In some embodiments of method 2800B and / or 2800C, prediction sample P(x,y) is an inter-predicted sample before a local illumination compensation (LIC) technique is applied to the video block, and the video block uses the LIC technique. In some embodiments of method 2800B and / or 2800C, prediction sample P(x,y) is an inter-predicted sample before a decoder-side motion vector refinement (DMVR) technique or a decoder-side motion vector derivation (DMVD) technique is applied to the video block, and the video block uses the DMVR technique or the DMVD technique. In some embodiments of method 2800B and / or 2800C, prediction sample P(x,y) is an inter-predicted sample before being multiplied by a weighting factor, and the video block uses a weighted prediction technique or a generalized Bi-prediction (GBi) technique.

[0236] In some embodiments of methods 2800B and / or 2800C, a first gradient component in a first direction and / or a second gradient component in a second direction are derived for a final predicted sample value, which, if no refinement process is applied, is added to a residual sample value Res(x,y) to obtain a reconstructed sample value Rec(x,y). In some embodiments of methods 2800B and / or 2800C, the final predicted sample value is a predicted sample P(x,y). In some embodiments of methods 2800B and / or 2800C, the first gradient component in a first direction and / or a second gradient component in a second direction are derived for an intermediate predicted sample value from which the final predicted sample value is to be derived. In some embodiments of methods 2800B and / or 2800C, in response to the video block being inter-predicted using a bi-predictive technique, a first gradient component in a first direction and / or a second gradient component in a second direction are derived from prediction samples from one reference picture list.

[0237] In some embodiments of methods 2800B and / or 2800C, in response to the video block being inter-predicted using a triangular prediction mode (TPM) technique, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from one reference picture list. In some embodiments of methods 2800B and / or 2800C, in response to the video block being inter-predicted using a geometric prediction mode (GPM) technique, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from one reference picture list. In some embodiments of methods 2800B and / or 2800C, in response to the video block being inter-predicted using a multiple hypothesis prediction technique, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from one reference picture list. In some embodiments of methods 2800B and / or 2800C, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from inter-predicted samples for an inter-inter jointly predicted video block.

[0238] In some embodiments of methods 2800B and / or 2800C, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from inter-predicted samples before a local illumination compensation (LIC) technique is applied to the video block, and the video block uses the LIC technique. In some embodiments of methods 2800B and / or 2800C, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from inter-predicted samples before a decoder-side motion vector refinement (DMVR) technique or a decoder-side motion vector derivation (DMVD) technique is applied to the video block, and the video block uses the DMVR technique or the DMVD technique. In some embodiments of methods 2800B and / or 2800C, the first gradient component in the first direction and / or the second gradient component in the second direction are derived from inter-predicted samples before being multiplied by a weighting factor, and the video block uses a weighted prediction technique or a generalized Bi-prediction (GBi) technique.

[0239] In some embodiments of methods 2800B and / or 2800C, the refined prediction sample P'(x,y) is further modified to obtain a final prediction sample value. In some embodiments of methods 2800B and / or 2800C, the first gradient component in the first direction is denoted as Gx(x,y), the second gradient component in the second direction is denoted as Gy(x,y), the first motion displacement is denoted as Vx(x,y), the second motion displacement is denoted as Vy(x,y), and the reconstructed sample value Rec(x,y) at position (x,y) is replaced with the refined reconstructed sample value Rec'(x,y), where Rec'(x,y) = Rec(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y). In some embodiments of methods 2800B and / or 2800C, the first gradient component Gx(x,y) and the second gradient component Gy(x,y) are derived based on the reconstructed sample values.

[0240] In some embodiments of method 2800B and / or 2800C, the method further includes determining a first motion displacement Vx(x,y) at a position (x,y) within the video block and a second motion displacement Vy(x,y) at a position (x,y) based on information from at least spatially neighboring video blocks of the video block or at least information from temporally neighboring video blocks located immediately temporally relative to the video block. In some embodiments of method 2800B and / or 2800C, the spatially neighboring video blocks are located adjacent to the video block and the temporally neighboring video blocks are located immediately temporally relative to the video block. In some embodiments of method 2800B and / or 2800C, the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are determined using bidirectional optical flow (BIO) techniques based on at least spatially neighboring video blocks of the video block or at least temporally neighboring video blocks located immediately temporally relative to the video block.

[0241] In some embodiments of methods 2800B and / or 2800C, the information from spatially neighboring video blocks or from temporally neighboring video blocks includes motion information, coding mode, neighboring coding unit (CU) size, or neighboring CU position. In some embodiments of methods 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to MVMix, which is equal to Wc(x,y)×MVc+WN1(x,y)×MVN1+WN2(x,y)×MVN2+...+WNk(x,y)×MVNk, where MVc is the motion vector of the video block, MVN1,...,MVNk are the motion vectors of the k spatially or temporally neighboring video blocks, N1,...,Nk are the spatial or temporal neighboring video blocks, and Wc, WN1,...,WNk are weighting values ​​that are integers or real numbers.

[0242] In some embodiments of method 2800B and / or 2800C, (Vx(x,y), Vy(x,y)) is equal to MVMix, which is equal to Shift(Wc(x,y)×MVc+WN1(x,y)×MVN1+WN2(x,y)×MVN2+...+WNk(x,y)×MVNk,n1), where the Shift() function refers to a binary shift operation, MVc is the motion vector of the video block, MVN1, ..., MVNk are the motion vectors of k spatially or temporally adjacent video blocks, N1, ..., Nk are spatially or temporally adjacent video blocks, and Wc, WN1, ..., WNk are weighting values ​​that are integers or real numbers, and n1 is an integer. In some embodiments of method 2800B and / or 2800C, (Vx(x,y), Vy(x,y)) is equal to MVMix, which is equal to SatShift(Wc(x,y)×MVc+WN1(x,y)×MVN1+WN2(x,y)×MVN2+...+WNk(x,y)×MVNk,n1), where the SatShift() function refers to a saturated binary shift operation, MVc is the motion vector of the video block, MVN1, ..., MVNk are the motion vectors of k spatially or temporally adjacent video blocks, N1, ..., Nk are spatially or temporally adjacent video blocks, and Wc, WN1, ..., WNk are weighting values ​​that are integers or real numbers, and n1 is an integer.

[0243] In some embodiments of method 2800B and / or 2800C, Wc(x,y) is equal to zero. In some embodiments of method 2800B and / or 2800C, k is equal to 1 and N1 is a spatially neighboring video block. In some embodiments of method 2800B and / or 2800C, N1 is a spatially closest neighboring video block to position (x,y) in the video block. In some embodiments of method 2800B and / or 2800C, WN1(x,y) increases in value as N1 approaches position (x,y) in the video block. In some embodiments of method 2800B and / or 2800C, k is equal to 1 and N1 is a temporally neighboring video block. In some embodiments of method 2800B and / or 2800C, N1 is a co-located video block in a co-located picture for position (x,y) in the video block. In some embodiments of methods 2800B and / or 2800C, in response to the prediction sample not being located at the top or left boundary of the video block, the prediction sample in the basic video block is not refined.

[0244] In some embodiments of method 2800B and / or 2800C, a prediction sample within a basic video block is refined in response to the prediction sample being located at the top boundary of the video block. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to MVMix, and MVMix for a basic video block located at the top boundary of the video block is derived based on a neighboring video block located above the video block. In some embodiments of method 2800B and / or 2800C, a prediction sample within a basic video block is refined in response to the prediction sample being located at the left boundary of the video block. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to MVMix, and MVMix for a basic video block located at the left boundary of the video block is derived based on a neighboring video block located to the left of the video block. In some embodiments of methods 2800B and / or 2800C, the video block and motion vector and the motion vectors of a number of spatially neighboring video blocks or a number of temporally neighboring video blocks are scaled to the same reference picture.

[0245] In some embodiments of method 2800B and / or 2800C, the same reference picture is a reference picture referenced by the motion vector of the video block. In some embodiments of method 2800B and / or 2800C, spatial or temporal neighboring video blocks are used to derive (Vx(x,y), Vy(x,y)) only if the spatial or temporal neighboring video blocks are not intra-coded. In some embodiments of method 2800B and / or 2800C, spatial or temporal neighboring video blocks are used to derive (Vx(x,y), Vy(x,y)) only if the spatial or temporal neighboring video blocks are not intra-block copying (IBC) predictively coded. In some embodiments of methods 2800B and / or 2800C, (Vx(x,y), Vy(x,y)) is derived using spatially adjacent or temporally adjacent video blocks only if the first motion vector of the spatially adjacent or temporally adjacent video block references the same reference picture as the second motion vector of the video block.

[0246] In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to f(MVMix,MVc), where f is a function and MVc is a motion vector for the video block. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to MVMix - MVc. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to MVc - MVMix. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to p x MVMix + q x MVc, where p and q are real numbers. In some embodiments of method 2800B and / or 2800C, (Vx(x,y),Vy(x,y)) is equal to Shift(p×MVMix+q×MVc,n) or SatShift(p×MVMix+p×MVc,n), where p, q, and n are integers, the Shift() function refers to a binary shift operation, and the SatShift() function refers to a saturated binary shift operation. In some embodiments of method 2800B and / or 2800C, the video block is unidirectionally inter predicted and the motion vector of the video block references reference picture list 0. In some embodiments of method 2800B and / or 2800C, the video block is unidirectionally inter predicted and the motion vector of the video block references reference picture list 1.

[0247] In some embodiments of methods 2800B and / or 2800C, the video block is inter-predicted using bi-predictive techniques, and the motion vector of the video block references reference picture list 0 or reference picture list 1. In some embodiments of methods 2800B and / or 2800C, the final predicted sample value is refined based on a first motion displacement Vx(x,y) and a second motion displacement Vy(x,y) derived from the motion vector of the video block that references reference picture list 0 or reference picture list 1. In some embodiments of methods 2800B and / or 2800C, the first predicted sample value from reference picture list 0 is refined based on a first motion displacement Vx(x,y) and a second motion displacement Vy(x,y) derived from the motion vector of the video block that references reference picture list 0. In some embodiments of methods 2800B and / or 2800C, a second predicted sample value from reference picture list 1 is refined based on a first motion displacement Vx(x,y) and a second motion displacement Vy(x,y) derived from a motion vector of a video block that references reference picture list 1.

[0248] In some embodiments of methods 2800B and / or 2800C, a final predicted sample value is derived using a first predicted sample value refined from reference picture list 0 and a second predicted sample value refined from reference picture list 1. In some embodiments of methods 2800B and / or 2800C, the final predicted sample value is equal to the average or weighted average of the first predicted sample value refined from reference picture list 0 and the second predicted sample value refined from reference picture list 1. In some embodiments of methods 2800B and / or 2800C, the video block is inter-predicted using a bi-prediction technique, a bidirectional optical flow (BIO) technique is applied to the video block, and the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are modified based on spatially or temporally neighboring video blocks.

[0249] In some embodiments of methods 2800B and / or 2800C, a first motion displacement vector V'(x,y) = (V'x(x,y), V'y(x,y)) is derived from BIO techniques, a second motion displacement vector derived using a method other than BIO techniques is denoted as V"(x,y) = (V"x(x,y), V"y(x,y)), and a third motion displacement vector V(x,y) = (Vx(x,y), Vy(x,y)) is derived from the first set of motion displacement vectors. and the second set of motion displacement vectors, and a third motion displacement vector V(x,y) is used to refine the prediction sample P(x,y). In some embodiments of methods 2800B and / or 2800C, the third set of motion displacement vectors is obtained using the following formula: V(x,y)=V'(x,y)×W'(x,y)+V"(x,y)×W"(x,y), where W'(x,y) and W"(x,y) are integers or real numbers. In some embodiments of methods 2800B and / or 2800C, the third set of motion displacement vectors is obtained using the following formula: V(x,y)=Shift(V'(x,y)×W'(x,y)+V"(x,y)×W"(x,y),n1), where the Shift() function refers to a binary shift operation, W'(x,y) and W"(x,y) are integers, and n1 is a non-negative integer.

[0250] In some embodiments of method 2800B and / or 2800C, the third set of motion displacement vectors is obtained using the following formula: V(x,y)=SatShift(V'(x,y)×W'(x,y)+V"(x,y)×W"(x,y),n1), where the SatShift() function refers to a saturated binary shift operation, W'(x,y) and W"(x,y) are integers, and n1 is a non-negative integer. In some embodiments of method 2800B and / or 2800C, the first motion displacement Vx(x,y) and the second motion displacement Vx(x,y) are The motion displacement Vy(x,y) is modified based on the spatially or temporally adjacent video blocks and based on the position (x,y). In some embodiments of methods 2800B and / or 2800C, the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) that are not located on the top boundary of the video block or the left boundary of the video block are not modified. In some embodiments of methods 2800B and / or 2800C, the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are clipped.

[0251] In some embodiments of method 2800B and / or 2800C, the spatially or temporally neighboring video block is a video block that is not adjacent to the video block. In some embodiments of method 2800B and / or 2800C, the spatially or temporally neighboring video block is a sub-block that is not adjacent to the video block. In some embodiments of method 2800B and / or 2800C, the spatially or temporally neighboring video block is a sub-block that is not adjacent to any of the video block, or a coding tree unit (CTU), or a video processing and distribution unit (VPDU), or a current region covering the sub-block. In some embodiments of method 2800B and / or 2800C, the motion vector of the spatially or temporally neighboring video block comprises an entry in a history-based motion vector. In some embodiments of methods 2800B and / or 2800C, determining the first motion displacement and the second motion displacement at the decoder side includes determining the presence of information related to the first motion displacement and the second motion displacement by parsing a bitstream representation of the video block.

[0252] In some embodiments of method 2800B and / or 2800C, the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived at a sub-block level of the video block. In some embodiments of method 2800B and / or 2800C, the first motion displacement and the second motion displacement are derived at a 2x2 block level. In some embodiments of method 2800B and / or 2800C, the first motion displacement and the second motion displacement are derived at a 4x1 block level. In some embodiments of method 2800B and / or 2800C, the video block is an 8x4 video block, or the sub-block is an 8x4 sub-block. In some embodiments of method 2800B and / or 2800C, the video block is a 4x8 video block, or the sub-block is a 4x8 sub-block. In some embodiments of methods 2800B and / or 2800C, the video blocks are 4x4 single-predicted video blocks, or the sub-blocks are 4x4 single-predicted sub-blocks. In some embodiments of methods 2800B and / or 2800C, the video blocks are 8x4, 4x8, and 4x4 single-predicted video blocks, or the sub-blocks are 8x4, 4x8, and 4x4 single-predicted sub-blocks.

[0253] In some embodiments of method 2800B and / or 2800C, the video block excludes a 4x4 bi-predictive video block or the sub-block excludes a 4x4 bi-predictive sub-block. In some embodiments of method 2800B and / or 2800C, the video block is a luma video block. In some embodiments of method 2800B and / or 2800C, the determining and performing are based on any one or more of: a color component of the video block, a block size of the video block, a color format of the video block, a block position of the video block, a motion type, a motion vector magnitude, a coding mode, a pixel gradient magnitude, a transform type, whether a bi-directional optical flow (BIO) technique is applied, whether a bi-predictive technique is applied, and whether a decoder-side motion vector refinement (DMVR) technique is applied.

[0254] 28D is a flowchart of an example video processing method. Method 2800D includes determining (2832) a first motion displacement Vx(x,y) at a position (x,y) and a second motion displacement Vy(x,y) at a position (x,y) within a video block to be coded using an optical flow-based method, where x and y are fractional numbers, where Vx(x,y) and Vy(x,y) are determined based on at least the position (x,y) and a center position of a basic video block of the video block, and performing (2834) a conversion between the video block and a bitstream representation of the current video block using the first motion displacement and the second motion displacement.

[0255] In some embodiments of method 2800D, Vx(x,y)=a×(x-xc)+b×(y-yc), Vy(x,y)=c×(x-xc)+d×(y-yc), where (xc,yc) is the center position of the basic image block of the image block, a, b, c and d are affine parameters, the basic image block has dimensions w×h, and the position of the basic image block includes position (x,y). In some embodiments of method 2800D, Vx(x,y)=Shift(a×(x-xc)+b×(y-yc),n1), Vy(x,y)=Shift(c×(x-xc)+d×(y-yc),n1), where (xc,yc) is the center position of the elementary video block of the image block, a, b, c and d are affine parameters, the elementary video block has dimensions w×h, and the position of the elementary video block comprises position (x,y), the Shift() function refers to a binary shift operation; and In some embodiments of method 2800D, Vx(x,y)=SatShift(a×(x-xc)+b×(y-yc),n1), Vy(x,y)=SatShift(c×(x-xc)+d×(y-yc),n1), where (xc,yc) is the center position of the elementary video block of the image block, a, b, c, and d are affine parameters, the elementary video block has dimensions w×h, the position of the elementary video block comprises position (x,y), the SatShift() function refers to a saturated binary shift operation, and n1 is an integer.

[0256] In some embodiments of method 2800D, Vx(x,y)=-a×(x-xc)-b×(y-yc), Vy(x,y)=-c×(x-xc)-d×(y-yc), where (xc,yc) is the center position of the basic video block of the image block, a, b, c, and d are affine parameters, the basic video block has dimensions w×h, and the position of the basic video block comprises position (x,y). In some embodiments of method 2800D, (xc,yc)=(x0+(w / 2),y0+(h / 2)), and the top left position of the basic video block is (x0,y0). In some embodiments of method 2800D, (xc,yc)=(x0+(w / 2)-1,y0+(h / 2)-1), and the top left position of the basic video block is (x0,y0). In some embodiments of method 2800D, (xc, yc)=(x0+(w / 2), y0+(h / 2)-1), and the top-left position of the elementary image block is (x0, y0).

[0257] In some embodiments of method 2800D, (xc, yc) = (x0 + (w / 2) - 1, y0 + (h / 2)), and the top-left location of the basic video block is (x0, y0). In some embodiments of method 2800D, c = -b and d = a, responsive to the video block being coded using a four-parameter affine mode. In some embodiments of method 2800D, a, b, c, and d may be derived from the control point motion vector (CPMV), the width (W) of the video block, and the height (H) of the video block. In some embodiments of method 2800D,

number

[0258] In some embodiments of method 2800D, a, b, c, and d may be obtained from history-based stored information. In some embodiments of method 2800D, Vx(x+1,y)=Vx(x,y)+a, and Vy(x+1,y)=Vy(x,y)+c. In some embodiments of method 2800D, Vx(x+1,y)=Shift(Vx(x,y)+a,n1) and Vy(x+1,y)=Shift(Vy(x,y)+c,n1), where the Shift() function refers to a binary shift operation, and n1 is an integer. In some embodiments of method 2800D, Vx(x+1,y) = SatShift(Vx(x,y) + a,n1), Vy(x+1,y) = SatShift(Vy(x,y) + c,n1), where the SatShift() function refers to a saturated binary shift operation, and n1 is an integer. In some embodiments of method 2800D, Vx(x+1,y) = Vx(x,y) + Shift(a,n1), Vy(x+1,y) = Vy(x,y) + Shift(c,n1), where the Shift() function refers to a binary shift operation, and n1 is an integer.

[0259] In some embodiments of method 2800D, Vx(x+1,y) = Vx(x,y) + SatShift(a,n1), Vy(x+1,y) = Vy(x,y) + SatShift(c,n1), where the SatShift() function refers to a saturated binary shift operation, and n1 is an integer. In some embodiments of method 2800D, Vx(x,y+1) = Vx(x,y) + b, and Vy(x+1,y) = Vy(x,y) + d. In some embodiments of method 2800D, Vx(x,y+1) = Shift(Vx(x,y) + b,n1), Vy(x,y+1) = Shift(Vy(x,y) + d,n1), where the Shift() function refers to a binary shift operation, and n1 is an integer. In some embodiments of method 2800D, Vx(x,y+1) = SatShift(Vx(x,y) + b,n1), Vy(x,y+1) = SatShift(Vy(x,y) + d,n1), where the SatShift() function refers to a saturated binary shift operation, and n1 is an integer. In some embodiments of method 2800D, Vx(x,y+1) = Vx(x,y) + Shift(b,n1), Vy(x,y+1) = Vy(x,y) + Shift(d,n1), where the Shift() function refers to a binary shift operation, and n1 is an integer.

[0260] In some embodiments of method 2800D, Vx(x,y+1) = Vx(x,y) + SatShift(b,n1), Vy(x,y+1) = Vy(x,y) + SatShift(d,n1), where the SatShift() function refers to a saturated binary shift operation, and n1 is an integer. In some embodiments of method 2800D, a, b, c, and d refer to reference picture list 0 or reference picture list 1 in response to the video block being affine predicted using bi-prediction techniques. In some embodiments of method 2800D, the final predicted sample is refined with a first motion displacement Vx(x,y) and a second motion displacement Vy(x,y), and the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived using a, b, c, and d that refer to either reference picture list 0 or reference picture list 1.

[0261] In some embodiments of method 2800D, the prediction samples for the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are from reference picture list 0 and are refined, and the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived using a, b, c, and d that refer to reference picture list 0. In some embodiments of method 2800D, the prediction samples for the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are from reference picture list 1 and are refined, and the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived using a, b, c, and d that refer to reference picture list 1. In some embodiments of method 2800D, the first prediction sample from reference picture list 0 is refined, and the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived using a, b, c, and d that refer to reference picture list 0. 0 x(x,y),V 0 y(x,y)), and a second prediction sample from reference picture list 1 is derived using a, b, c, and d that refer to reference picture list 1. 1 x(x,y),V 1y(x,y)), and the final predicted sample is obtained by combining the first predicted sample with the second predicted sample.

[0262] In some embodiments of method 2800D, the motion displacement vector (Vx(x,y), Vy(x,y)) has a first motion vector precision that is different from the second motion vector precision of the motion vector of the video block. In some embodiments of method 2800D, the first motion vector precision is 1 / 8 pixel precision. In some embodiments of method 2800D, the first motion vector precision is 1 / 16 pixel precision. In some embodiments of method 2800D, the first motion vector precision is 1 / 32 pixel precision. In some embodiments of method 2800D, the first motion vector precision is 1 / 64 pixel precision. In some embodiments of method 2800D, the first motion vector precision is 1 / 128 pixel precision. In some embodiments of method 2800D, the motion displacement vector (Vx(x,y), Vy(x,y)) is determined based on a float pixel precision technique.

[0263] 28E is a flowchart of an example video processing method. The method 2800E includes determining (2842) a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within the video block and a second gradient component Gy(x,y) in a second direction estimated at a position (x,y) within the video block, the first gradient component and the second gradient component being based on a final predicted sample value of a prediction sample P(x,y) at the position (x,y), where x and y are integers, and performing (2844) a conversion between the video block and a bitstream representation of the current video block using a reconstructed sample value Rec(x,y) at the position (x,y) that is obtained based on a residual sample value Res(x,y) plus a final predicted sample value of the prediction sample P(x,y) refined using the gradients Gx(x,y), Gy(x,y).

[0264] FIG. 28F is a flowchart of an exemplary video processing method. Method 2800F includes determining (2852) a first gradient component Gx(x,y) in a first direction estimated at a position (x,y) within the video block and a second gradient component Gy(x,y) in a second direction estimated at a position (x,y) within the video block, where the first gradient component and the second gradient component are based on a final predicted sample value of a prediction sample P(x,y) at the position (x,y), where x and y are integers; and encoding (2854) a bitstream representation of the video block to include a residual sample value Res(x,y) that is based on a reconstructed sample value Rec(x,y) at the position (x,y), where the reconstructed sample value Rec(x,y) is based on the final predicted sample value of the prediction sample P(x,y) refined using the gradients Gx(x,y), Gy(x,y) plus the residual sample value Res(x,y).

[0265] 28G is a flowchart of an example video processing method. Method 2800G includes determining (2862) a first gradient component Gx(x,y) in a first direction estimated at a location (x,y) within the video block and a second gradient component Gy(x,y) in a second direction estimated at a location (x,y) within the video block, where the first gradient component and the second gradient component are based on intermediate predicted sample values ​​of prediction samples P(x,y) at location (x,y), where the final predicted sample values ​​of prediction samples P(x,y) are based on the intermediate predicted sample values, and x and y are integers, and performing (2864) a conversion between the video block and a bitstream representation of the current video block using reconstructed sample values ​​Rec(x,y) at location (x,y) obtained based on the final predicted sample values ​​of prediction samples P(x,y) and residual sample values ​​Res(x,y).

[0266] 28H is a flowchart of an example video processing method. The method 2800H includes determining (2872) a first gradient component Gx(x,y) in a first direction estimated at a location (x,y) within the video block and a second gradient component Gy(x,y) in a second direction estimated at a location (x,y) within the video block, where the first gradient component and the second gradient component are based on intermediate predicted sample values ​​of prediction samples P(x,y) at the location (x,y), where the final predicted sample values ​​of prediction samples P(x,y) are based on the intermediate predicted sample values, and x and y are integers; and encoding (2874) a bitstream representation of the video block to include residual sample values ​​Res(x,y) that are based on reconstructed sample values ​​Rec(x,y) at the location (x,y), where the reconstructed sample values ​​Rec(x,y) are based on the final predicted sample values ​​of prediction samples P(x,y) and the residual sample values ​​Res(x,y).

[0267] In some embodiments of methods 2800G and / or 2800H, in response to the video block being affine predicted using bi-prediction techniques, the first gradient component and / or the second gradient component are based on intermediate predicted sample values ​​from one reference picture list. In some embodiments of methods 2800G and / or 2800H, when the video block uses affine mode and local illumination compensation (LIC), the first gradient component and / or the second gradient component are based on inter-predicted sample values ​​before LIC is applied. In some embodiments of methods 2800G and / or 2800H, when the video block uses affine mode with either a weighted prediction technique or a CU-level weighted bi-prediction (BCW) technique, the first gradient component and / or the second gradient component are based on inter-predicted sample values ​​before being multiplied by a weighting factor. In some embodiments of methods 2800G and / or 2800H, when the video block uses affine mode and local illumination compensation (LIC), the first gradient component and / or the second gradient component are based on inter-predicted sample values ​​to which LIC has been applied. In some embodiments of methods 2800G and / or 2800H, when the video block uses affine mode with either a weighted prediction technique or a CU-level weighted bi-prediction (BCW) technique, the first gradient component and / or the second gradient component are based on inter-predicted sample values ​​multiplied by a weighting factor.

[0268] 28I is a flowchart of an example video processing method. Method 2800I determines (2822) a refined prediction sample P'(x,y) at a location (x,y) in an affine-coded video block by modifying the prediction sample P(x,y) at the location (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the location (x,y), a second gradient component Gy(x,y) in a second direction estimated at the location (x,y), a first motion displacement Vx(x,y) estimated for the location (x,y), and a second motion displacement Vy(x,y) estimated for the location (x,y), where the first direction is orthogonal to the second direction, and x and y are integers; determining (2884) a reconstructed sample value Rec(x,y) at location (x,y) based on the sample P'(x,y) and the residual sample value Res(x,y); determining (2886) a refined reconstructed sample value Rec'(x,y) at location (x,y) within the affine-coded video block, where Rec'(x,y) = Rec(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y), and performing (2888) a conversion between the affine-coded video block and the bitstream representation of the affine-coded video block using the refined reconstructed sample value Rec'(x,y).

[0269] 28J is a flowchart of an example video processing method. Method 2800J determines (2892) a refined prediction sample P'(x,y) at a location (x,y) within an affine-coded video block by modifying the prediction sample P(x,y) at the location (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the location (x,y), a second gradient component Gy(x,y) in a second direction estimated at the location (x,y), a first motion displacement Vx(x,y) estimated for the location (x,y), and a second motion displacement Vy(x,y) estimated for the location (x,y), where the first direction is orthogonal to the second direction, and x and y are and determining (2894) a reconstructed sample value Rec(x,y) at position (x,y) based on the refined prediction sample P'(x,y) and the residual sample value Res(x,y), where P'(x,y) is an integer, determining (2896) a refined reconstructed sample value Rec'(x,y) at position (x,y) within the affine-coded video block, where Rec'(x,y)=Rec(x,y)+Gx(x,y)×Vx(x,y)+Gy(x,y)×Vy(x,y), and encoding (2898) a bitstream representation of the affine-coded video block to include the residual sample value Res(x,y).

[0270] In some embodiments of method 2800I and / or 2800J, the first gradient component Gx(x,y) and / or the second gradient component Gy(x,y) are based on the reconstructed sample values ​​Rec(x,y). In some embodiments of method 2800I and / or 2800J, the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) are derived at the 2x2 block level of the affine-coded video block.

[0271] 28K is a flowchart of an example video processing method. Method 2800K includes determining (28102) motion vectors with 1 / N pixel accuracy for a video block in affine mode, determining (28104) estimated motion displacement vectors (Vx(x,y), Vy(x,y)) for a location (x,y) within the video block, where the motion displacement vectors are derived with 1 / M pixel accuracy, where N and M are positive integers, and where x and y are integers, and performing (28106) a conversion between the video block and a bitstream representation of the video block using the motion vectors and motion displacement vectors.

[0272] In some embodiments of method 2800K, N is equal to 8 and M is equal to 16, or N is equal to 8 and M is equal to 32, or N is equal to 8 and M is equal to 64, or N is equal to 8 and M is equal to 128, or N is equal to 4 and M is equal to 8, or N is equal to 4 and M is equal to 16, or N is equal to 4 and M is equal to 32, or N is equal to 4 and M is equal to 64, or N is equal to 4 and M is equal to 128.

[0273] 28L is a flowchart of an example video processing method. Method 2800L includes determining (28112) two sets of motion vectors for a video block or for a sub-block of the video block, each set of the two sets of motion vectors having a different motion vector pixel precision, and the two sets of motion vectors being determined using a temporal motion vector prediction (TMVP) technique or a sub-block-based temporal motion vector prediction (SbTMVP) technique, and performing (28114) a conversion between the video block and a bitstream representation of the video block based on the two sets of motion vectors.

[0274] In some embodiments of method 2800L, the two sets of motion vectors include a first set of motion vectors and a second set of motion vectors, where the first set of motion vectors has 1 / N pixel accuracy and the second set of motion vectors has 1 / M pixel accuracy, and N and M are positive integers. In some embodiments of method 2800L, N is equal to 8 and M is equal to 16, or N is equal to 8 and M is equal to 32, or N is equal to 8 and M is equal to 64, or N is equal to 8 and M is equal to 128, or N is equal to 4 and M is equal to 8, or N is equal to 4 and M is equal to 16, or N is equal to 4 and M is equal to 32, or N is equal to 4 and M is equal to 64, or N is equal to 4 and M is equal to 128, or N is equal to 16 and M is equal to 32, or N is equal to 16 and M is equal to 64, or N is equal to 16 and M is equal to 128. In some embodiments of method 2800L, refinement is applied to two sets of motion vectors by applying an optical flow-based method, the two sets of motion vectors including a first set of motion vectors and a second set of motion vectors, a predicted sample in the optical flow-based method is obtained using the first set of motion vectors, and an estimated motion displacement for a position within the video block in the optical flow-based method is obtained by subtracting a second motion vector in the second set of motion vectors from a first motion vector in the first set of motion vectors.

[0275] In some embodiments of Method 2800L, an optical flow-based method is applied by determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers, and a reconstructed sample value Rec(x,y) at the position (x,y) is obtained based on the refined prediction sample P'(x,y) and the residual sample value Res(x,y). In some embodiments of Method 2800L, the first direction and the second direction are orthogonal to each other. In some embodiments of method 2800L, the first motion displacement represents a direction parallel to the first direction and the second motion displacement represents a direction parallel to the second direction. In some embodiments of method 2800L, P'(x,y) = P(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y).

[0276] In some embodiments of any one or more of methods 2800C through 2800L, the video block is an 8x4 video block, or the sub-block is an 8x4 sub-block. In some embodiments of any one or more of methods 2800C through 2800L, the video block is a 4x8 video block, or the sub-block is a 4x8 sub-block. In some embodiments of any one or more of methods 2800C through 2800L, the video block is a 4x4 uni-predicted video block, or the sub-block is a 4x4 uni-predicted sub-block. In some embodiments of any one or more of methods 2800C through 2800L, the video block is an 8x4, 4x8, and 4x4 uni-predicted video block, or the sub-block is an 8x4, 4x8, and 4x4 uni-predicted sub-block. In some embodiments of any one or more of methods 2800C through 2800L, the video block excludes a 4x4 bi-predicted video block, or the sub-block excludes a 4x4 bi-predicted sub-block. In some embodiments of any one or more of methods 2800C through 2800L, the video block is a luma video block. In some embodiments of any one or more of methods 2800C through 2800L, the determining and performing are performed based on any one or more of the following: a color component of the video block, a block size of the video block, a color format of the video block, a block position of the video block, a motion type, a motion vector magnitude, a coding mode, a pixel gradient magnitude, a transform type, whether a bidirectional optical flow (BIO) technique is applied, whether a bi-prediction technique is applied, and whether a decoder-side motion vector refinement (DMVR) technique is applied.

[0277] 28M is a flowchart of an example video processing method. Method 2800M performs (28122) an interweave prediction technique on a video block to be coded using an affine coding mode by dividing the video block into multiple partitions using K different sub-block patterns, where K is an integer greater than 1, and generates (28124) predicted samples for the video block by motion compensation using a first of the K different sub-block patterns, denoted as P(x,y), where x and y are integers, and generates (28125) predicted samples for the video block by motion compensation using a first of the K different sub-block patterns, denoted as P(x,y), where x and y are integers, and generates (28126) predicted samples for the video block using a second of the K different sub-block patterns, denoted as P(x,y), where x and y are integers. for at least one of the K subblock patterns, determining (28126) an offset value OL(x,y) at position (x,y) based on the predicted sample derived with the first subblock pattern and the difference between the motion vector derived using the first of the K subblock patterns and the motion vector derived using the Lth pattern; determining (28128) a final predicted sample for position (x,y) as a function of OL(x,y) and P(x,y); and performing (28130) a conversion between the bitstream representation of the video block and the video block using the final predicted sample.

[0278] In some embodiments of method 2800M, K=2, L=1, and the final predicted sample is determined using P(x,y)+((O1(x,y)+1)>>1), where >> represents a binary shift operation. In some embodiments of method 2800M, K=2, L=1, and the final predicted sample is determined using P(x,y)+(O1(x,y)>>1), where >> represents a binary shift operation. In some embodiments of method 2800M, the final predicted sample is determined using P(x,y)+(O1(x,y)+...+OK(x,y)+K / 2) / K. In some embodiments of method 2800M, the final predicted sample is determined using P(x,y)+(O1(x,y)+...+OK(x,y)) / K. In some embodiments of method 2800M, OL(x,y) is generated from prediction samples derived in the first sub-block pattern after horizontal and vertical interpolation and before converting the prediction samples to the bit depth of the input samples. In some embodiments of method 2800M, OL(x,y) is generated for each prediction direction. In some embodiments of method 2800M, the motion displacement at position (x,y) in the Lth pattern is derived as the difference between a first motion vector at position (x,y) from the Lth pattern and a second motion vector at position (x,y) from the first of the K sub-block patterns.

[0279] In some embodiments of method 2800M, motion vectors derived using a first of the K sub-block patterns have 1 / N pixel precision, and motion vectors derived using the Lth pattern have 1 / ML pixel precision, where N and M are integers. In some embodiments of method 2800M, N=16 and ML=32, or N=16 and ML=64, or N=16 and ML=128, or N=8 and ML=16, or N=8 and ML=32, or N=8 and ML=64, or N=8 and ML=128, or N=4 and ML=8, or N=4 and ML=16, or N=4 and ML=32, or N=4 and ML=64, or N=4 and ML=128. In some embodiments of method 2800M, ML is different for each of the remaining K sub-block patterns.

[0280] 28N is a flowchart of an example video processing method. Method 2800N includes performing (28132) a conversion between a bitstream representation of the video block and the video block using final prediction samples, where the final prediction samples are derived from refined intermediate prediction samples by (a) performing an interweave prediction technique and a subsequent optical flow-based prediction refinement technique based on a rule, or (b) performing a motion compensation technique.

[0281] In some embodiments of method 2800N, the video block is coded using an affine coding mode by dividing the video block into multiple partitions using K different sub-block patterns, where K is an integer greater than 1, and for each sub-block pattern, a prediction sample is generated for each sub-block by performing a motion compensation technique, the prediction sample is refined using an optical flow-based prediction refinement technique to obtain updated prediction samples, and a final prediction sample is generated by combining the refined prediction samples of each sub-block pattern. In some embodiments of method 2800N, the video block is a uni-predicted video block, and an interweave prediction technique and an optical flow-based prediction refinement technique are applied to the video block. In some embodiments of method 2800N, a first tap filter is applied to a video block that is coded using an interweave prediction technique and / or an optical flow-based prediction refinement technique, and the first tap filter is shorter than a second tap filter used in an interpolation filter used for other video blocks that are not coded using an interweave prediction technique and / or an optical flow-based prediction refinement technique.

[0282] 82. The method of claim 81, wherein the affine sub-block size for the video block is 8x4 or 4x8. 82. In some embodiments of method 2800N, a first sub-block size is used for a video block that is coded with interweave prediction and / or optical flow-based prediction refinement techniques, the first sub-block size being different from a second sub-block size for other video blocks that are not coded with interweave prediction and / or optical flow-based prediction refinement techniques. 82. The method of claim 82, wherein the rule is based on coding information of the video block, the coding information including a prediction direction for the video block, reference picture information for the video block, or color components of the video block.

[0283] 28O is a flowchart of an example video processing method. Method 2800O includes, when bi-prediction is applied, performing (28142) a conversion between a bitstream representation of a video block and the video block using final prediction samples, where the final prediction samples are derived from refined intermediate prediction samples by (a) disabling interweave prediction techniques and performing optical flow-based prediction refinement techniques, or (b) performing motion compensation techniques.

[0284] 28P is a flowchart of an example video processing method. Method 2800P shows a fifteenth example of a video processing method that includes performing (28152) a conversion between a bitstream representation of a video block and the video block using prediction samples, the prediction samples being derived from refined intermediate prediction samples by performing an optical flow-based prediction refinement technique that depends on only one of a first set of motion displacements Vx(x,y) estimated in a first direction for the video block or a second set of motion displacements Vy(x,y) estimated in a second direction for the video block, where x and y are integers and the first direction is orthogonal to the second direction.

[0285] In some embodiments of method 2800P, the prediction sample is based only on the first set of motion displacements Vx(x,y), and the second set of motion displacements Vy(x,y) is zero. In some embodiments of method 2800P, the prediction sample is based only on the second set of motion displacements Vy(x,y), and the first set of motion displacements Vx(x,y) is zero. In some embodiments of method 2800P, in response to the sum of the absolute values ​​of the first set of motion displacements Vx(x,y) being greater than or equal to the sum of the absolute values ​​of the second set of motion displacements Vy(x,y), the prediction sample is based only on the first set of motion displacements Vx(x,y). In some embodiments of method 2800P, in response to the sum of the absolute values ​​of the first set of motion displacements Vx(x,y) being less than or equal to the sum of the absolute values ​​of the second set of motion displacements Vy(x,y), the prediction sample is based only on the second set of motion displacements Vy(x,y).

[0286] In some embodiments of method 2800P, in response to the sum of the absolute values ​​of the first gradient components in the first direction being greater than or equal to the sum of the absolute values ​​of the second gradient components in the second direction, the prediction sample is based solely on the first set of motion displacements Vx(x,y). In some embodiments of method 2800P, in response to the sum of the absolute values ​​of the first gradient components in the first direction being greater than or equal to the sum of the absolute values ​​of the second gradient components in the second direction, the prediction sample is based solely on the second set of motion displacements Vy(x,y).

[0287] 28Q is a flowchart of an example video processing method. Method 2800Q includes obtaining (28162) a refined motion vector for the video block by refining the motion vector of the video block, where the motion vector is refined before performing a motion compensation technique, the refined motion vector having 1 / N pixel accuracy, and the motion vector having 1 / M pixel accuracy; obtaining (28164) a final prediction sample by performing an optical flow-based prediction refinement technique on the video block; applying the optical flow-based prediction refinement technique to a difference between the refined motion vector and the motion vector; and performing (28166) a conversion between the bitstream representation of the video block and the video block using the final prediction sample.

[0288] In some embodiments of method 2800Q, M is 16 and N is 1, or M is 8 and N is 1, or M is 4 and N is 1, or M is 16 and N is 2, or M is 8 and N is 2, or M is 4 and N is 2. In some embodiments of method 2800Q, the motion vector is in either a first direction or a second direction, the first direction being orthogonal to the second direction, and the optical flow-based prediction refinement technique is performed in either the first direction or the second direction. In some embodiments of method 2800Q, the video block is a bi-predictive video block, and the motion vector is refined in a first direction or a second direction, the first direction being orthogonal to the second direction. In some embodiments of method 2800Q, the video block is a bi-predictive video block, and the motion vector is refined in a first direction and a second direction, the first direction being orthogonal to the second direction. In some embodiments of method 2800Q, the optical flow-based prediction refinement technique is performed on a first number of partial motion vector components of the motion vector, the first number of partial motion vector components being less than or equal to a second number of partial motion vector components of the motion vector.

[0289] 28R is a flowchart of an example video processing method. Method 2800R includes determining (28172) a final motion vector using a multi-step decoder-side motion vector refinement process for a video block, the final motion vector having 1 / N pixel accuracy, and performing (28174) a conversion between the current block and a bitstream representation using the final motion vector.

[0290] In some embodiments of method 2800R, N is equal to 32, 64, or 128. In some embodiments of method 2800R, the final motion vector is a refined motion vector obtained by refining a motion vector for the video block, the motion vector being refined prior to performing a motion compensation technique, and the final prediction sample being obtained by performing an optical flow-based prediction refinement technique on the video block, and the optical flow-based prediction refinement technique being applied to the difference between the final motion vector and the motion vector. In some embodiments of method 2800R, an optical flow-based prediction refinement technique is applied by determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) with a first gradient component Gx(x,y) in a first direction estimated at the position (x,y), a second gradient component Gy(x,y) in a second direction estimated at the position (x,y), a first motion displacement Vx(x,y) estimated for the position (x,y), and a second motion displacement Vy(x,y) estimated for the position (x,y), where x and y are integers, and a final prediction sample Rec(x,y) at the position (x,y) is obtained based on the refined prediction sample P'(x,y) and the residual sample value Res(x,y). In some embodiments of method 2800R, the first direction and the second direction are orthogonal to each other. 105. The method of claim 104, wherein in some embodiments of method 2800R, the first movement displacement represents a direction parallel to a first direction and the second movement displacement represents a direction parallel to a second direction.

[0291] In some embodiments of method 2800R, P'(x,y) = P(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y). In some embodiments of method 2800R, refinement motion vectors with 1 / 32 pixel precision are rounded to 1 / 16 pixel precision prior to performing the motion compensation technique. In some embodiments of method 2800R, refinement motion vectors with 1 / 32 pixel precision are rounded to 1 pixel precision prior to performing the motion compensation technique. In some embodiments of method 2800R, refinement motion vectors with 1 / 64 pixel precision are rounded to 1 / 16 pixel precision prior to performing the motion compensation technique. In some embodiments of method 2800R, refinement motion vectors with 1 / 64 pixel precision are rounded to 1 pixel precision prior to performing the motion compensation technique.

[0292] 28S is a flowchart of an example video processing method. Method 2800S includes obtaining (28182) refined intermediate prediction samples of a video block by performing an interweave prediction technique and an optical flow-based prediction refinement technique on the intermediate prediction samples of the video block, deriving (28184) final prediction samples from the refined intermediate prediction samples, and performing (28186) a conversion between a bitstream representation of the video block and the video block using the final prediction samples.

[0293] In some embodiments of Method 2800S, an optical flow-based prediction refinement technique is first performed, and then an interweave prediction technique is performed. In some embodiments of Method 2800S, refined intermediate prediction samples having two different subblock partitioning patterns are first obtained using the optical flow-based prediction refinement technique, and the refined intermediate prediction samples are weighted-averaged using the interweave prediction technique to obtain a final prediction sample. In some embodiments of Method 2800S, an interweave prediction technique is first performed, and then an optical flow-based prediction refinement technique is performed. In some embodiments of Method 2800S, intermediate prediction samples having two different subblock partitioning patterns are weighted-averaged using the interweave prediction technique to obtain a refined intermediate prediction sample, and then the optical flow-based prediction refinement technique is performed on the refined intermediate prediction sample to obtain a final prediction sample. In some embodiments of any one or more of methods 2800N through 2800S, the optical flow based prediction refinement technique is a prediction refinement with optical flow (PROF) technique.

[0294] 28T is a flowchart of an example video processing method. Method 2800T includes obtaining (28192) refined intermediate prediction samples of a video block by performing an interweave prediction technique and a phase-variational affine sub-block motion compensation (PAMC) technique on the intermediate prediction samples of the video block, deriving (28194) final prediction samples from the refined intermediate prediction samples, and performing (28196) a conversion between a bitstream representation of the video block and the video block using the final prediction samples.

[0295] In some embodiments of Method 2800T, the PAMC technique is performed first, and then the interweave prediction technique is performed. In some embodiments of Method 2800T, refined intermediate prediction samples with two different subblock division patterns are first obtained using the interpolation method of the PAMC technique, and the refined intermediate prediction samples are weighted-averaged using the interweave prediction technique to obtain a final prediction sample. In some embodiments of Method 2800T, the interweave prediction technique is performed first, and then the PAMC technique is performed.

[0296] 28U is a flowchart of an example video processing method. Method 2800U includes obtaining (28202) refined intermediate prediction samples of a video block by performing an optical flow-based prediction refinement technique and a phase variational affine sub-block motion compensation (PAMC) technique on the intermediate prediction samples of the video block, deriving (28204) final prediction samples from the refined intermediate prediction samples, and performing (28206) a conversion between a bitstream representation of the video block and the video block using the final prediction samples. In some embodiments of method 2800U, the PAMC technique is performed first, and then the optical flow-based prediction refinement technique is performed. In some embodiments of method 2800U, refined intermediate prediction samples with two different subblock division patterns are first obtained using the interpolation method of the PAMC technique, and the refined intermediate prediction samples are then processed using the optical flow-based prediction refinement technique to obtain final prediction samples. In some embodiments of method 2800U, the optical flow-based prediction refinement technique is performed first, and then the PAMC technique is performed.

[0297] In this document, the term "video processing" may refer to video encoding (including transcoding), video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion of a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are collocated or spread differently within the bitstream, e.g., as specified by the syntax. For example, a macroblock may be coded with respect to a transformed coded error residual value, or may be coded using bits in a header and other fields within the bitstream.

[0298] It will be appreciated that the disclosed methods and techniques will be beneficial in embodiments of video encoders and / or decoders integrated into video processing devices such as smartphones, laptops, desktops, and similar devices by enabling the use of the techniques disclosed herein.

[0299] The following list of sections provides further mechanisms and embodiments that use the techniques disclosed in this document.

[0300] 1. A method of video processing, comprising, in converting between a video block and a bitstream representation of the video block, determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying the prediction sample P(x,y) at the position (x,y) as a function of a gradient in a first direction and / or a second direction estimated at the position (x,y) and a first motion displacement and / or a second motion displacement estimated for the position (x,y), and performing the conversion using a reconstructed sample value Rec(x,y) from the refined prediction sample P'(x,y).

[0301] 2. The method of claim 1, wherein the first direction and the second direction are perpendicular to each other.

[0302] 3. The method of any of paragraphs 1 or 2, wherein the first movement displacement is in a direction parallel to the first direction and the second movement displacement is in a direction parallel to the second direction.

[0303] 4. The method according to any one of clauses 1 to 3, wherein the gradient in the first direction is represented as Gx(x,y), the gradient in the second direction is represented as Gy(x,y), the first motion displacement is represented as Vx(x,y), the second motion displacement is represented as Vy(x,y), and P'(x,y) = P(x,y) + Gx(x,y) × Vx(x,y) + Gy(x,y) × Vy(x,y).

[0304] 5. The method according to any one of clauses 1 to 3, wherein the gradient in the first direction is denoted as Gx(x,y), the gradient in the second direction is denoted as Gy(x,y), the first motion displacement is denoted as Vx(x,y), the second motion displacement is denoted as Vy(x,y), and P'(x,y)=α(x,y)×P(x,y)+β(x,y)×Gx(x,y)×Vx(x,y)+γ(x,y)×Gy(x,y)×Vy(x,y), where α(x,y), β(x,y) and γ(x,y) are weighting values ​​at the position (x,y) and are integers or real numbers.

[0305] 6. The method according to any one of items 1 to 5, wherein the predicted sample P(x, y) is a one-way predicted sample at the position (x, y).

[0306] 7. The method according to any one of clauses 1 to 5, wherein the predicted sample P(x, y) is the final result of bi-prediction at the position (x, y).

[0307] 8. The method according to any one of clauses 1 to 5, wherein the predicted sample P(x,y) satisfies one of the following: - The result of multiple hypothesis inter prediction (inter prediction using three or more MVs) Affine prediction results, Intra prediction results, Intra-Block Copy (IBC) prediction results, Generated by triangular prediction mode (TPM), Inter-intra joint prediction results, Global inter-prediction results where the regions share the same motion model and parameters, Palette coding mode results in Inter-view prediction results in multiview or 3D video coding, · Inter-layer prediction results in scalable video coding, The result of the filtering operation.

[0308] The result of a refinement process to improve the accuracy of the prediction at said location (x,y).

[0309] 9. The method of any of clauses 1 to 8, wherein the reconstructed sample values ​​Rec(x,y) are further refined prior to use in the transformation.

[0310] Item 1 listed in Section 4 provides further example embodiments for Methods 1-9.

[0311] 10. A method of video processing, comprising the steps of determining a first displacement vector Vx(x,y) and a second displacement vector Vy(x,y) at a position (x,y) within a video block corresponding to an optical flow-based method of encoding the video block based on information from adjacent blocks or basic blocks, and performing a conversion between the video block and a bitstream representation of a current video block using the first displacement vector and the second displacement vector.

[0312] 11. The method of claim 10, wherein the adjacent blocks are spatially adjacent blocks.

[0313] 12. The method of claim 10, wherein the adjacent blocks are time-adjacent blocks.

[0314] 13. The method of any of clauses 10 to 12, wherein the neighboring blocks are selected based on a coding unit size, or a coding mode, or the locations of possible candidate blocks neighboring the video block.

[0315] 14. The method according to any of clauses 10 to 13, wherein the first displacement vector and the second displacement vector are calculated using a combination of motion vectors for a plurality of neighboring blocks.

[0316] 15. The method of any one of clauses 11 to 13, wherein the basic blocks have predetermined dimensions.

[0317] 16. The method according to clause 10, wherein Vx(x,y) and Vy(x,y) are determined as Vx(x,y) = a × (x-xc) + b(y-yc), Vy(x,y) = c × (x-xc) + d(y-yc), where (xc,yc) is the center position of the basic block of dimensions w × h that covers the position (x,y).

[0318] 17. The method according to clause 10, wherein Vx(x,y) and Vy(x,y) are determined as Vx(x,y) = Shift(a × (x - xc) + b (y - yc), n1), Vy(x,y) = Shift(c × (x - xc) + d (y - yc), n1), where n1 is an integer and (xc,yc) is the center position of the basic block of dimensions w × h that covers the position (x,y).

[0319] 18. The method of clause 10, wherein Vx(x,y) and Vy(x,y) are determined as Vx(x+1,y) = Vx(x,y) + a and Vy(x+1,y) = Vy(x,y) + c, and the parameters a and c are obtained based on information from the neighboring blocks or from a history-based storage device.

[0320] 19. The method according to any of clauses 10 to 18, wherein the precision of Vx(x,y) and Vy(x,y) is different from the precision of the motion vectors of the basic blocks.

[0321] Section 4, item 2 provides further examples of the embodiments of items 10-19.

[0322] 20. A method of video processing, comprising: determining a refined prediction sample P'(x,y) at a position (x,y) within the video block by modifying a prediction sample P(x,y) at the position (x,y), wherein a gradient in a first direction and a gradient in a second direction at the position (x,y) are determined based on the refined prediction sample P'(x,y) and a final prediction value determined from a residual sample value at the position (x,y); and performing the conversion using the gradient in the first direction and the gradient in the second direction.

[0323] 21. The method of claim 20, wherein the gradient in the first direction and the gradient in the second direction correspond to a horizontal gradient and a vertical gradient.

[0324] 22. The method of any of clauses 20 to 21, wherein the gradient in the first direction and the gradient in the second direction are derived from the result of an intermediate prediction.

[0325] 23. The method of any of clauses 20-21, wherein the gradient in the first direction and the gradient in the second direction are derived from one reference picture list for the video block, which is an affine-predicted bi-predictive video block.

[0326] Items 3 and 4 of Section 4 provide further example embodiments of the techniques described in Items 20-23.

[0327] 24. A method of video processing, comprising the steps of determining a reconstructed sample Rec(x,y) at a position (x,y) within a video block to be affine coded, refining Rec(x,y) using first and second displacement vectors and first and second gradients at the position (x,y) to obtain a refined reconstructed sample Rec'(x,y), and using the refined reconstructed sample to perform a conversion between the video block and a bitstream representation of a current video block.

[0328] 25. The method according to clause 24, wherein Rec'(x,y)=Rec(x,y)+Gx(x,y)×Vx(x,y)+Gy(x,y)×Vy(x,y), where Gx(x,y) and Gy(x,y) represent the first and second gradients, and Vx(x,y) and Vy(x,y) represent the first and second displacement vectors at the position (x,y).

[0329] 26. The method of clause 25, wherein Vx(x,y) and Vy(x,y) are derived at the sub-block level.

[0330] 27. The method of clause 25, wherein Vx(x,y) and Vy(x,y) are derived with 1 / M pel precision, where M is an integer.

[0331] 28. The method of any of clauses 1 to 27, wherein the method is applied as a result of the video block having a particular size and / or a particular coding mode.

[0332] 29. The method of claim 28, wherein the specific dimensions are 8x4.

[0333] 30. The method of clause 28, wherein the video block is a 4x4 single-prediction video block.

[0334] Items 5-12 of Section 4 provide examples of the embodiments described in Items 24-30.

[0335] 31. A method according to any of clauses 1 to 27, wherein the method is applied as a result of the video block having particular color components or a particular color format, or having a particular position within a video picture, or using a particular transformation type.

[0336] 32. The method of any of clauses 1 to 31, wherein the transforming includes generating the current block from the bitstream representation or generating the bitstream representation from the current block.

[0337] 33. A method of image processing comprising: a motion vector derived using the first one of the K subblock patterns and a motion vector derived using the L pattern; determining an offset value OL(x,y) at the location (x,y) based on P(x,y) and a difference between the motion vector derived using the first one of the K subblock patterns and the motion vector derived using the L pattern; determining a final predicted sample for the location (x,y) as a function of OL(x,y) and P(x,y); and performing the conversion using the final predicted sample.

[0338] 34. The method of clause 33, wherein K=2, L=1, and the final predicted sample is determined using P(x,y)+((O1(x,y)+1)>>1), where >> represents a binary shift operation.

[0339] 35. The method of clause 33, wherein K=2, L=1, and the final predicted sample is determined using P(x,y)+(O1(x,y)>>1), where >> denotes a binary shift operation.

[0340] 36. The method of any of clauses 33 to 35, wherein OL(x,y) is generated from P(x,y) after performing horizontal and vertical interpolation.

[0341] 37. The method of any of clauses 33 to 35, wherein motion compensation for the first pattern can use 1 / N pixel accuracy, and motion compensation for the L pattern can use 1 / ML pixel accuracy, where N and M are integers.

[0342] 38. The method of item 37, wherein N=16 and ML=32, 64, or 128.

[0343] 39. The method of item 37, wherein N=8 and ML=16, 32, 64, or 128.

[0344] 40. The method of claim 37, wherein N=4 and ML=8, 16, 32, 64, or 128.

[0345] 41. A video encoder or re-encoder having a processor configured to implement the method of any of clauses 1 to 40.

[0346] 42. A video decoder having a processor configured to implement the method of any of clauses 1 to 40.

[0347] 43. A computer readable medium having code for implementing the method of any of clauses 1 to 40.

[0348] 33 is a block diagram illustrating an example of a video processing system 2100 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 2100. System 2100 may include an input 2102 that receives video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 2102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0349] System 2100 may include an encoding component 2104 that may implement various coding or encoding methods described herein. The encoding component 2104 may reduce the average bitrate of the video from input 2102 to the output of the encoding component 2104, generating an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression techniques or video transcoding techniques. The output of the encoding component 2104 may be stored or transmitted via a communication connection, as represented by component 2106. The stored or communicated bitstream (or encoded) representation of the video received at input 2102 may be used by component 2108 to generate pixel values ​​or displayable video that are sent to display interface 2110. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Also, while certain video processing operations may be referred to as “encoding” operations or tools, it is understood that the encoding tools or operations are used in an encoder, and corresponding decoding tools or operations that reverse the results of the encoding are performed in a decoder.

[0350] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be embodied in a variety of electronic devices, such as, for example, mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0351] Some embodiments of the disclosed technology include making a decision to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder will use or implement that tool or mode in processing blocks of video, but will not necessarily modify the resulting bitstream based on the use of that tool or mode. That is, conversion from blocks of video to a bitstream representation of video will use the video processing tool or mode when it is enabled based on the decision. In another example, when a video processing tool or mode is enabled, a decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, conversion from a bitstream representation of video to blocks of video will be performed using the video processing tool or mode that was enabled based on the decision.

[0352] Some embodiments of the disclosed technology include making a decision to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder will not use that tool or mode in converting blocks of video into a video bitstream. In another example, when a video processing tool or mode is disabled, a decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was disabled based on the decision.

[0353] Figure 34 is a block diagram illustrating an example of a video encoding system 100 that can utilize the techniques of this disclosure. As shown in Figure 34, the video encoding system 100 can include a source device 110 and a destination device 120. The source device 110 generates encoded video data and can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and can be referred to as a video decoding device. The source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0354] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may comprise one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a series of bits forming a coded representation of the video data. The bitstream may include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted via the I / O interface 116 directly over the network 130a to the destination device 120. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0355] The destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0356] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120 configured to interface with an external display device.

[0357] Video encoder 114 and video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other current and / or future standards.

[0358] FIG. 35 is a block diagram illustrating an example of a video encoder 200, which may be video encoder 114 in system 100 shown in FIG.

[0359] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. In the example of FIG. 35, video encoder 200 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video encoder 200. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0360] The functional components of the video encoder 200 may include a division unit 201, a prediction unit 202, which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0361] In other examples, video encoder 200 may include more, fewer, or different functional components. In one example, prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0362] Also, some components, such as the motion estimation unit 204 and the motion compensation unit 205, although shown separately in the example of Figure 35 for illustrative purposes, may be highly integrated.

[0363] Division unit 201 may divide a picture into one or more video blocks. Video encoder 200 and video decoder 300 may support a variety of video block sizes.

[0364] The mode selection unit 203 may select one of a plurality of coding modes, intra or inter, based on, for example, an error result, and provide the resulting intra- or inter-coded block to a residual generation unit 207, which generates residual block data, and to a reconstruction unit 212, which reconstructs the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combination of intra and inter predication (CIIP) mode, in which prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 may also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision) in the case of inter prediction.

[0365] To perform inter prediction on a current video block, motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 may determine a prediction video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture associated with the current video block.

[0366] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice, for example.

[0367] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search reference pictures in list 0 or list 1 for reference video blocks for the current video block. Motion estimation unit 204 may then generate a reference index that points to the reference picture in list 0 or list 1 that contains the reference video block, and a motion vector that indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Based on the reference video block indicated by the motion information of the current video block, motion compensation unit 205 may generate a prediction video block for the current block.

[0368] In another example, motion estimation unit 204 may perform bidirectional prediction on the current video block, where motion estimation unit 204 may search reference pictures in list 0 for a reference video block for the current video block and may also search reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 204 may then generate reference indices that point to the reference pictures in lists 0 and 1 that contain the reference video blocks, and motion vectors that indicate spatial displacements between the reference video blocks and the current video block. Motion estimation unit 204 may output the reference indices and the motion vector for the current video block as motion information for the current video block. Based on the reference video blocks indicated by the motion information for the current video block, motion compensation unit 205 may generate a prediction video block for the current block.

[0369] In some examples, the motion estimation unit 204 may output a complete set of motion information for the decoding process of the decoder.

[0370] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video block. Rather, motion estimation unit 204 may signal motion information for the current video block by reference to motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0371] In one example, motion estimation unit 204 may point to a value within a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0372] In another example, motion estimation unit 204 may identify another video block and a motion vector difference (MVD) within a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the pointed-to video block. Video decoder 300 may use the motion vector of the pointed-to video block and the motion vector difference to determine the motion vector of the current video block.

[0373] As mentioned above, video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0374] Intra prediction unit 206 may perform intra prediction on the current video block. When intra prediction unit 206 performs intra prediction on the current video block, intra prediction unit 206 may generate predictive data for the current video block based on decoded samples of other video blocks in the same picture. The predictive data for the current video block may include a predictive video block and various syntax elements.

[0375] Residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks that correspond to different sample components of the samples in the current video block.

[0376] In other examples, for example, in skip mode, residual data for the current video block may not exist and residual generation unit 207 may not perform a subtraction operation.

[0377] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0378] After transform processing unit 208 generates the transform coefficient video block for the current video block, quantization unit 209 may quantize the transform coefficient video block for the current video block based on one or more quantization parameter (QP) values ​​for the current video block.

[0379] Inverse quantization unit 210 and inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. Reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by prediction unit 202 to generate a reconstructed video block for the current block that is stored in buffer 213.

[0380] After reconstruction unit 212 reconstructs the video blocks, a loop filtering operation may be performed to reduce video blocking artifacts in the video blocks.

[0381] An entropy encoding unit 214 may receive data from other functional components of the video encoder 200. Once the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0382] FIG. 36 is a block diagram illustrating an example of a video decoder 300, which may be video decoder 124 in system 100 shown in FIG.

[0383] Video decoder 300 may be configured to perform any or all of the techniques described in this disclosure. In the example of FIG. 36, video decoder 300 includes multiple functional components. The techniques described in this disclosure may be shared among various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.

[0384] In the example of Figure 36, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. Video decoder 300 may, in some examples, perform a decoding pass that is generally inverse to the encoding pass described with respect to video encoder 200 (Figure 35).

[0385] An entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-encoded video data, and from the entropy-decoded video data, a motion compensation unit 302 may determine motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 may determine such information by, for example, implementing AMVP and merge mode.

[0386] The motion compensation unit 302 may optionally perform interpolation based on an interpolation filter to generate the motion-compensated blocks. An identifier for the interpolation filter used with sub-pixel precision may be included in the syntax element.

[0387] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of the reference block using the interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 according to the received syntax information and generate the predictive block using the interpolation filters.

[0388] The motion compensation unit 302 may use some of the syntax information to determine the size of the blocks used to encode the frames and / or slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is divided, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence.

[0389] The intra prediction unit 303 may form a prediction block from spatially neighboring blocks using an intra prediction mode, e.g., received in the bitstream. The inverse quantization unit 303 inverse quantizes, or dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0390] A reconstruction unit 306 may add the residual block with a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303 to form a decoded block. If desired, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for later motion compensation / intra prediction and also generates decoded video for presentation on a display device.

[0391] In this patent document, the term "sample" or "samples" may refer to one or more samples of a video block. From the foregoing, it will be understood that, although specific embodiments of the presently disclosed technology have been described herein for purposes of illustration, various modifications may be made without departing from the scope of the invention. Accordingly, the presently disclosed technology is not to be limited except as by the appended claims.

[0392] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document, including the structures disclosed herein and their structural equivalents, can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in combinations of one or more of these. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by or for controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter producing a machine-readable propagated signal, or a combination of one or more of these. The term "data processing apparatus" encompasses any apparatus, device, and machine that processes data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus can include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of these. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver device.

[0393] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program at hand, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or on multiple computers, either collocated or distributed across multiple locations and interconnected by a communications network.

[0394] The processes and logic flows described in this document may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. These processes and logic flows may also be performed by, and apparatus may be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0395] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or is operatively coupled to receive data from or transfer data to the mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0396] While this patent document contains numerous details, these should not be construed as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular technology. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operating in a particular combination, and even initially claimed as such, in some cases one or more features from a claimed combination may be removed from the combination, or the claimed combination may be subject to subcombinations or variations of the subcombination.

[0397] Similarly, although the figures may depict operations in a particular order, this should not be understood as requiring that those operations be performed in the particular order or sequence shown, or that all of the operations shown be performed, to achieve desired results. Additionally, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0398] Only a few implementations and examples have been described, and other implementations, extensions and variations may be made based on what is described and illustrated in this patent document.

Claims

1. 1. A method for processing video data, comprising: determining at least one control point motion vector for an affine-coded video block of the video; determining a motion vector for a sub-block having a position (x, y) of the affine-coded video block based on the at least one control point motion vector; determining a first motion displacement Vx(x,y) in a first direction and a second motion displacement Vy(x,y) in a second direction for the position (x,y) based on the at least one control point motion vector; determining a first gradient component Gx(x,y) in the first direction and a second gradient component Gy(x,y) in the second direction for the location (x,y); determining a refined prediction sample P′(x,y) for the position (x,y) by modifying a prediction sample P(x,y) derived for the position (x,y) with the first gradient component Gx(x,y), the second gradient component Gy(x,y), the first motion displacement Vx(x,y), and the second motion displacement Vy(x,y), wherein the prediction sample P(x,y) is derived based on the motion vector for the sub-block; using the refined prediction samples P′(x,y) to convert between the affine-coded video block and the video bitstream; and the precision of the first motion displacement Vx(x, y) and the second motion displacement Vy(x, y) is different from the precision of the motion vector for the sub-block; Vx(x,y)=a×(x−xc)+b×(y−yc), Vy(x,y)=c×(x−xc)+d×(y−yc), (xc,yc) is based on the center position or size of the sub-block, and a, b, c, and d are affine parameters; a, b, c, and d may be derived from the control point motion vector, the width (W) of the affine coded video block, and the height (H) of the affine coded video block; a, b, c and d are shifted, method.

2. The method of claim 1 , wherein the color component of the affine-coded video block is a luma component.

3. The method of claim 1 or 2, wherein the accuracy of the first motion displacement Vx(x,y) and the second motion displacement Vy(x,y) is 1 / 32 pixel accuracy.

4. The method of claim 1 , wherein the precision of the motion vectors for the sub-blocks is 1 / 16 pixel precision.

5. The method of claim 1 , wherein Vx(x, y) and Vy(x, y) are determined based on at least the position (x, y) and a center position of the sub-block.

6. The method of claim 1 , wherein Vx(x,y) and Vy(x,y) are determined based on at least the position (x,y) and the size of the sub-block.

7. 2. The method of claim 1, wherein c=-b and d=a in response to the affine-coded video block being coded using a four-parameter affine mode. 【Request 8】 【Number 1】 and music video 0 , mv 1 , and mv 2 is the control point motion vector, a motion vector component having a superscript h indicates that the motion vector component is in the first direction; another motion vector component having a superscript v indicates that the another motion vector component is in the second direction; the first direction is perpendicular to the second direction; The method of claim 1.

9. The method according to any one of claims 1 to 8, wherein the method is used for the luma component and not for the chroma components.

10. 10. A method according to any one of claims 1 to 9, wherein the method is used for affine modes and not for non-affine modes.

11. The method of claim 1 , wherein decoder-side motion vector refinement and / or bidirectional optical flow methods are not applied to the affine-coded video blocks.

12. The method of claim 1 , wherein the transforming comprises encoding the affine-coded video blocks into the bitstream.

13. The method of claim 1 , wherein the transforming comprises decoding the affine-coded video blocks from the bitstream.

14. 1. An apparatus for processing video data comprising a processor and a non-transitory memory having instructions that, when executed by the processor, cause the processor to: determining at least one control point motion vector for an affine-coded video block of the video; determining a motion vector for a sub-block having a position (x, y) of the affine-coded video block based on the at least one control point motion vector; determining a first motion displacement Vx(x,y) in a first direction and a second motion displacement Vy(x,y) in a second direction for the position (x,y) based on the at least one control point motion vector; determining a first gradient component Gx(x,y) in the first direction and a second gradient component Gy(x,y) in the second direction for the location (x,y); modifying the derived prediction sample P(x,y) for the position (x,y) with the first gradient component Gx(x,y), the second gradient component Gy(x,y), the first motion displacement Vx(x,y), and the second motion displacement Vy(x,y) to generate a refined prediction sample P′(x,y) for the position (x,y), wherein the prediction sample P(x,y) is derived based on the motion vector for the sub-block; converting between the affine-coded video block and the video bitstream using the refined prediction samples P′(x,y); the precision of the first motion displacement Vx(x, y) and the second motion displacement Vy(x, y) is different from the precision of the motion vector for the sub-block; Vx(x,y)=a×(x−xc)+b×(y−yc), Vy(x,y)=c×(x−xc)+d×(y−yc), (xc,yc) is based on the center position or size of the sub-block, and a, b, c, and d are affine parameters; a, b, c, and d may be derived from the control point motion vector, the width (W) of the affine coded video block, and the height (H) of the affine coded video block; a, b, c and d are shifted, Device.

15. A non-transitory computer-readable storage medium having stored thereon instructions that cause a processor to: determining at least one control point motion vector for an affine-coded video block of the video; determining a motion vector for a sub-block having a position (x, y) of the affine-coded video block based on the at least one control point motion vector; determining a first motion displacement Vx(x,y) in a first direction and a second motion displacement Vy(x,y) in a second direction for the position (x,y) based on the at least one control point motion vector; determining a first gradient component Gx(x,y) in the first direction and a second gradient component Gy(x,y) in the second direction for the location (x,y); modifying the derived prediction sample P(x,y) for the position (x,y) with the first gradient component Gx(x,y), the second gradient component Gy(x,y), the first motion displacement Vx(x,y), and the second motion displacement Vy(x,y) to generate a refined prediction sample P′(x,y) for the position (x,y), wherein the prediction sample P(x,y) is derived based on the motion vector for the sub-block; converting between the affine-coded video block and the video bitstream using the refined prediction samples P′(x,y); the precision of the first motion displacement Vx(x, y) and the second motion displacement Vy(x, y) is different from the precision of the motion vector for the sub-block; Vx(x,y)=a×(x−xc)+b×(y−yc), Vy(x,y)=c×(x−xc)+d×(y−yc), (xc,yc) is based on the center position or size of the sub-block, and a, b, c, and d are affine parameters; a, b, c, and d may be derived from the control point motion vector, the width (W) of the affine coded video block, and the height (H) of the affine coded video block; a, b, c and d are shifted, A computer-readable storage medium.

16. A method for storing a video bitstream, comprising: determining at least one control point motion vector for an affine-coded video block of the video; determining a motion vector for a sub-block having a position (x, y) of the affine-coded video block based on the at least one control point motion vector; determining a first motion displacement Vx(x,y) in a first direction and a second motion displacement Vy(x,y) in a second direction for the position (x,y) based on the at least one control point motion vector; determining a first gradient component Gx(x,y) in the first direction and a second gradient component Gy(x,y) in the second direction for the location (x,y); determining a refined prediction sample P′(x,y) for the position (x,y) by modifying a prediction sample P(x,y) derived for the position (x,y) with the first gradient component Gx(x,y), the second gradient component Gy(x,y), the first motion displacement Vx(x,y), and the second motion displacement Vy(x,y), wherein the prediction sample P(x,y) is derived based on the motion vector for the sub-block; generating the bitstream using the refined prediction samples P′(x,y); storing the bitstream on a non-transitory computer-readable recording medium; and the precision of the first motion displacement Vx(x, y) and the second motion displacement Vy(x, y) is different from the precision of the motion vector for the sub-block; Vx(x,y)=a×(x−xc)+b×(y−yc), Vy(x,y)=c×(x−xc)+d×(y−yc), (xc,yc) is based on the center position or size of the sub-block, and a, b, c, and d are affine parameters; a, b, c, and d may be derived from the control point motion vector, the width (W) of the affine coded video block, and the height (H) of the affine coded video block; a, b, c and d are shifted, method.

Citation Information

Patent Citations

  • Multiple predictor candidates for motion compensation

    WO2019002215A1

  • Systems, apparatus and methods for inter prediction refinement with optical flow

    WO2020163319A1