Determination Based on Sub-Region for Refinement of Motion Information

By refining motion information using sub-regions and advanced techniques like optical flow and DMVR, the method addresses the bandwidth challenges in current video compression technologies, achieving improved coding efficiency and compression performance.

JP7697765B2Active Publication Date: 2025-06-24DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024010643
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-08
Filing Date
2024-01-29
Publication Date
2025-06-24
Estimated Expiration
2040-05-18

AI Technical Summary

Technical Problem

Current video compression technologies consume a significant portion of bandwidth due to the large size of digital video files, and existing video coding standards struggle to efficiently manage motion information across different video blocks.

Method used

The proposed method involves refining motion information using sub-regions within video blocks through techniques such as optical flow-based methods, decoder-side motion vector refinement (DMVR), and bidirectional optical flow (BIO), allowing for more precise conversion between video blocks and their coded representations.

Benefits of technology

This approach enhances video coding efficiency by improving motion information refinement, leading to reduced bandwidth requirements and better compression performance across various video coding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697765000069
    Figure 0007697765000069
  • Figure 0007697765000070
    Figure 0007697765000070
  • Figure 0007697765000071
    Figure 0007697765000071
Patent Text Reader

Abstract

To provide a device, a system, and a method for video processing.SOLUTION: A video processing method according to an embodiment includes the steps of determining that motion information of the current video block is refined using an optical flow-based method in which at least one motion vector offset is derived for a region within the current video block for conversion between the current video block of the video and the coded representation of the video, clipping at least one motion vector offset to the range [-N,M], where N and M are integers, and performing transformation on the basis of the at least one clipped motion vector offset.SELECTED DRAWING: Figure 12A
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Cross - Reference to Related Applications] This application is a divisional application of Japanese Patent Application No. 2021 - 564118, filed on October 27, 2021, based on International Patent Application No. PCT / CN2020 / 090802, filed on May 18, 2020. The above - mentioned international patent application claims priority and the benefit thereof with respect to International Patent Application No. PCT / CN2019 / 0871937, filed on May 16, 2019; International Patent Application No. PCT / CN2019 / 090037, filed on June 4, 2019; International Patent Application No. PCT / CN2019 / 090903, filed on June 12, 2019; International Patent Application No. PCT / CN2019 / 093616, filed on June 28, 2019; International Patent Application No. PCT / CN2019 / 093973, filed on June 29, 2019; International Patent Application No. PCT / CN2019 / 094282, filed on July 1, 2019; International Patent Application No. PCT / CN2019 / 104489, filed on September 5, 2019; and International Patent Application No. PCT / CN2019 / 116757, filed on November 8, 2019. The entire contents of the above - mentioned patent applications are incorporated herein by reference in their entirety.

[0002] [Technical Field] This patent document relates to video processing technologies, devices, and systems.

Background Art

[0003] Despite the progress of video compression, digital video still occupies the maximum bandwidth usage on the Internet and other digital communication networks. As the number of user devices capable of receiving and displaying video increases, the bandwidth demand for digital video utilization is expected to continue to grow.

Summary of the Invention

[0004] Devices, systems, and methods related to digital video coding including motion information refinement based on sub-regions are described. The described methods may be applicable to both existing video coding standards (e.g., HEVC (High Efficiency Video Coding)) and future video coding standards (e.g., VVC (Versatile Video Coding)) or codecs.

[0005] In one representative aspect, the disclosed technology is a method of video processing, determining that motion information of a current video block is refined using an optical flow-based method in which at least one motion vector offset is derived for a region within the current video block for conversion between the current video block of the video and the coded representation of the video; clipping the at least one motion vector offset to a range [-N, M], where N and M are integers based on a rule; performing the conversion based on the at least one motion vector offset and may be used to provide a method having.

[0006] In another aspect, the disclosed technology is a method of video processing, selecting, as refined motion vectors, motion information equal to a result of applying a similarity matching function using one or more motion vector differences associated with a current video block of a video during a decoder-side motion vector refinement (DMVR) operation used to refine motion information; performing a conversion between the current video block of the video and the coded representation of the video using the refined motion vectors and may be used to provide a method having.

[0007] In yet another exemplary aspect, a method of video processing is disclosed. The method is Deriving motion information associated with the current video block for conversion between the current video block of the video and the coded representation of the video; Applying a refinement operation to the current video block including a first sub-region and a second sub-region according to a rule, the rule allowing the first sub-region and the second sub-region to have different motion information from each other by the refinement operation; Performing the conversion using the refined motion information of the current video block; and including.

[0008] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes: Deriving motion information associated with the current video block for conversion between the current video block of the video and the coded representation of the video; Determining the applicability of a refinement operation using bidirectional optical flow (BIO) based on the output of decoder-side motion vector refinement (DMVR) used to refine the motion information for sub-regions of the current video block; Performing the conversion based on the determination; and including.

[0009] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes: Deriving motion information associated with a current video block coded in merge mode by motion vector difference (MMVD) including a motion vector representation including a distance table defining a distance between two motion candidates; Applying decoder-side motion vector refinement (DMVR) to the current video block to refine the motion information according to a rule specifying how to refine the distance used for the MMVD; Performing a conversion between the current video block and the coded representation of the video; and including.

[0010] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining the applicability of bidirectional optical flow (BDOF) for samples or sub-blocks of the current video block of the video, based on the derived motion information according to rules, wherein the derived motion information is refined using spatial and / or temporal gradients; executing a conversion between the current video block and the coded representation of the video based on the determination; and including.

[0011] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes deriving a motion vector difference (MVD) for conversion between the current video block of the video and the coded representation of the video; applying a clipping operation to the derived motion vector difference to generate a clipped motion vector difference; calculating the cost of the clipped motion vector difference using a cost function; determining, according to rules, that a bidirectional optical flow (BDOF) operation in which the derived motion vector difference is refined using spatial and / or temporal gradients is not allowed based on at least one of the derived motion vector difference, the clipped motion vector difference, or the cost; executing the conversion based on the determination; and including.

[0012] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes deriving a motion vector difference for conversion between the current video block of the video and the coded representation of the video; refining the derived motion vector difference based on one or more motion vector refinement tools and candidate motion vector differences (MVDs); performing the conversion using the refined motion vector difference and including

[0013] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes limiting a set of candidate derived motion vector differences associated with a current video block of a video, the derived motion vector differences being used for a refinement operation that refines motion information associated with the current video block, the limiting step; performing a conversion between the current video block and the coded representation of the video using the derived motion vector differences as a result of the limiting step and including

[0014] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes applying a clipping operation using a set of clipping parameters determined based on utilization of the video unit and / or coding tools in the video unit according to rules during a bidirectional optical flow (BDOF) operation used to refine motion information associated with a current video block of a video unit of a video; performing a conversion between the current video block and the coded representation of the video and including

[0015] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes applying a clipping operation according to rules to clip an x component and / or a y component of a motion vector difference (vx, vy) during a refinement operation used to refine motion information associated with a current video block of a video; performing a conversion between the current video block and the coded representation of the video using the motion vector difference and including The rule defines to convert the motion vector difference to a value taking the form of zero or K before or after the clipping operation, where m is an integer. m

[0016] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes selecting a search region used to derive or refine motion information during a decoder-side motion derivation operation or a decoder-side motion refinement operation according to rules for conversion between a current video block of a video and a coded representation of the video; performing the conversion based on the derived or refined motion information.

[0017] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with the current video block by using a search region used in refinement including an integer position of a best match for conversion between the current video block of the video and a coded representation of the video; performing the conversion by using the refined motion vector difference. Applying the DMVR operation includes deriving a sub-pel motion vector difference (MVD) according to rules.

[0018] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with a current video block of a video unit of the video; performing a conversion between the current video block and a coded representation of the video by using the refined motion vector difference. ​​​​The application of the DMVR operation includes determining whether to permit or not permit sub-pel motion vector difference (MVD) derivation according to the use of bidirectional optical flow (BDOF) for the video unit.

[0019] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes deriving a motion vector difference of a first sample of a current video block of a video during a refinement operation using optical flow; determining a motion vector difference of a second sample based on the derived motion vector difference of the first sample; and performing a conversion between the current video block and a coded representation of the video based on the determination.

[0020] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes deriving a prediction refinement sample by applying a refinement operation video to a current video block of the video using bidirectional optical flow (BDOF); determining the applicability of a clipping operation to clip the derived prediction refinement sample to a predetermined range [-M, N], where M and N are integers; performing a conversion between the current video block and a coded representation of the video; and

[0021] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining a coding group size for a current video block of a video, the current video block including a first coding group and a second coding group that are coded using different residual coding modes such that the first coding group and the second coding group are aligned according to a rule; ​​performing a conversion between the current video block and the coded representation of the video based on the determination; including.

[0022] In yet another exemplary aspect, other methods of video processing are disclosed. The method includes determining the applicability of a predictive refinement optical flow (PROF) tool in which motion information is refined using optical flow based on coded information and / or decoded information associated with the current video block for conversion between the current video block of the video and the coded representation of the video; performing the conversion based on the determination; including.

[0023] In yet another representative aspect, the above method is embodied in the form of processor-executable code and stored in a computer-readable program medium.

[0024] In yet another representative aspect, a device configured or operable to perform the above method is disclosed. The device may include a processor programmed to implement this method.

[0025] In yet another representative aspect, a video decoder device may implement the methods described herein.

[0026] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the specification, and the claims.

Brief Description of the Drawings

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5a

Figure 5b

Figure 6

Figure 7a

Figure 7b

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 12E

Figure 12F

Figure 13A

Figure 13B

Figure 13C

Figure 13D

Figure 13E

Figure 13F

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 15

Figure 16

Modes for Carrying Out the Invention

[0028] The technologies and devices disclosed herein provide motion information refinement. Some implementations of the disclosed technologies relate to motion information refinement based on sub-regions. Some implementations of the disclosed technologies may be applied to motion compensation in video encoding and decoding.

[0029] Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created the H.262 / MPEG-2 Video, H264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, using temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, the JVET (Joint Video Exploration Team) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been introduced by the JVET and placed in the reference software named JEM (Joint Exploration Model). In April 2018, the JVET (Joint Video Expert Team) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MEPG) was formed to study the VVC standard with the goal of a 50% bitrate reduction compared to HEVC.

[0030] 1.1. Pattern-Matched Motion Vector Derivation The pattern matched motion vector derivation (PMMVD) mode is a special merge mode based on the Frame-Rate Up Conversion (FRUC) technique. According to this mode, the motion information of the block is not signaled but is derived at the decoder side.

[0031] The FRUC flag is signaled for the CU when its merge flag is true. When the FRUC flag is false, the merge index is signaled and the regular marge mode is used. When the FRUC flag is true, an additional FRUC mode flag is signaled to indicate which method (such as bilateral matching or template matching) should be used to derive the motion information for that block.

[0032] On the encoder side, the decision on whether to use the FRUC merge mode for the CU is based on the RD cost selection as is done for normal merge candidates. That is, both of the two matching modes (bilateral matching and template matching) are checked for the CU by using the RD cost selection. The one that results in the minimum cost is further compared with other CU modes. If the FRUC matching mode is the most efficient mode, the FRUC flag is set to true for that CU and the relevant matching mode is used.

[0033] The motion derivation process in FRUC merge mode has two steps. Motion search at the CU level is first performed, followed by motion refinement at the sub-CU level. At the CU level, an initial motion vector is derived for the entire CU based on bilateral matching or template matching. First, a list of MC candidates is generated, and the candidate that results in the minimum matching cost is selected as the starting point for further CU level refinement. Then, local search based on bilateral matching or template matching around the starting point is performed, and the MV that results in the minimum matching cost is regarded as the MV for the entire CU. After that, the motion information is further refined at the sub-CU level using the derived CU motion vector as the starting point.

[0034] For example, the following derivation process is performed for the motion information of a W×H CU. In the first stage, the MV for the entire W×H CU is derived. In the second stage, the CU is further divided into M×M sub-CUs. The value of M is calculated as seen in (16), and D is a predefined split depth that is set to 3 by default in JEM. Then, the MV for each sub-CU is derived.

Number

[0035] As shown in FIG. 1, bilateral matching is used to derive the motion information of the current CU by finding the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures. Under the assumption of a continuous motion trajectory, the motion vectors MV0 and MV1 indicating the two reference blocks should be proportional to the temporal distances between the current picture and the two reference pictures, i.e., TD0 and TD1. As a special case, when the current picture is temporally between the two reference pictures and the temporal distances from the current picture to the two reference pictures are the same, the bilateral matching results in a mirror-based bidirectional MV.

[0036] As shown in Figure 2, template matching is used to derive the motion information of the current CU by finding the closest match between a template (blocks above and / or to the left of the current CU) within the current picture and a block (of the same size as the template) within the reference picture. Except for the above-mentioned FRUC merge mode, template matching is also applied to the AMVP mode. In JEM, like in HEVC, AMVP has two candidates. A new candidate is derived by the template matching method. If the newly derived candidate by template matching is different from the first existing AMVP candidate, it is inserted at the very beginning of the AMVP candidate list, and then the list size is set to 2 (i.e., the second existing AMVP candidate is removed). When applied to the AMVP mode, only CU-level search is applied.

[0037] CU Level MV Candidate Set The set of MV candidates at the CU level consists of i) the original AMVP candidates when the current CU is in the AMVP mode, ii) all merge candidates, iii) some MVs in the interpolated MV field introduced in Section 2.1.1.3, iv) the motion vectors above and to the left and is composed of.

[0038] When using bilateral matching, each valid MV of the merge candidates is used as input to generate MV pairs under the assumption of bilateral matching. For example, one valid MV of the merge candidates is (MVa, refa) in reference list A. Then, the reference picture refb of the corresponding bilateral MV is found in the other reference list B such that refa and refb are on different sides of the current picture in terms of time. If such a refb is not available in reference list B, refb is determined as a reference different from refa and its time distance to the current picture is the minimum within list B. After refb is determined, MVb is derived by scaling MVa based on the time distances between the current picture and refa, refb.

[0039] Four MVs from the interpolated MV field are also added to the CU-level candidate list. More specifically, the interpolated MVs at the positions (0,0), (W / 2,0), (0,H / 2), and (W / 2,H / 2) of the current CU are added.

[0040] When FRUC is applied in AMVP mode, the original AMVP is also added to the CU-level MV candidate set.

[0041] At the CU level, up to 15 MVs of the AMVP CU and up to 13 MVs of the merge CU are added to the candidate list.

[0042] Sub - CU Level MV Candidate Set The MV candidate set at the sub-CU level consists of i) the MVs determined from the CU-level search, ii) the upper, left, upper-left, and upper-right adjacent MVs, iii) the scaled versions of the MVs at the same position from the reference pictures iv) up to 4 ATMVP candidates, v) up to 4 STMVP candidates and is composed of.

[0043] The scaled MV from the reference picture is derived as follows. All reference pictures in both lists are traversed. The MVs at the same position of the sub-CUs in the reference picture are scaled according to the reference of the start CU level MV.

[0044] The ATMVP and STMVP candidates are limited to the first four candidates.

[0045] At the sub-CU level, up to 17 MVs are added to the candidate list.

[0046] Generation of Interpolated MV Field Before coding a frame, the interpolated motion field is generated for the entire picture based on unilateral ME. Then, the motion field may be used later as a CU level or sub-CU level MV candidate.

[0047] First, the motion field of each reference picture in both reference lists is traversed at the 4×4 block level. For each 4×4 block, if the motion associated with the block passes through a 4×4 block in the current picture (shown in Figure 3 which shows an example of unilateral ME in FRUC) and the block has not been assigned any interpolated motion, the motion of the reference block is scaled according to the temporal distances TD0 and TD1 to match the current picture (similar to that of MV scaling of TMVP in HEVC), and the scaled motion is assigned to that block in the current frame. If the scaled MV cannot be assigned to a 4×4 block, the motion of that block is marked as unavailable in the interpolated motion field.

[0048] Interpolation and Matching Cost When the motion vector indicates a fractional sample position, motion-compensated interpolation is required. To suggest complexity, bilinear interpolation is used for both bilateral matching and template matching instead of regular 8-tap HEVC interpolation.

[0049] The calculation of the matching cost is slightly different in different steps. When selecting a candidate from the candidate set at the CU level, the matching cost is the absolute sum difference (SAD) of bilateral matching or template matching. After the starting MV is determined, the matching cost C of bilateral matching in the sub-CU level search is calculated as follows.

Number

[0050] In the FRUC mode, the MV is derived by using only the luma samples. The derived motion will be used for both luma and chroma for MC inter prediction. After the MV is determined, the final MC is performed using an 8-tap interpolation filter for luma and a 4-tap interpolation filter for chroma.

[0051] MV Refinement MV refinement is pattern-based MV search based on the criteria of bilateral matching cost or template matching cost. In JEM, two search patterns are supported. They are the unrestricted center-biased diamond search (UCBDS) and adaptive cross search for MV refinement at the CU level and sub-CU level, respectively. For MV refinement at both the CU and sub-CU levels, the MV is directly searched with a quarter luma sample MV accuracy, followed by an eighth luma sample MV refinement. The search range for MV refinement at the CU and sub-CU steps is set equal to 8 luma samples.

[0052] Selection of Prediction Direction in Template Matching FRUC Merge Mode In the bilateral matching merge mode, the motion information of the CU is derived based on the closest match between two blocks along the motion trajectory of the current CU in two different reference pictures, so that dual prediction is always applied. There is no such restriction in the template matching merge mode. In the template matching merge mode, the encoder can select from single prediction from list 0, single prediction from list 1, or dual prediction for the CU. The selection is based on the template matching cost as follows: When costBi <= factor × min(cost0, cost1), dual prediction is used; otherwise, when cost0 <= cost1, single prediction from list 0 is used; otherwise, single prediction from list 1 is used.

[0053] Here, cost0 is the SAD of list 0 template matching, cost1 is the SAD of list 1 template matching, and costBi is the SAD of dual prediction template matching. The value of factor is equal to 1.25, that is, the selection process is biased towards dual prediction.

[0054] Inter prediction direction selection is only applied to the template matching process at the CU level.

[0055] 1.2. Intra and Inter Composite Prediction In JVET-L0100, multi-hypothesis prediction has been proposed, and intra and inter composite prediction is one way to generate multiple hypotheses.

[0056] When multiple hypothesis prediction is applied to improve the intra mode, the multiple hypothesis prediction combines one intra prediction and one merged indexed prediction. In the merge CU, one flag is signaled for the merge mode to select the intra mode from the intra candidate list if the flag is true. For the luma component, the intra candidate list is derived from four intra prediction modes including the DC, planar, horizontal, and vertical modes, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is removed from the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra list mode. One intra prediction mode selected by the intra mode index and one merged indexed prediction selected by the merge index are combined using weighted average. For the chroma component, DM is always applied regardless of the extra signaling. The weights for combining the predictions are described as follows. Equal weights are applied when the DC or planar mode is selected, or when the CB width or height is less than 4. For a CB where the CB width and height are 4 or more, when the horizontal / vertical mode is selected, one CB is first divided into four equal-area regions in the vertical / horizontal direction. Let i be from 1 to 4, and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5), and (w_intra4, w_inter4) = (2, 6), then (w_intra i , w_inter iEach set of weights represented as ( ) is to be applied to the corresponding region. (w_intra1, w_inter1) is for the region closest to the reference sample, and (w_intra4, w_inter4) is for the region farthest from the reference sample. Then, the combined prediction can be calculated by summing the two weighted predictions and performing a 3-bit right shift. Further, the intra prediction mode for the predictor's intra hypothesis can be saved for reference to the next adjacent CU.

[0057] 1.3. Bidirectional Optical Flow BIO is also known as BDOF (Bi Directional Optical Flow). In BIO, first, motion compensation is performed to generate the first prediction (in each prediction direction) of the current block. The first prediction is used to derive the spatial gradient, temporal gradient, and optical flow of each sub-block / pixel within the block, which are then used to generate the second prediction, i.e., the final prediction of the sub-block / pixel. The details are described as follows.

[0058] Bidirectional Optical flow (BIO) is a sample-level motion refinement that is performed in addition to block-level motion compensation for bi-prediction. Sample-level motion refinement does not use signaling.

[0059] I (k) be the luma value from the reference k (k = 0, 1) after block motion compensation, then ∂I (k) / ∂x, ∂I (k) / ∂y are the horizontal and vertical components of the gradient of I (k) respectively. If the optical flow is valid, the motion vector field (v x , v y ) is given by the following equation:

Equation

[0060] By combining this optical flow equation with Hermite interpolation for the motion trajectories of each sample result, the function value I (k) and the derivative, ∂I (k) / ∂x, ∂I (k) / ∂y, a unique cubic polynomial that finally matches both is obtained. The value of this polynomial at t = 0 is the BIO prediction:

Number

[0061] Here, τ0 and τ1 represent the distances to the reference frame, as shown in Figure 4 which shows an example of the optical flow trajectory. The distances τ0 and τ1 are calculated based on the POC for Ref0 and Ref1: τ0 = POC(current) - POC(Ref0) τ1 = POC(Ref1) - POC(current) If both predictions come from the same time direction (either both come from the past or both come from the future), the signs are different (i.e., τ0·τ1 < 0). In this case, BIO is applicable only when the predictions are not from the same time point (i.e., τ0 ≠ τ1), both of the referenced regions have non-zero motion (MV x0 , MV y0 , MV x1 , Mv y1 ≠0), and the block motion vectors are proportional to the time distance (MV x0 / MV x1 = MV y0 / Mv y1 = -τ0 / τ1).

[0062] The motion vector field (v x , v y ) is determined by minimizing the difference Δ between the values at points A and B (the intersections of the motion trajectory in Figure 4 and the reference frame plane). The model uses only the first linear term of the local Taylor expansion for Δ:

Number

[0063] All values of Equation (5) depend on the sample positions (i’, j’), which have been omitted from the notation so far. If the motion is consistent in the local neighborhood area, assuming M is equal to 2, Δ is minimized within the (2M + 1) × (2M + 1) square window Ω centered on the currently predicted point (i, j):

Number

[0064] For this optimization problem, JEM uses a simplified approach that first minimizes in the vertical direction and then in the horizontal direction. As a result, the following is obtained:

Number

Number

Number

[0065] To avoid division by zero or extremely small values, the normalization parameters r and m are introduced in Equations (7) and (8): r = 500·4 d-8 (10) m = 700·4 d-8 (11) Here, d is the bit depth of the video sample.

[0066] To keep the memory access for BIO the same as in the case of normal dual-prediction motion compensation, all prediction and gradient values I (k) , ∂I (k) / ∂x, ∂I (k)∂ / ∂y is calculated only for the positions within the current block. In Equation (9), the (2M + 1)×(2M + 1) square window Ω centered at the currently predicted point on the boundary of the predicted block needs to access points outside the block as shown in FIG. 5A. In JEM, for I (k) , ∂I (k) / ∂x, ∂I (k) / ∂y values are set to be equal to the nearest available value within that block. For example, this can be implemented as padding as shown in FIG. 5B. FIGS. 5A and 5B show an example of BIO without block extension. FIG. 5A shows an example of an access position outside the block, and FIG. 5B shows an example of the padding used to avoid extra memory access and calculation.

[0067] According to BIO, it is possible for the motion field to be refined for each sample. To reduce the computational complexity, a block-based design of BIO is used in JEM. Motion refinement is calculated based on 4×4 blocks. In block-based BIO, the value of s n in Equation (9) for all samples within a 4×4 block is aggregated, and the aggregated value of s n is used to derive the BIO motion vector offset for the 4×4 block. More specifically, the following equation is used for block-based BIO derivation.

Equation

[0068] In some cases, the MV regime of BIO may be unreliable due to noise or irregular motion. Therefore, in BIO, the magnitude of the MV regime is clipped to a threshold thBIO. The threshold is determined based on whether all the reference pictures of the current picture are from one direction. When all the reference pictures of the current picture are from one direction, the threshold is set to 12×2 14-d and otherwise, it is set to 12×2 13-d .

[0069] The gradient of BIO is calculated simultaneously with motion compensation interpolation using an operation that matches the HEVC motion compensation process (2D separable FIR). The input for this 2D separable FIR is the same reference frame samples as in the motion compensation process and the fractional positions (fracX, fracY) according to the fractional part of the block motion vector. In the case of the horizontal gradient ∂I / ∂x, the signal is first interpolated vertically using BIOfilterS corresponding to the fractional position fracY with a de-scaling shift d-8, and then the gradient filter BIOfilterG is applied horizontally corresponding to the fractional position fracX with a de-scaling shift by 18-d. In the case of the horizontal gradient ∂I / ∂Y, first the gradient filter is applied vertically using BIOfilterG corresponding to the fractional position fracY with a de-scaling shift d-8, and then the signal displacement is performed horizontally using BIOfilterS corresponding to the fractional position fracX with a de-scaling shift by 18-d. The length of the interpolation filters for the gradient calculation BIOfilterG and the signal displacement BIOfilterF is shorter (6 taps) to maintain reasonable complexity. Table 1 shows the filters used for gradient calculation for various fractional positions of block motion in BIO. Table 2 shows the interpolation filters used for prediction signal generation in BIO.

Table 1

Table 2

[0070] In JEM, BIO is applied to all doubly predicted blocks when the two predictions are from different reference pictures. BIO is disabled when LIC is available for the CU.

[0071] In JEM, OBMC is applied to the block after the normal MC process. To reduce computational complexity, BIO is not applied during the OBMC process. That is, BIO is used in the MC process only when the block uses its own MV, and is not applied in the MC process when the MV of adjacent blocks is used during the OBMC process.

[0072] A two-stage early termination method is used to conditionally disable the BIO operation according to the similarity between two prediction signals. Early termination is applied first at the CU level and then at the sub-CU level. Specifically, the proposed method first calculates the SAS between the L0 and L1 prediction signals at the CU level. If BIO is applied only to luma, only luma samples need to be considered for SAD calculation. If the SAD at the CU level is no longer greater than a predefined threshold, the BIO process is completely disabled for the entire CU. The CU-level threshold is set to 2 (BDepth-9) per sample. If the BIO process is not disabled at the CU level and the current CU contains multiple sub-CUs, the SAD of each sub-CU within the CU will be calculated. Then, a decision on whether to enable or disable the BIO process is made at the sub-CU level based on a predefined sub-CU level SAD threshold. The sub-CU level SAD threshold is set to 3×2 (BDepth-10) per sample.

[0073] 1.4. Specification of BDOF in VVC (In JVET-N1001-v2) The specification of BDOF is as follows: 8.5.7.4 Bidirectional Optical Flow Prediction Process The inputs to this process are: · Two variables nCbW and nCbH that specify the width and height of the current coding block, · Two (nCbW + 2) × (nCbH + 2) arrays of luma prediction samples predSamplesL0 and predSamplesL1, · Prediction list utilization flags predFlagL0 and predFlagL1, · Reference indices refIdxL0 and refIdxL1, · Bidirectional optical flow utilization flag bdofUtilizationFlag[xIdx][yIdx] with xIdx = 0..(nCbW >> 2) - 1, yIdx = 0..(nCbH >> 2) - 1 That is. The output of this process is an (nCbW) × (nCbH) array pbSamples of luma prediction samples. The variables bitDepth, shift1, shift2, shift3, shift4, offset4, and mvRefineThres are derived as follows: · The variable bitDepth is set equal to BitDepth Y is set equal to. · The variable shift1 is set equal to Max(2, 14 - bitDepth). · The variable shift2 is set equal to Max(8, bitDepth - 4). · The variable shift3 is set equal to Max(5, bitDepth - 7). · The variable shift4 is set equal to Max(3, 15 - bitDepth), and the variable offset4 is set equal to 1 << (shift4 - 1). · The variable mvRefineThres is set equal to Max(2, 1 << (13 - bitDepth)). For xIdx = 0..(nCbW >> 2) - 1, yIdx = 0..(nCbH >> 2) - 1, the following holds: · The variable xSb is set equal to (xIdx << 2) + 1, and ySb is set equal to (yIdx << 2) + 1. · If bdofUtilizationFlag[xSbIdx][yIdx] is equal to false, for x = xSb - 1..xSb + 2, y = ySb - 1..ySb + 2, the predicted sample values of the current block are derived as follows: [Number] · Otherwise (bdofUtilizationFlag[xSbIdx][yIdx] is equal to true), the predicted sample values of the current block are derived as follows: · For x = xSb - 1..xSb + 4, y = ySb - 1..ySb + 4, the following ordered steps are applied: 1. For each corresponding sample position (x, y) in the predicted sample array, the positions (h x , v y ) are derived as follows: h x = Clip3(1, nCbW, x) (8 - 853) v y = Clip3(1, nCbH, y) (8 - 854) 2. The variables gradientHL0[x][y], gradientVL0[x][y], gradientHL1[x][y] and gradientVL1[x][y] are derived as follows: [Number] 3. The variables temp[x][y], tempH[x][y] and temp[x][y] are derived as follows: [Number] · The variables sGx2, sGy2, SGxGy, sGxdI and sGydI are derived as follows: [Number] · The horizontal and vertical motion offsets of the current sub-block are derived as follows:

Equation

Equation

Equation

Equation

Equation

[0074] 1.5. Decoder-side Motion Vector Refinement In the dual-prediction operation, for the prediction of the area of one block, two prediction blocks respectively formed using the motion vector (MV) of list 0 and the MV of list 1 are combined to form a single prediction signal. In JVET-K0217, in the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the dual prediction are further refined by the bilateral matching process.

[0075] In the proposed method, DMVR is applied only in the merge mode and the skip mode when the following conditions are met: (POC - POC0)×(POC - POC1) < 0 Here, POC is the picture order count of the current picture to be encoded, and POC0 and POC1 are the reference picture order counts for the current picture.

[0076] The signaled merge candidate pairs are used as inputs to the DMVR process and are represented by the initial motion vectors (MV0, MV1). The search points explored by DMVR follow the motion vector difference mirroring condition. That is, the candidate vector pairs (MV0’, MV1’) represented and checked by DMVR Any points follow the following two equations: MV0’ = MV0 + MV diff MV1’ = MV1 - MV diff Here, MV diff represents a point within the search space in one of the reference pictures.

[0077] After constructing the search space, the unilateral prediction is configured using a regular 8 - tap DCTIF interpolation filter. The bilateral matching cost function is calculated by the MRSAD (Mean Removed Sum of Absolute Differences) between the two predictions (see Figure 6 showing an example of bilateral matching by 6 - point search), and the search point that results in the minimum cost is selected as the refined MV pair. For MRSAD calculation, 16 - bit precision of the samples is used (which is the output of the interpolation filtering), and no clipping operation and rounding operation are applied before the MRSAD calculation. The reason for not applying rounding and clipping is to reduce the internal buffer requirements.

[0078] In the proposed method, the integer precision search points are selected by the adaptive pattern method. The cost corresponding to the central point (indicated by the initial motion vector) is calculated first. The other four costs (sign shape) are calculated by two predictions located on opposite sides of each other at the central point. The last sixth point in terms of angle is selected by the gradient of the costs calculated previously, as shown in FIGS. 7a and 7b. FIG. 7a shows an example of an adaptive integer search pattern, and FIG. 7b shows an example of a half-sample search pattern.

[0079] The output of the DMVR process is the refined motion vector pair corresponding to the minimum cost.

[0080] After one iteration, if the minimum cost is achieved at the central point of the search space, i.e., if the motion vector does not change, the refinement process ends. Otherwise, further, the best cost is considered as the center and the process continues, while the minimum cost does not correspond to the central point and the search range cannot be exceeded.

[0081] Half-sample precision search is applied only when the application of half-pel search does not exceed the search range. In this case, only 4 MRSAD calculations are performed. This corresponds to the plus-shaped points around the central point and is selected as the optimal one during integer precision search. Finally, the refined motion vector pair is output, which corresponds to the minimum cost point.

[0082] Some simplifications and improvements are further proposed in JVET-L0163.

[0083] Reference Sample Padding The reference sample padding is applied to expand the reference sample block indicated by the initial motion vector. It is assumed that when the size of the coding block is given by "w" and "h", a block of size w + 7 and h + 7 is read from the reference picture buffer. The read buffer is then expanded by two samples in each direction by iterative sample padding using the nearest sample. Thereafter, the expanded reference sample block is used to generate the final prediction when a refined motion vector is obtained (which can be displaced by two samples in either direction from the initial motion vector).

[0084] It can be seen that this change completely removes the extra memory access requirements of DMVR without any coding loss.

[0085] Bilinear Interpolation Instead of 8 - Tap DCTIF According to the proposal, bilinear interpolation is applied during the DMVR search process. That is, the prediction used in the MRSAD calculation is generated using bilinear interpolation. When the final refined motion vector is obtained, a regular 8-tap DCTIF interpolation filter is applied to generate the final prediction.

[0086] Disable DMVR for Small Blocks DMVR is disabled for blocks 4×4, 4×8, and 8×4.

[0087] Early Termination Based on MV Difference between Merge Candidates Additional conditions are imposed on DMVR to limit the MV refinement process. Thereby, DMVR is conditionally disabled when the following conditions are met.

[0088] The MV difference between any of the previous merge candidates in the same merge list as the selected merge candidate is smaller than a predefined threshold (i.e., intervals of 1 / 4, 1 / 2, and 1 pixel wide for CUs having less than 64 pixels, less than 256 pixels, and at least 256 pixels respectively).

[0089] Early Termination Based on SAD Cost at Central Search Coordinates Currently, the sum of absolute differences (SAD) between two prediction signals (L0 and L1 predictions) using the initial motion vector of the CU is calculated. If the SAD is no longer greater than a predefined threshold, i.e., 2 per sample (BDepth-9) , then DMVR is skipped; otherwise, DMVR is still applied to refine the two motion vectors of the current block.

[0090] DMVR Application Conditions The DMVR application condition that it is implemented in BMS2.1, i.e., (POC - POC1)×(POC - POC2)<0, is replaced by the new condition (POC - POC1)==(POC2 - POC). This means that DMVR is applied only when the reference picture is in the opposite temporal direction and equidistant from the current picture.

[0091] MRSAD Calculation Used Every Other Row The MRSAD cost is calculated only for the odd - numbered rows of the block, and the even - numbered sample rows are not considered. Therefore, the number of operations for MRSAD calculation is halved.

[0092] Sub - Pixel Offset Estimation Based on Parametric Error Surface In JVET - K0041, a parametric error surface adapted using integer distance position evaluated costs was proposed to determine sub - pixel offsets with 1 / 16 - pel accuracy with very low computational complexity.

[0093] This method has been introduced into VVC and is summarized as follows: 1. The parametric error surface fit is calculated only if the minimum matching cost of the integer MVD is not equal to 0 and the matching cost of zero MVD is greater than the threshold. 2. The best integer position is considered the center position, and the cost at the center position and the costs at positions (-1,0), (0,-1), (1,0), and (0,1) (in units of integer pixels) relative to the center position are given by the equation: E(x,y)=A(x - x0) 2 +B(y - y0) 2 +C which is used to fit a two - dimensional parametric error surface equation. Here, (x0,y0) corresponds to the position with the minimum cost, and C corresponds to the minimum cost value. By solving five equations for five unknowns, (x0,y0) is: x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y0=(E(0,-1)-E(0,1)) / (2(E(0,-1)+E(0,1)-2E(0,0))) and is calculated as such. (x0,y0) can be calculated to any required sub - pixel accuracy by adjusting the precision at which the division is performed (i.e., how many bits of quotient are calculated). For 1 / 16 - pel accuracy, only 4 bits of the absolute value of the quotient need to be calculated, which is useful for a fast shift - subtract - based implementation of the two divisions required per CU. 3. The calculated (x0,y0) is added to the integer - distance refined MV to obtain a sub - pixel accuracy refined delta MV.

[0094] On the other hand, for a 5×5 search space, the parametric error surface fit is only performed if one of the nine central positions is the best integer position, as shown in Figure 8.

[0095] JVET - N0236: Prediction Refinement by Optical Flow This contribution proposes a method for refining the sub-block based affine motion compensated prediction by optical flow. After the sub-block based affine motion compensation is performed, the prediction sample is refined by adding the difference derived by the optical flow equation. This is called Prediction Refinement with Optical Flow (PROF). The proposed method can achieve inter prediction at pixel level granularity without increasing the memory access bandwidth.

[0096] To achieve finer granularity motion compensation, this contribution proposes a method for refining the sub-block based affine motion compensated prediction by optical flow. After the sub-block based affine motion compensation is performed, the luma prediction sample is refined by adding the difference derived by the optical flow equation. The proposed PROF (Prediction Refinement with Optical Flow) is described as the following four steps.

[0097] Step 1) Sub-block based affine motion compensation is performed to generate the sub-block prediction I(i,j).

[0098] Step 2) The spatial gradients g x (i,j) and g y (i,j) of the sub-block prediction are calculated at each sample position using a 3-tap filter [-1,0,1]. g x (i,j)=I(i + 1,j)-I(i - 1,j) g y (i,j)=I(i,j + 1)-I(i,j - 1)

[0099] The sub-block accuracy is extended by one pixel on each side for gradient calculation. To reduce the memory bandwidth and complexity, the pixels on the extended boundary are copied from the nearest integer pixel position in the reference picture. Thus, additional interpolation for area padding is avoided.

[0100] Step 3) The refinement of the luma prediction (represented by ΔI) is calculated by the optical flow formula. ΔI(i,j)=g x (i,j)×Δv x (i,j)+g y (i,j)×Δv y (i,j) Here, the delta MV (represented by Δv(i,j)) is the difference between the pixel MV calculated for the sample position (i,j), represented by v(i,j), and the sub-block MV of the sub-block to which the pixel (i,j) belongs, as shown in FIG. 10.

[0101] Since the affine model parameters and the pixel positions relative to the sub-block center do not change for each sub-block, Δv(i,j) is calculated for the first sub-block and can be reused for other sub-blocks within the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the center of the sub-block, then Δv(x,y) can be derived by the following formula.

Equation

[0102] For the 4-parameter affine model,

Equation

Equation

[0103] Step 4) Finally, the luma prediction refinement is added to the sub-block prediction I(i,j). The final prediction I’ is generated as the following equation. I’(i,j)=I(i,j)+ΔI(i,j)

[0104] Some Details in JVET - N0236 a) Method for deriving the gradient of PROF In JVET-N0263, the gradient is calculated for each sub-block of each reference list (4×4 sub-blocks of VTM-4.0). For each sub-block, the closest integer samples of the reference block are fetched to pad the lines outside the four sides of the samples (see the previous figure). Assume that the MV of the current sub-block is (MV x ,MV y ). Then, the fractional part is calculated as (FracX,FracY)=(MVx&15,MVy&15). The integer part is calculated as (IntX,IntY)=(MVx>>4,MVy>>4). The offset (OffsetX,OffsetY) is: OffsetX=FracX>7?1:0; OffsetY=FracY>7?1:0; derived as such. Assume that the upper-left coordinate of the current sub-block is (xCur,yCur), and the size of the current sub-block is W×H. In that case, (xXor0,yCor0), (xCor1,yCor1), (xCor2,yCor2) and (xCor3,yCor3) are: (xXor0,yCor0)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY-1); (xCor1,yCor1)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY+H-1); (xCor2,yCor2)=(xCur+IntX+OffsetX-1,yCur+IntY+OffsetY); (xCor3, yCor3) = (xCur + IntX + OffsetX + W, yCur + IntY + OffsetY); It is calculated as follows. PredSample[x][y] (where x = 0..W - 1 and y = 0..H - 1) holds the predicted samples of the sub - block. Then, the padding samples are: PredSample[x][-1] = (Ref(xCor0 + x, yCor0) << Shift0) - Rounding (when x = -1..W); PredSample[x][H] = (Ref(xCor1 + x, yCor1) << Shift0) - Rounding (when x = -1..W); PredSample[-1][y] = (Ref(xCor2, yCor2 + y) << Shift0) - Rounding (when y = 0..H - 1); PredSample[W][y] = (Ref(xCor3, yCor3 + y) << Shift0) - Rounding (when y = 0..H - 1); are derived as follows. Here, Rec represents the reference picture. Rounding is an integer equal to 2 in an exemplary PROF implementation. Shift0 = Max(2, (14 - BitDepth)). 13 is an integer equal to 2 in an exemplary PROF implementation. Shift0 = Max(2, (14 - BitDepth)). Unlike the BIO of VTM - 4.0 where the gradient is output with the same regime as the input luma samples, PROF attempts to increase the accuracy of the gradient. The gradient of PROF is as follows: Shift1 = Shfit0 - 4 gradientH[x][y] = (predSample[x + 1][y] - predSample[x - 1][y]) >> Shift1 gradientV[x][y] = (predSample[x][y + 1] - predSample[x][y - 1]) >> Shift1 is calculated as such. It should be noted that predSample[x][y] maintains its accuracy after interpolation.

[0105] b) Method for deriving Δv of PROF The derivation of Δv (represented by dMvH[posX][posY] and dMvV[posX][posY]) (where posX = 0..W - 1 and posY = 0..H - 1) can be described as follows. Assume that the dimensions of the current block are cbWidth × cbHeight, the number of control point motion vectors is numCpMv, and the control point motion vectors are cpMvLX[cpIdx]. Here, cpIdx = 0..numCpMv - 1, and X is 0 or 1 representing two reference lists. The variables log2CbW and log2CbH are derived as follows: log2CbW = Log2(cbWidth) log2CbH = Log2(cbHeight) The variables mvScaleHor, mvScaleVer, dHorX, and dVerX are derived as follows: mvScaleHor = cpMvLX[0][0] << 7 mvScaleVer = cpMvLX[0][1] << 7 dHorX = (cpMvLX[1][0] - cpMvLX[0][0]) << (7 - log2CbW) dVerX = (cpMvLX[1][1] - cpMvLX[0][1]) << (7 - log2CbW) The variables dHorY and dVerY are derived as follows: · When numCpMv is equal to 3, the following applies: dHorY = (cpMvLX[2][0] - cpMvLX[0][0]) << (7 - log2CbH) dVerY = (cpMvLX[2][1] - cpMvLX[0][1]) << (7 - log2CbH) · Otherwise (when numCpMv is equal to 2), the following applies: dHorY = -dVerX dVerY = dHorX The variables qHorX, qVerX, qHorY, and qVerY are: qHorX = dHorX << 2; qVerX = dVerX << 2; qHorY = dHorY << 2; qVerY = dVerY << 2; It is derived as follows. dMvH[0][0] and dMvV[0][0] are: dMvH[0][0] = ((dHorX + dHorY) << 1) - ((qHorX + qHorY) << 1); dMvV[0][0] = ((dVerX + dVerY) << 1) - ((qVerX + qVerY) << 1); It is calculated as follows. For dMvH[xPos][0] and dMvV[xPos][0] where xPos ranges from 1 to W - 1: dMvH[xPos][0] = dMvH[xPos - 1][0] + qHorX; dMvV[xPos][0] = dMvV[xPos - 1][0] + qVerX; It is derived as follows. For yPos ranging from 1 to H - 1, the following applies: dMvH[xPos][yPos] = dMvH[xPos][yPos - 1] + qHorY (where xPos = 0..W - 1) dMvV[xPos][yPos] = dMvV[xPos][yPos - 1] + qVerY (where xPos = 0..W - 1) Finally, for dMvH[xPos][yPos] and dMvV[xPos][yPos] (where posX = 0..W - 1, posY = 0..H - 1): dMvH[xPos][yPos] = SatShift(dMvH[xPos][yPos], 7 + 2 - 1); dMvV[xPos][yPos] = SatShift(dMvV[xPos][yPos], 7 + 2 - 1); They are right-shifted as follows. Here, SatShift(x, n) and Shift(x, n) are:

Number

[0106] c) Method for deriving ΔI of PROF For the position (posX, posY) within the sub - block, the corresponding Δv(i, j) is represented as (dMvH[posX][posY], dMvV[posX][posY]). The corresponding gradient is represented as (gradientH[posX][posY], gradientV[posX][posY]). Next, ΔI(posX, posY) is derived as follows. (dMvH[posX][posY], dMvV[posX][posY]) is:[[]] dMvH[posX][posY]=Clip3(-32768, 32767, dMvH[posX][posY]); dMvV[posX][posY]=Clip3(-32768, 32767, dMvV[posX][posY]); Clipped as such. ΔI(posX, posY)=dMvH[posX][posY]×gradientH[posX][posY]+dMvV[posX][posY]×gradientV[posX][posY]; ΔI(posX, posY)=Shift(ΔI(posX, posY), 1 + 1+4); ΔI(posX, posY)=Clip3(-(2 13 -1), 2 13 -1, ΔI(posX, posY)); (PROF - eq2)

[0107] d) Method for deriving I’ of PROF When the current block is not coded as double - prediction or weighted - prediction: I’(posX, posY)=Shift((I(posX, posY)+ΔI(posX, posY)), Shift0), I’(posX, posY)=ClipSample(I’(posX, posY)) Here, ClipSample clips the sample value to a valid output sample value. Next, I’(posX, posY) is output as an inter prediction value. Otherwise (when the current block is coded as dual prediction or weighted prediction), I’(posX, posY) is stored and used to generate an inter prediction value according to other prediction values and / or weight values.

[0108] Related method In the prior PCT applications PCT / CN2018 / 096384, PCT / CN2018 / 098691, PCT / CN2018 / 104301, PCT / CN2018 / 106920, PCT / CN2018 / 109250, and PCT / CN2018 / 109425, an MV update method and a two-step inter prediction method are proposed. The derived MV between reference block 0 and reference block 1 in BIO is scaled and added to the original motion vectors of list 0 and list 1. On the other hand, the updated MV is used for motion compensation and a second inter prediction is generated as the final prediction.

[0109] On the other hand, in those prior PCT applications, the temporal gradient is changed by removing the average difference between reference block 0 and reference block 1.

[0110] In another prior PCT application PCT / CN2018 / 092118, for several different sub-blocks, only one set of MVs is generated for the chroma component.

[0111] In JVET-N0145, DMVR is coordinated with MMVD (Merge with MVD). DMVR is applied to the motion vector derived by MMVD when the distance of MMVD is greater than 2 pels.

[0112] 1.7. DMVR of VVC draft 4 The use of DMVR in JVET-M1001_v7 (VVC working draft 4, version 7) is defined as follows: ·dmvrFlag is set to 1 when all of the following conditions are met: ·sps_dmvr_enabled_flag is equal to 1 ·The current block is not coded in triangular prediction mode, AMVR affine mode, sub - block modes (including merge affine mode and ATMVP mode). ·merge_flag[xCb][yCb] is equal to 1 ·Both predFlagL0[0][0] and predFlagL1[0][0] are equal to 1 ·mmvd_flag[xCb][yCb] is equal to 0 ·DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) is equal to DiffPicOrderCnt(RefPicList[1][refIdxL1],currPic) ·cbHeight is 8 or more ·cbHeight × cbWidth is 64 or more

[0113] Disadvantages of existing implementations The current design of DMVR / BIO has the following problems: 1) When DMVR is applied to one block or one unit (e.g., 16×16) for processing DMVR, the whole unit / block shares the same refined motion information. However, the refined motion information may not be beneficial for some of the samples within the whole unit / block.

[0114] 2) In JVET-N0145, DMVR may also be applied to the MMVD coding block. However, DMVR is restricted to be applied only when the distance of MMVD is greater than 2 pels. The selected MMVD distance is first used to derive the first refined motion vector of the current block. When DMVR is further applied, the second refined MMVD distance by DMVR may be included in the MMVD distance table. This is inappropriate.

[0115] 3) DMVR / BIO is always applied to blocks that satisfy the conditions of block size and POC distance. However, it does not fully consider the decoded information of one block that may result in the worst performance after refinement.

[0116] 4) A 5×5 square MVD space is searched in DMVR. This is computationally complex. The horizontal MVD can take values in the range from -2 to 2, and the vertical MVD can take values in the range from -2 to 2.

[0117] 5) The parametric error surface fit is only performed for the central nine integer positions. This may be inefficient.

[0118] 6) The MVD values in BDOF / PROF (i.e., the difference between the decoded MV and the refined MV), that is, v derived from Equations 8-867 and 8-868 in 8.5.7.4 of BDOF x and v y , or Δv y (x,y) in PROF, may take large values, so it is impossible to replace multiplication by shifting in the BDOF / PROF process.

[0119] 7) The derived sample prediction refinement of BDOF (e.g., the offset to be added to the predicted sample) may not be suitable for storage as it is not clipped. On the other hand, although PROF uses the same OF flow as BDOF, PROF calls a clipping operation for the derived offset.

[0120] Problems of Residual Coding 1) The coding group (CG) sizes used in the transform skip mode and the regular residual coding mode (e.g., the transform mode) are different. In regular residual coding, a 2×2 CG size is used for 2×2, 2×4, and 4×2 residual blocks, and 2×8 and 8×2 CG sizes are used for 2×N and N×2 (N >= 8) residual blocks, respectively. On the other hand, in the transform skip mode, a 2×2 CG size is always used for 2×N and N×2 (N >= 2) residual blocks.

[0121] Exemplary Method for Coding Tools by Adaptive Resolution Transformation Embodiments of the disclosed technology address the drawbacks of existing implementations. The examples of the disclosed technology given below are discussed to aid understanding of the disclosed technology and should not be construed as limiting the disclosed technology. The various features described in these examples may be combined if no explicit contrary indication is given.

[0122] Let MV0 and MV1 be represented as the MVs of the blocks in prediction directions 0 and 1 respectively. For MVX (x = 0 or 1), MVX[0] and MVX[1] represent the horizontal and vertical components of MVX respectively. When MV’ is equal to (MV0 + MV1), MV’[0] = MV0[0] + MV1[0] and MV’[1] = MV0[1] + MV1[1]. When MV’ is equal to a × MV0, MV’[0] = a × MV1[0] and MV’[1] = a × MV0[1]. Assuming that the reference pictures in list 0 and list 1 are Ref0 and Ref1 respectively, the POC distance between the current picture and Ref0 is PocDist0 (i.e., the absolute value of the result of subtracting the POC of Ref0 from the POC of the current picture), and the POC distance between Ref1 and the current picture is PocDist1 (i.e., the absolute value of the result of subtracting the POC of the current picture from the POC of Ref1). Let the width and height of the block be represented as W and H respectively. Assume that the function abs(x) returns the absolute value of x.

[0123] The horizontal MVD of DMVR can take values in the range from -MVD_Hor_TH1 to MVD_Hor_TH2, and the vertical MVD can take values in the range from -MVD_ver_TH1 to MVD_Ver_TH2. Let the width and height of the block be represented as W and H respectively. Assume that abs(X)) returns the absolute value of X, and Max(X, Y) returns the larger of X and Y. Assume that sumAbsHorMv = abs(MV0[0]) + abs(MV1[0]), sumAbsVerMv = abs(MV0[1]) + abs(MV1[1]), maxHorVerMv = Max(abs(MV[0]), abs(MV1[0])), and maxAbsVerMv = Max(abs(MV[1]), abs(MV1[1])). Let the clipped vx and vy be represented as clipVx and clipVy respectively.

[0124] In the following disclosure, the term "absolute MV" of MV = (MVx, MVy) can refer to abs(MVx), or abs(MVy), or abs(MVx) + abs(MVy), or Max(abs(MVx), abs(MVy)). The term "absolute horizontal MV" refers to abs(MVx), and "absolute vertical MV" refers to abs(MVy).

[0125] The function logK(x) returns the logarithm of x to the base K, the function ceil(y) returns the smallest integer greater than or equal to y (i.e., the integer value inty when y <= inty < y + 1), and floor(y) returns the largest integer less than or equal to y (i.e., the integer value inty when y - 1 < inty <= y). The function sign(x) returns the sign of x, for example, 0 when x >= 0 and 1 when x < 0. Let x^y be equal to x y to. The function

Number

[0126] Note that the proposed method may also be applicable to other types of decoder-side motion vector derivation / refinement and prediction / reconstruction sample refinement methods.

[0127] DMVR is applied to one block or one unit (e.g., 16×16) for processing DMVR. In the following description, the sub-region may be a part smaller than the processing unit.

[0128] BIO is also known as BDOF (Bi-Directional Optical Flow).

[0129] In the following discussion, the MVDs used in the decoder-side motion derivation process (e.g., BDOF, PROF) can be represented by v x , v y , Δv x (x, y) and Δv y (x, y). In one example, v x and vy In BDOF, "v x and v y " derived from Equations 8-867 and 8-868 in 8.5.7.4, or in PROF, "Δv x (x,y) and Δv y (x,y)" may be referred to. In one example, Δv x (x,y) and Δv y (x,y) in BDOF, "v x and v y " derived from Equations 8-867 and 8-868 in 8.5.7.4, or in PROF, "Δv x (x,y) and Δv y (x,y)" may be referred to.

[0130] The proposed method for PROF / BDOF may be applicable to other types of coding methods that use optical flow.

[0131] 1. The results of MRSAD or other rules (e.g., SAD, SATD) using a given MVD (or a given pair of MVDs in two lists) may be further modified before being used to select refined motion vectors in the decoder-side motion derivation process. a. In one example, the results may be modified by multiplying by a scaling factor. i. Alternatively, the results may be modified by dividing by a scaling factor. ii. Alternatively, the results may be modified by adding / subtracting a scaling factor. iii. In one example, the scaling factor may depend on the value of the MVD. 1. In one example, the scaling factor may be increased for larger MVDs (e.g., the larger absolute value of the horizontal and / or vertical components of the MVD). 2. In one example, the scaling factor may depend on the sub-pel position of the candidate MV having the MVD tested on the initialized MV plus. iv. In one example, the scaling factor may depend on the permitted MVD set (e.g., the horizontal or vertical components of the candidates within the permitted MVD set are in the range [-M, N] assuming M and N are non-negative values).

[0132] 2. It is proposed that whether and / or how to apply DMVR and / or BIO may depend on the coding information of the block. a. In one example, the coding information may include the motion vectors and POC values of the reference pictures. i. In one example, the coding information may include whether the block is coded in the AMVP mode or the merge mode. 1. In one example, for a block coded in the AMVP, BIO may be disabled. ii. In one example, whether to enable DMVR / BIO may depend on the sum of two motion vectors (represented by MV0 and MV1 and represented by MV0' and MV1') or the absolute value of the MV components. 1. In one example, when abs(MV0'[0] + MV1'[0]) > T1 or abs(MV0'[1] + MV1'[1]) > T1, DMVR and / or BIO may be disabled. For example, T1 = 10 integer pixels. 2. In one example, when abs(MV0'[0] + MV1'[0]) > T1 and abs(MV0'[1] + MV1'[1]) > T1, DMVR and / or BIO may be disabled. 3. In one example, when abs(MV0'[0] + MV1'[0]) + abs(MV0'[1] + MV1'[1]) > T2, DMVR and / or BIO may be disabled. For example, T2 = 15 integer pixels. 4. In one example, when abs(MV0'[0] + MV1'[0]) < T1 or abs(MV0'[1] + MV1'[1]) < T1, DMVR and / or BIO may be disabled. For example, T1 = 10 integer pixels. 5. In one example, when abs(MV0’[0] + MV1’[0]) > T1 and abs(MV0’[1] + MV1’[1]) > T1, DMVR and / or BIO can be disabled. 6. In one example, when abs(MV0’[0] + MV1’[0]) + abs(MV0’[1] + MV1’[1]) < T2, DMVR and / or BIO can be disabled. For example, T1 = 15 integer pixels. 7. In one example, MV0’ is set equal to MV0 and MV1’ is set equal to MV1. 8. In one example, MV0’ is set equal to MV0 and MV1 can be scaled to generate MV1’. a. In one example, MV1’ = MV1 × PocDist0 / PocDist1. b. In one example, MV1’ = MV1. 9. In one example, MV1’ is set equal to MV1 and MV0 can be scaled to generate MV0’. a. In one example, MV0’ = MV0 × PocDist1 / PocDist0. b. In one example, MV0’ = MV0. 10. In one example, MV0’ can be set equal to MV0 × (POCRef1 - POCcur) and MV1’ is set equal to MV1 × (POCcur - POCRef0). POCcur, POCRef0, and POCRef1 can represent the POC values of the current picture, the reference picture of MV0, and the reference picture of MV1, respectively. b. In one example, the coding information may include a merge index and / or an MVP index. i. For example, DMVR and / or BDOF may be permitted only for a specific merge index. 1. For example, DMVR and / or BDOF may be permitted only for even merge indices. 2. For example, DMVR and / or BDOF may be permitted only for odd merge indices. 3. For example, DMVR and / or BDOF may be permitted only for merge indices smaller than a threshold T1. 4. For example, DMVR and / or BDOF may be permitted only for merge indices greater than a threshold value T1. ii. For example, BDOF may be permitted only for a specific MVP index, for example, an even or odd MVP index. c. In one example, the coding information may include whether the SBT mode is being used by a block. i. In one example, DMVR and / or BDOF may not be permitted when the SBT mode is being used. ii. In one example, DMVR and / or BDOF may be permitted only for sub - partitions having non - zero residuals in the SBT mode. iii. In one example, DMVR and / or BDOF may not be permitted only for sub - partitions having non - zero residuals in the SBT mode.

[0133] 3. It is proposed that whether and / or how to apply DMVR and / or BIO may depend on the relationship between the pre - refined MV referring to the first reference list (for example, denoted as MV0 in Section 2.5) and the pre - refined MV referring to the second reference list (for example, denoted as MV1 in Section 2.5). a. In one example, when MV0 and MV1 are symmetric in a block, DMVR and / or BIO may be disabled. i. For example, when the block has symmetric motion vectors (for example, (MV0 + MV1) has only zero components), BIO may be disabled. In one example, the block may be coded in the AMVP mode. Alternatively, the block may be coded in the merge mode. ii. Whether MV0 and MV1 are symmetric may also depend on the POC. 1. For example, MV0 and MV1 are symmetric when MV1×(POCcur - POC0)+MV0×(POC1 - POCcur) is equal to the zero motion vector. POCcur, POCRef0, and POCRef1 may represent the POC values of the current picture, the reference picture of MV0, and the reference picture of MV1, respectively. b. In one example, when MV0 and MV1 are approximately symmetric in a block, DMVR or / and BIO can be disabled. i. For example, MV0 and MV1 are approximately symmetric when abs(MV0 + MV1) < Th1. Th1 is a number such as 1 integer pixel. ii. Whether MV0 and MV1 are symmetric may depend on the POC. 1. For example, MV0 and MV1 are approximately symmetric when abs(MV1 × (POCcur - POC0) + MV0 × (POC1 - POCcur)) < Th2. Th2 is a number that may depend on the POC. c. In one example, when MV0 and MV1 are far from being symmetric in a block, DMVR or / and BIO can be disabled. i. For example, MV0 and MV1 are not symmetric when abs(MV0 + MV1) > Ths1. Th1 is a number such as 16 integer pixels. ii. Whether MV0 and MV1 are symmetric may depend on the POC. 1. For example, MV0 and MV1 are not approximately symmetric when abs(MV1 × (POCcur - POC0) + MV0 × (POC1 - POCcur)) > Th2. Th2 is a number that may depend on the POC.

[0134] 4. Multiple-step refinement may be applied. As the step size increases, the refined region becomes smaller. a. In one example, the k-th step should refine one or more Mk × Nk units, and the j-th step should refine one or more Mj × Nj units. Here, k > j, Mk × Nk is not equal to Mj × Nj, and at least one of the two conditions Mk <= Mj, Nk < -Mj holds.

[0135] 5. The refined motion information within one block may be different for each K × L sub-region. a. In one example, the first K×L sub-region may use the decoded motion information, and the second K×L sub-region may use the refined motion information derived by DMVR applied to the unit covering the second sub-region.

[0136] 6. Whether to apply the sub-region BIO may depend on the output of the DMVR process. a. In one example, for the sub-regions within the block / unit to which DMVR is applied, if the motion information does not change after the DMVR process, BIO can be disabled for that sub-region.

[0137] 7. N (N>1) pairs of MVDs (e.g., (MV diff , -MV diff )) in Section 2.5 may be selected by DMVR to generate the final prediction sample, and it is proposed that one pair includes two motion vector differences including one from list 0 and one from list 1. a. In one example, N = 2, and it includes the selected MVD pair in the DMVR process and the all-zero MVD pair. b. In one example, N = 2. Two pairs of MVDs at the discretion of the DMVR process (e.g., two MVDs with the minimum cost) can be used. c. In one example, each K×L sub-region within the block may determine its own MVD pair. d. In one example, each pair of MVDs can identify two reference blocks. The cost (e.g., SAD / SSE) between each sample or block (e.g., 2×2 block or 4×4 block, etc.) within the two reference blocks can be calculated. Then, the final prediction sample can be generated as a combination of all reference block pairs according to the cost values. Let predKLX (K = 1, 2,..., N) represent the reference sample in list X (X = 0 or 1) identified by the Kth pair of MVDs, and costK represent the cost of the Kth pair of MVDs. Both predKLX and costK can be functions of (x, y) (e.g., the position of the sample). e. In one example, the weighting of each sample may be inversely proportional to the cost. i. In one example,

Number

Number

[0138] 8. When DMVR and MMVD are used together, DMVR may not be permitted to refine the MMVD distance to a specific value. a. In one example, the MMVD distance refined by DMVR should not be included in the MMVD distance table. i. Alternatively, further, if the MMVD distance refined by DMVR is included in the MMVD distance table, DMVR is not permitted and the original MMVD distance (e.g., the signaled MMVD distance) is used. b. In one example, after generating the best integer MVD with DMVR, it is added to the original MMVD distance to generate a coarsely refined MMVD distance. DMVR may not be permitted if the coarsely refined MMVD distance is included in the MMVD distance table. c. In one example, dMVr may be permitted for all MMVD distances with the above constraints. i. For example, DMVR may be permitted for the 2-per-MMVD distance.

[0139] 9. It is proposed that BDOF may not be permitted for samples or sub-blocks depending on the derived MVD, and / or the spatial gradient, and / or the temporal gradient, etc. a. In one example, BDOF may not be permitted for a sample when the derived part widens the distance between the predicted sample in reference list 0 and the predicted sample in reference list 1. i. For example, [Number] When it is (as shown in Equation (5) of Section 2.3), BDOF may not be permitted for the position (x, y). ii. For example, [Number] When it is, BDOF may not be permitted for the position (x, y). iii. For example, [Number] When it is (as shown in Equation (5) of Section 2.3), BDOF may not be permitted for the position (x, y). iv. For example, [Number] When it is, BDOF may not be permitted for the position (x, y). v. Alternatively, the derived offset may be scaled by a coefficient represented as fbdof (fbdof < 1.0) when the derived portion widens the distance between the predicted samples in reference list 0 and the predicted samples in reference list 1. 1. For example, [Number] b. In one example, BDOF may not be permitted for a sample or a K×L (e.g., 2×2, 1×1) sub-block when the derived portion widens the distance between the predicted samples in reference list 0 and the predicted samples in reference list 1. i. For example, BDOF may not be permitted when Δ1 > Δ2. Here, [Number] where Ω represents the samples belonging to the sub-block. ii. In one example, BDOF may not be permitted when Δ1 ≥ Δ2. iii. In one example, the difference can be measured by SSE (Sum of Squared Error), or mean removed SSE, or SATD (Sum of Absolute Transformed Difference). For example,

Number

Number

[0140] 10. When the MVD (for example, v x and v y ) derived by BDOF changes after the clipping operation, the cost, for example,

Number

Number

[0141] 11. The motion vector differences (i.e., MVDs) derived by BDOF and other decoder refinement tools (e.g., PROF) may be further refined based on one or more sets of candidate MVDs that depend on the derived MVDs. And the block can be reconstructed based on the signaled motion vector and the refined MVD. Let the derived MVD be (v x ,v y) represented by, and refined MVD (refinedV x , refinedV y ) is represented by. a. In one example, a two-step refinement may be applied. The first set of candidate MVDs used in the first step is the unchanged v y and the changed v x . Based on this, the second set of candidate MVDs is the unchanged v y and the changed v y . Or vice versa. b. In one example, the set of candidate MVDs is defined to include (v x + offsetX, v y ), where offsetX is within the range [-ThX1, ThX2], and ThX1 and ThX2 are non-negative integer values. In one example, ThX1 = ThX2 = ThX, and ThX is a non-negative integer. In other examples, ThX1 = ThX2 + 1 = ThX, and ThX is a non-negative integer. i. In one example, ThX = 2. In one example, MVDs including (v x - 2, v y ), (v x - 1, v y ), (v x , v y ), (v x + 1, v y ), and (v x + 2, v y ) can be checked after deriving v x . ii. Alternatively, ThX = 1. In one example, MVDs including (v x - 1, v y ), (v x , v y ), and (v x + 1, v y ) can be checked after deriving v x . iii. In one example, only some of the values within the range [-ThX1, ThX2] (e.g., offsetX) may be permitted. iv. Alternatively, v y is set equal to 0. v. Alternatively, vx is set equal to refinedV y . c. In one example, the set of candidate MVDs is defined to include (v x , v y + offsetY), where offsetY is in the range [-ThY1, ThY2], and ThY1 and ThY2 are non-negative integer values. In one example, ThY1 = ThY2 = ThY, where ThY is a non-negative integer. In other examples, ThY1 = ThY2 + 1 = ThY, where ThY is a non-negative integer. i. In one example, ThY = 2. In one example, an MVD including (v x , v y - 2), (v x , v y - 1), (v x , v y ), (v x , v y + 1), and (v x , v y + 2) can be checked after deriving v y . ii. In one example, ThY = 1. In one example, an MVD including (v x , v y - 1), (v x , v y ), and (v x , v y + 1) can be checked after deriving v y . iii. In one example, only some of the values within the range [-ThY1, ThY2] (e.g., offsetY) may be permitted. iv. Alternatively, v x is set equal to refinedV y . v. Alternatively, v x is set equal to 0. d. In one example, v x and v y may be refined in order. i. For example, v x is first refined, for example, according to bullet 10.b. 1. Alternatively, further, v yis assumed to be equal to 0. ii. For example, v y is first refined, for example, according to burette 10.c. 1. Alternatively, further, v x is assumed to be equal to 0. e. In one example, v x and v y may be refined together. For example, a set of candidate MVDs is defined to include (v x + offsetX, v y + offsetY), where offsetX is within the range [-ThX1, ThX2] and offsetY is within the range [-ThY1, ThY2]. i. For example, ThX1 = ThX2 = 2 and ThY1 = ThY2 = 2. ii. For example, ThX1 = ThX2 = 1 and ThY1 = ThY2 = 1. iii. In one example, only some values of (v x + offsetX, v y + offsetY) may be included in the candidate MVD set. f. In one example, the number of candidate MVDs derived from v x and v y may be the same, for example, ThX1 = ThY1 and ThX2 = ThY2. i. Alternatively, v x and v y may derive a different number of candidate MVDs, for example, ThX1!= ThY1 or ThX2!= ThY2. g. Alternatively, further, the MVD that achieves the minimum cost (for example, the cost defined in burette 9) may be selected as the refined MVD.

[0142] 12. It is proposed that the MVDs used in BDOF and other decoder refinement methods (e.g., PROF) may be restricted to be within a given candidate set. a. The maximum and minimum values of the horizontal MVD are represented by MVD maxX and MVD minX respectively, and the maximum and minimum values of the vertical MVD are represented by MVDmaxY and MVD minY is represented by. The number of horizontal MVD candidates allowed within a given candidate set must be less than (1 + MVD maxX - MVD minX ), and / or the number of vertical MVD candidates allowed within a given candidate set must be less than (1 + MVD maxY - MVD minY ). b. In one example, a given candidate set has only candidates where v x and / or v y takes the form of K m , where m is an integer. For example, K = 2. i. Alternatively, further, a given candidate set may also include candidates where v x and / or v y is equal to zero. ii. Alternatively, a given candidate set for Δv x (x, y) or / and Δv y (x, y) has only candidates where Δv x (x, y) or / and Δv y (x, y) is zero or takes the form of K m , where m is an integer. For example, K = 2. c. In one example, a given candidate set has only candidates where v m and / or v x and / or v y or / and Δv x (x, y) or / and Δv y (x, y) has an absolute value of zero or takes the form of K d. In one example, a given candidate set has only candidates where v x and / or v y or / and Δv x (x, y) or / and Δv y (x, y) is zero or takes the form of K m , where m is an integer. For example, K = 2. e. In one example, MVD may first be derived using the prior art (e.g., v derived in equations 8-867 and 8-868 of 8.5.7.4 in BDOF), and then may be modified to be one of the allowed candidates within a given candidate set. x and v y ) and then may be modified to be one of the allowed candidates within a given candidate set. i. In one example, X is set equal to abs(v x ) or abs(v y ), X' is derived from X, and v x or v y is changed to sign(v x ) × X' or sign(v y ) × X'. 1. For example, X' may be set equal to K^ceil(logK(X)). 2. For example, X' may be set equal to K^floor(logK(X)). 3. For example, X' may be set equal to 0 when X is less than a threshold T. 4. For example, X' may be set equal to K^ceil(logK(X)) when ceil(logK(X)) - X <= X - floor(logK(X)). a. Alternatively, further, X' may be set equal to K^floor(logK(X)) otherwise. 5. For example, X' may be set equal to K ceil(logK(X)) - X <= X - K floor(logK(X)) and in that case may be set equal to K ceil(logK(X)) . a. Alternatively, further, X' may be set equal to K ceil(logK(X)) - X > X - K floor(logK(X)) and in that case may be set equal to K floor(logK(X)) . 6. For example, X' may be set equal to K ceil(logK(X)) - X < X - K floor(logK(X)) and in that case may be set equal to K ceil(logK(X)) . a. Alternatively, further, X' may be set equal to K ceil(logK(X)) - X >= X - K floor(logK(X)) and in that case may be set equal to K floor(logK(X)) . 7. In one example, K is set to 2. 8. In one example, X’ may be set equal to K ceil(logK(X+offset)) and offset is an integer. For example, offset may be equal to X - 1 or X >> 1. a. In one example, offset may depend on X. For example, offset may be equal to ceil(logK(X)) K / 2 + P. P is an integer such as 0, 1, -1. 9. In one example, X’ may be set equal to K floor(logK(X+offset)) and offset is an integer. For example, offset may be equal to X - 1 or X >> 1. a. In one example, offset may depend on X. For example, offset may be equal to floor(logK(X)) K / 2 + P. P is an integer such as 0, 1, -1. 10. In one example, the conversion from X to X’ may be implemented by a predefined lookup table. a. In one example, the lookup table may be derived from the above method. b. In one example, the lookup table may not be derived from the above method. c. In one example, when converting from X to X’, X may be used as an index for accessing the lookup table. d. In one example, when converting from X to X’, the first N (N = 3) most significant bits of X may be used as an index for accessing the lookup table. e. In one example, when converting from X to X’, the first N (N = 3) least significant bits of X may be used as an index for accessing the lookup table. f. In one example, when converting from X to X’, a specific N (N = 3) consecutive bits of X may be used as an index for accessing the lookup table. 11. The above method replaces v x with Δv x (x, y) to obtain Δv x(x,y) can also be applicable. 12. The above method is for v y to be replaced by Δv y in (x,y) so that Δv y can also be applicable to (x,y). ii. In one example, a set of candidate values that are powers of K is used to calculate a cost (e.g., the cost defined in bullet 9) for v x or v y and may be selected accordingly. The value that achieves the minimum cost is selected as the final v x or v y Let lgKZ = floor(logJ(abs(vz))), where Z = x or y. 1. For example, the set of values may be sign(vz) × {K^lgKZ, K^(lgKZ + 1)}. 2. For example, the set of values may be sign(vz) × {K^lgKZ, K^(lgKZ + 1), 0}. 3. For example, the set of values may be sign(vz) × {K^lgKZ, 0}. 4. For example, the set of values may be sign(vz) × {K^(lgKZ - 1), K^lgKZ, K^(lgKZ + 1)}. 5. For example, the set of values may be sign(vz) × {K^(lgKZ - 1), K^lgKZ, K^(lgKZ + 1), 0}. 6. For example, the set of values may be sign(vz) × {K^(lgKZ - 2), K^(lgKZ - 1), K^lgKZ, K^(lgKZ + 1), K^(lgKZ + 2), 0}. 7. For example, the set of values may be sign(vz) × {K^(lgKZ - 1), K^lgKZ, K^(lgKZ + 1)}. iii. In one example, v x and v y may be changed in order. 1. In one example, v x may be changed first assuming that v y is equal to zero. a. Alternatively, further, v y is the changed v xIt is derived based on 2. In one example, v y may be initially modified assuming that v x is equal to zero. a. Alternatively, further, v x is derived based on the modified v y . f. Alternatively, only the allowed candidates within a given candidate set may be checked to derive the MVD. i. In one example, in BDOF, instead of explicitly deriving the MVD (e.g., v x and v y ) from equations (7) and (8), as derived in equations 8 - 867 and 8 - 868 of 8.5.7.4, v x and v y may be directly selected from a set of candidate MVD values. ii. In one example, v x and v y may be determined in sequence. 1. In one example, v x may be initially determined assuming that v y is equal to zero. a. Alternatively, further, v y is derived based on the determined v x . 2. In one example, v y may be initially determined assuming that v x is equal to zero. a. Alternatively, further, v x is derived based on the determined v y . iii. In one example, v x and v y are determined together. iv. In one example, the candidate MVD value set may contain only values of the form K^N or -K^N or zero, where N is an integer and K > 0, e.g., K = 2. 1. In one example, the candidate MVD value set may contain values of the form M / (K^N), where M is an integer, N >= 0, K > 0, e.g., k = 2. 2. In one example, the candidate MVD value set may include values in the form of M / (K^N), where M and N are integers, K > 0, for example, K = 2. 3. In one example, the candidate MVD value set may include the value zero or values in the form of (K^M) / (K^N), where M and N are integers, K > 0, for example, K = 2. g. Alternatively, only permitted MVD values need not be derived for v x and v y either. i. In one example, vx is derived as follows. Here, the functions F1() and F2() can be floor() or ceil().

Number

Number

Number

[0143] 13. v in BDOF x or / and v y or / and Δv in PROF x (x, y) or / and Δv y (x, y) is first within a predefined range (v x or / and v y or / and Δv x (x, y) or / and Δv y(x, y) may be clipped to different ranges and then converted to a value in the form of zero or K m or -K m by the proposed method. a. Alternatively, v in BDOF x or / and v y or / and Δv in PROF x (x, y) or / and Δv y (x, y) is first converted to a value in the form of zero or K m or -K m by the proposed method, and then clipped to a predefined range (v x or / and v y or / and Δv x (x, y) or / and Δv y (x, y) may be clipped to different ranges). i. Alternatively, further, the range should be in the form of [-K m1 , K n or [K m2 , K n . b. In one example, the numerator and / or denominator in the above method may be clipped to a predefined range before use.

[0144] 14. The clipping parameter in BDOF (e.g., thBIO in equations (7) and (8)) may be different for different sequences / pictures / slices / tile groups / tiles / CTUs / CUs. a. In one example, different thBIO may be used when clipping horizontal MVD and vertical MVD. b. In one example, thBIO may depend on the dimensions of the picture. For example, a larger thBIO may be used for pictures with larger dimensions. c. In one example, thBIO may depend on the decoded motion information of the block. i. For example, a larger thBIO may be used for a block having a larger absolute MV, such as the "absolute MV" of MV[0] or the "absolute MV" of MV[1] or sumAbsHorMv or sumAbsVerMv or sumAbsHorMv + sumAbsVerMv. ii. For example, a larger thBIO may be used to clip the horizontal MVD for a block having a larger absolute horizontal MV, such as the "absolute horizontal MV" of MV[0] or the "absolute horizontal MV" of MV[1] or sumAbsHorMv or maxHorVerMv. iii. For example, a larger thBIO may be used to clip the horizontal MVD for a block having a larger absolute vertical MV, such as the "absolute vertical MV" of MV[0] or the "absolute vertical MV" of MV[1] or sumAbsVerMv or maxAbsVerMv. d. The clipping parameter may be determined by the encoder and signaled to the decoder in the sequence parameter set (SPS) or / and the video parameter set (VPS) or / and the adaptive parameter set (APS) or / and the picture parameter set (PPS) or / and the slice header or / and the tile group header.

[0145] 15. The clipping parameter in BDOF may depend on whether the coding tool X is enabled. a. In one example, X is DMVR. b. In one example, X is the affine inter mode. c. In one example, a larger clipping parameter, such as the thresholds thBIO in equations (7) and (8), may be used when X is disabled for a sequence / video / picture / slice / tile group / tile / brick / coding unit / CTB / CU / block. i. Alternatively, furthermore, a smaller clipping parameter may be used when X is enabled for a sequence / video / picture / slice / tile group / tile / brick / coding unit / CTB / CU / block.

[0146] 16. The MVDs derived from BDOF and PROF (e.g., for PROF, (dMvH, dMvV) in Section 2.6 or / and for BDOF, (v x , v y )) may be clipped to the same range [-N, M], where N and M are integers. d. In one example, N = M = 31. e. In one example, N = M = 63. f. In one example, N = M = 15. g. In one example, N = M = 7. h. In one example, N = M = 3. i. In one example, N = M = 127. j. In one example, N = M = 255. k. In one example, M = N and 2×M is not equal to the value of 2 K (K is an integer). l. In one example, M is not equal to N and (M + N) is not equal to the value of 2 K (K is an integer). m. Alternatively, the MVDs derived from BDOF and PROF may be clipped to different ranges. i. For example, the MVD derived from BDOF may be clipped to [-31, 31]. ii. For example, the MVD derived from BDOF may be clipped to [-15, 15]. iii. For example, the MVD derived from BDOF may be clipped to [-63, 63]. iv. For example, the MVD derived from BDOF may be clipped to [-127, 127]. v. For example, the MVD derived from PROF may be clipped to [-63, 63]. vi. For example, the MVD derived from PROF may be clipped to [-31, 31]. vii. For example, the MVD derived from PROF may be clipped to [-15, 15]. viii. For example, the MVD derived from PROF may be clipped to [-127, 127].

[0147] 17. The MVD associated with one sample (e.g., (dMvH, dMvV) in Section 2.6) may be used to derive the MVDs of other samples by an optical flow-based method (e.g., PROF). n. In one example, the MVD in PROF may be derived using an affine model only for specific positions (e.g., according to PROF-eq1 in Section 2.6), and such an MVD may be used to derive the MVDs of other positions. Assume that the MVD is derived in PROF with size W×H. o. In one example, the MVD may be derived using an affine model only for the upper W×H / 2 part, and the MVD of the lower W×H / 2 part may be derived from the MVD of the upper W×H / 2 part. The example is shown in FIG. 11. p. In one example, the MVD may be derived using an affine model only for the lower W×H / 2 part, and the MVD of the upper W×H / 2 part may be derived from the MVD of the lower W×H / 2 part. q. In one example, the MVD may be derived using an affine model only for the left W×H / 2 part, and the MVD of the right W×H / 2 part may be derived from the MVD of the left W×H / 2 part. r. In one example, the MVD may be derived using an affine model only for the right W×H / 2 part, and the MVD of the left W×H / 2 part may be derived from the MVD of the right W×H / 2 part. The example is shown in FIG. 11. s. In one example, further, the MVD derived using an affine model may be rounded to a predefined accuracy and / or clipped to a predefined range before being used to derive the MVDs of other positions. i. For example, when the MVD of the upper W×H / 2 part is derived using an affine model, such an MVD may be rounded to a predefined accuracy and / or clipped to a predefined range before being used to derive the MVD of the lower W×H / 2 part. In one example, for x = 0, ..., W - 1 and y = H / 2, ..., H - 1, the MVD at position (x, y) is derived as follows. MVD h and MVD v are the horizontal and vertical MVDs, respectively. i. MVD h (x, y) = -MVD h (W - 1 - x, H - 1 - y) ii. MVD v (x, y) = -MVD v (W - 1 - x, H - 1 - y)

[0148] 18. The derived sample accuracy refinement in BDOF (e.g., the offset to be added to the predicted sample) may be clipped to a predefined range [-M, N] (where M and N are integers). a. In one example, N may be equal to M - 1. i. In one example, M is equal to K 2. For example, K may be equal to 11, 12, 13, 14, or 15. ii. Alternatively, N may be equal to M. b. In one example, the predefined range may depend on the bit depth of the current color component. i. In one example, M is equal to K 2, and K = Max(K1, BitDepth + K2), where BitDepth is the bit depth of the current color component (e.g., luma, Cb or Cr, or R, G, or B). Max(X, Y) returns the larger of X and Y. ii. In one example, M is equal to K 2, and K = Min(K1, BitDepth + K2), where BitDepth is the bit depth of the current color component (e.g., luma, Cb or Cr, or R, G, or B). Min(X, Y) returns the smaller of X and Y. iii. For example, K1 may be equal to 11, 12, 13, 14, or 15. iv. For example, K2 may be equal to -2, -1, 0, 1, 2, or 3. In one example, since the three color components share the same bit depth, the range depends entirely on the internal bit depth or the input bit depth of the samples within the block. c. In one example, the predefined range is the same as the range used in PROF. d. Alternatively, furthermore, the refined samples in BDOF (e.g., samples with an offset added to the samples before refinement) may not need to be clipped further.

[0149] 19. It is proposed to align the coding group (CG) sizes in different residual coding modes (e.g., Transform Skip (TS) mode and regular residual coding mode (non - TS mode)). a. In one example, for a block coded with transform skip (where the transform is bypassed or an identity transform is applied), the CG size may depend on whether the block contains more samples than (2×M) (e.g., M = 8) and / or the residual block size. i. In one example, when the block size is A×N (N>=M), the CG size is set to A×M. ii. In one example, when the block size is N×A (N>=M), the CG size is set to M×A. iii. In one example, when the block size is A×N (N<M), the CG size is set to A×A. iv. In one example, when the block size is N×A (N<M), the CG size is set to A×A. v. In one example, A = 2. b. In one example, for a non - TS coded block, the CG size may depend on whether either the width or the height of the block is equal to K. i. In one example, when either the width or the height is equal to K, the CG size is set to K×K (e.g., K = 2). ii. In one example, when the minimum value of the width and the height is equal to K, the CG size is set to K×K (e.g., K = 2). c. In one example, the 2×2 CG size can be used for 2×N or / and N×2 residual blocks in both the transform skip mode and the regular residual coding. d. In one example, the 2×8 CG size can be used for 2×N (N >= 8) residual blocks in both the transform skip mode and the regular residual coding. i. Alternatively, the 2×4 CG size can be used for 2×N (N >= 8) residual blocks in both the transform skip mode and the regular residual coding. ii. Alternatively, the 2×4 CG size can be used for 2×N (N >= 4) residual blocks in both the transform skip mode and the regular residual coding. e. In one example, the 8×2 CG size can be used for 2×N (N >= 8) residual blocks in both the transform skip mode and the regular residual coding. i. The 4×2 CG size can be used for 2×N (N >= 8) residual blocks in both the transform skip mode and the regular residual coding. ii. The 4×2 CG size can be used for 2×N (N >= 4) residual blocks in both the transform skip mode and the regular residual coding.

[0150] 20. Whether PROF should be enabled / disabled for an affine-coded block may depend on the decoded information associated with the current affine-coded block. a. In one example, PROF may be disabled when one reference picture is associated with a different resolution (width or height) than the current picture. b. In one example, if the current block is in uni-prediction coding, PROF may be disabled when the reference picture is associated with a different resolution (width or height) than the current picture. c. In one example, if the current block is in bi-prediction coding, PROF may be disabled when at least one reference picture is associated with a different resolution (width or height) than the current picture. d. In one example, if the current block is in dual prediction coding, PROF may be disabled if all reference pictures are associated with a different resolution (width or height) than the current picture. e. In one example, if the current block is in dual prediction coding, PROF may be enabled if all reference pictures are associated with the same resolution (width or height) even if they are different from the current picture. f. Whether to enable the above method may depend on the coded information of the current picture, such as the CPMV of the current block and the resolution ratio between the reference picture and the current picture.

[0151] 21. PROF may be disabled when certain conditions are met. The specific conditions are, for example, as follows: a. Generalized bi-prediction (GBi, also known as BCW) is enabled. b. Weighted prediction is enabled. c. An alternative half-pixel interpolation filter is applied.

[0152] 22. The above method applied to DMVR / BIO can also be applicable to other decoder-side motion vector derivation (DMVD) methods, such as prediction refinement based on optical flow for the affine mode.

[0153] 23. Whether to use the above method and / or which method to use may be signaled at the sequence / picture / slice / tile / block / video unit level.

[0154] 24. Non-square MVD regions (e.g., diamond regions, MVD regions based on octagons) may be searched by DMVR and / or other decoder-side motion vector derivation methods. a. For example, the following 5×5 rhombus MVD region may be searched. For example, the MVD values listed below are in integer pixel units.

Number

Number

Number

[0155] 25. The MVD region to be searched by DMVR and / or other decoder-side motion vector derivation methods may depend on the block size and / or block shape. Assume that the size of the current block is W×H. a. In one example, when W >= H + TH (TH >= 0), the horizontal MVD search range (e.g., MVD_Hor_TH1 + MVD_Hor_TH2 + 1) may be greater than or equal to the vertical MVD search range (e.g., MVD_Ver_TH1 + MVD_Ver_TH2 + 1). i. Alternatively, when W + TH <= H, the horizontal MVD search range may be larger than the vertical MVD search range. ii. For example, MVD_Hor_TH1 and MVD_Hor_TH2 are set equal to 2, and MVD_Ver_TH1 and MVD_Ver_TH2 are set equal to 1. iii. For example, the following 5×3 MVD region may be searched. The MVD values are in integer pixel units.

Number

Number

Number

[0156] 26. The MVD space to be searched by DMVR and / or other decoder-side motion vector derivation methods may depend on the motion information of the block before refinement. a. In one example, when sumAbsHorMv >= sumAbsVerMv + TH (TH >= 0), the horizontal MVD search range may be larger than the vertical MVD search range. i. Alternatively, when sumAbsHorMv + TH <= sumAbsVerMv, the horizontal MVD search range may be larger than the vertical MVD search range. ii. For example, MVD_Hor_TH1 and MVD_Hor_TH2 are set equal to 2, and MVD_Ver_TH1 and MVD_Ver_TH2 are set equal to 1. iii. For example, the following 5×3 MVD region may be searched, and the MVD values are in integer pixel units.

Number

Number

Number

[0157] 27. Even if the best integer position is at the boundary of the integer search region, in DMVR, a sub - pel MVD may be derived. a. In one example, when the positions on both the left and right of the best integer position are within the search range, a horizontal sub - pel MVD can be derived, for example, using a parametric error surface fitting method. The example is illustrated in FIG. 9. b. In one example, when the positions above and below the best integer position are both within the search range, a vertical sub - pel MVD can be derived, for example, using a parametric error surface fitting method. The example is illustrated in FIG. 9. c. In one example, the sub - pel MVD may be derived for any position. i. For example, the costs of the best integer position and its N (e.g., N = 4) closest adjacent positions can be used to derive the horizontal and vertical sub - pel MVDs. ii. For example, the costs of the best integer position and its N (e.g., N = 4) closest adjacent positions can be used to derive the horizontal sub - pel MVD. iii. For example, the costs of the best integer position and its N (e.g., N = 4) closest adjacent positions can be used to derive the vertical sub - pel MVD. iv. Parametric error surface fitting may be used to derive the sub - pel MVD. d. Alternatively, further, if the sub - pel MVD is not derived in a specific direction (e.g., horizontal or / and vertical direction), it is set equal to zero.

[0158] 28. Sub - pel MVD derivation may not be permitted in DMVR when BDOF is permitted for a picture / slice / tile group / tile / CTU / CU. a. Alternatively, sub - pel MVD derivation may be permitted in DMVR when BDOF is not permitted for a picture / slice / tile group / tile / CTU / CU.

[0159] 29. Which bullet to apply may depend on coded information such as whether BDOF or PRF is applied to one block.

[0160] Embodiment Deleted text is marked by double brackets (e.g., [[a]] represents the deletion of the character 'a'), and newly added parts are highlighted in bold italics.

[0161] Example An example of not permitting sub-pel MVD in DMVR when BDOF is permitted.

Number

[0162] Example An example of an octagonal search area in DMVR.

Number

[0163] Example Restricting the motion vector refinement in BDOF and PROF to be zero or in the form of K m or -K m An example of such a restriction.

[0164] The proposed changes added to JVET-O0070-CE4.2.1a-WD-r1.docx are highlighted in bold italics, and the deleted parts are marked by double brackets (e.g., [[a]] represents the deletion of the character 'a').

Number

[0165] Example An example of clipping the MVD to [-31, 31] in BDOF and PROF is described.

[0166] The proposed changes added to JVET-O2001-vE.docx are highlighted in bold italic, and the deleted parts are marked by double brackets (e.g., [[a]] represents the deletion of the character 'a').

Number

[0167] Example An example of deriving the MVD of the lower W×H / 2 part from the MVD of the upper W×H / 2 part in PROF.

[0168] The proposed changes added to JVET-O2001-vE.docx are highlighted in bold italic, and the deleted parts are marked by double brackets (e.g., [[a]] represents the deletion of the character 'a').

Number

[0169] FIG. 12A shows a flowchart of an exemplary method for video processing. Referring to FIG. 12A, method 1210 includes, at step 1212, determining that motion information of a current video block is refined using an optical flow-based method in which at least one motion vector offset is derived for a region within the current video block for conversion between the current video block of the video and the coded representation of the video. Method 1210 further includes, at step 1214, clipping at least one motion vector offset to the range [-N,M], where N and M are integers based on rules. Method 1210 further includes, at step 1216, performing a conversion based on at least one motion vector offset.

[0170] FIG. 12B shows a flowchart of an exemplary method for video processing. Referring to FIG. 12B, method 1220 includes, at step 1222, selecting motion information equal to a result of applying a similarity matching function using one or more motion vector differences associated with the current video block of the video as refined motion vectors during decoder-side motion vector refinement (DMVR) operations used to refine the motion information. Method 1220 further includes, at step 1224, performing a conversion between the current video block of the video and the coded representation of the video using the refined motion vectors.

[0171] FIG. 12C shows a flowchart of an exemplary method for video processing. Referring to FIG. 12C, method 1230 includes, at step 1232, deriving motion information associated with the current video block of the video for conversion between the current video block of the video and the coded representation of the video. Method 1230 further includes, at step 1234, applying a refinement operation to the current video block including a first sub-region and a second sub-region according to rules, where the rules allow the first sub-region and the second sub-region to have different motion information from each other by the refinement operation. Method 1230 further includes, at step 1236, performing a conversion using the refined motion information of the current video block.

[0172] Figure 12D shows a flowchart of an exemplary method for video processing. Referring to Figure 12D, method 1240 includes, at step 1242, deriving motion information associated with a current video block for conversion between the current video block of the video and the coded representation of the video. Method 1240 further includes, at step 1244, determining the applicability of a refinement operation using bidirectional optical flow (BIO) based on the output of decoder-side motion vector refinement (DMVR) used to refine the motion information for a sub-region of the current video block. Method 1240 further includes, at step 1246, performing the conversion based on the determination.

[0173] Figure 12E shows a flowchart of an exemplary method for video processing. Referring to Figure 12E, method 1250 includes, at step 1252, deriving motion information associated with a current video block coded in merge mode by motion vector difference (MMVD) including a motion vector representation including a distance table that defines the distance between two motion candidates. Method 1250 further includes, at step 1254, applying decoder-side motion vector refinement (DMVR) to the current video block to refine the motion information according to rules specifying how the distance used for MMVD should be refined. Method 1250 further includes, at step 1256, performing a conversion between the current video block and the coded representation of the video.

[0174] FIG. 12F shows a flowchart of an exemplary method for video processing. Referring to FIG. 12D, method 1260 includes, at step 1262, determining the applicability of bidirectional optical flow (BDOF) for samples or sub-blocks of the current video block of the video, based on the derived motion information, where the derived motion information is refined using spatial and / or temporal gradients, according to rules. Method 1260 further includes, at step 1264, performing a conversion between the current video block of the video and the coded representation of the video, based on the determination.

[0175] FIG. 13A shows a flowchart of an exemplary method for video processing. Referring to FIG. 13A, method 1310 includes, at step 1311, deriving a motion vector difference (MVD) for conversion between the current video block of the video and the coded representation of the video. Method 1310 further includes, at step 1312, applying a clipping operation to the derived motion vector difference to generate a clipped motion vector difference. Method 1310 further includes, at step 1313, calculating the cost of the clipped motion vector difference using a cost function. Method 1310 further includes, at step 1314, determining, according to rules, that a bidirectional optical flow (BDOF) operation in which the derived motion vector difference is refined using spatial and / or temporal gradients is not allowed, based on at least one of the derived motion vector difference, the clipped motion vector difference, or the cost. Method 1310 further includes, at step 1315, performing the conversion based on the determination.

[0176] FIG. 13B shows a flowchart of an exemplary method for video processing. Referring to FIG. 13B, method 1320 includes, at step 1322, deriving a motion vector difference for conversion between a current video block of a video and a coded representation of the video. Method 1320 further includes, at step 1324, refining the derived motion vector difference based on one or more motion vector refinement tools and candidate motion vector differences (MVDs). In method 1320, at step 1326, the method further includes performing a conversion using the refined motion vector difference.

[0177] FIG. 13C shows a flowchart of an exemplary method for video processing. Referring to FIG. 13C, method 1330 includes, at step 1332, restricting a set of candidates to derived motion vector differences associated with a current video block of a video, where the derived motion vector differences are used for a refinement operation that refines motion information associated with the current video block. Method 1330 further includes, at step 1334, performing a conversion between the current video block and a coded representation of the video using the derived motion vector differences as a result of the restricting step.

[0178] FIG. 13D shows a flowchart of an exemplary method for video processing. Referring to FIG. 13D, method 1340 includes, at step 1342, applying a clipping operation using a set of clipping parameters determined based on rules during a bidirectional optical flow (BDOF) operation used to refine motion information associated with a current video block of a video unit of a video, based on utilization of the video unit and / or coding tools in the video unit. Method 1340 further includes, at step 1344, performing a conversion between the current video block and a coded representation of the video.

[0179] FIG. 13E shows a flowchart of an exemplary method for video processing. Referring to FIG. 13E, method 1350 includes, at step 1352, applying a clipping operation according to a rule to clip the x-component and / or y-component of a motion vector difference (vx, vy) during a refinement operation used to refine motion information associated with a current video block of a video. Method 1350 further includes, at step 1354, performing a conversion between the current video block and a coded representation of the video using the motion vector difference. In some implementations, the rule defines converting the motion vector difference to a value taking the form of zero or K m where m is an integer.

[0180] FIG. 13F shows a flowchart of an exemplary method for video processing. Referring to FIG. 13F, method 1360 includes, at step 1362, selecting a search region used to derive or refine motion information during a decoder-side motion derivation operation or a decoder-side motion refinement operation according to a rule for conversion between a current video block of a video and a coded representation of the video. Method 1360 further includes, at step 1364, performing a conversion based on the derived or refined motion information.

[0181] FIG. 14A shows a flowchart of an exemplary method for video processing. Referring to FIG. 14A, method 1410 includes, at step 1412, applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with a current video block by using a search region used in a refinement including an integer position of a best match for conversion between the current video block of the video and a coded representation of the video. Method 1410 further includes, at step 1414, performing a conversion using the refined motion vector difference, and the application of the DMVR operation includes deriving a sub-pel motion vector difference (MVD) according to a rule.

[0182] FIG. 14B shows a flowchart of an exemplary method for video processing. Referring to FIG. 14B, method 1420 includes, at step 1422, applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with the current video block of the video unit of the video with respect to the current video block. Method 1420 further includes, at step 1424, performing a transformation between the current video block and the coded representation of the video using the refined motion vector difference. In some implementations, the application of the DMVR operation includes determining whether to allow or disallow sub-pel motion vector difference (MVD) derivation depending on the use of bidirectional optical flow (BDOF) for the video unit.

[0183] FIG. 14C shows a flowchart of an exemplary method for video processing. Referring to FIG. 14C, method 1430 includes, at step 1432, deriving a motion vector difference of a first sample of the current video block of the video during a refinement operation using optical flow. Method 1430 further includes, at step 1434, determining a motion vector difference of a second sample based on the derived motion vector difference of the first sample. Method 1430 further includes, at step 1436, performing a transformation between the current video block and the coded representation of the video based on the determination.

[0184] FIG. 14D shows a flowchart of an exemplary method for video processing. Referring to FIG. 14D, method 1440 includes, at step 1442, deriving a prediction refinement sample by applying a refinement operation using a bidirectional optical flow (BDOF) operation to the current video block of the video. Method 1440 further includes, at step 1444, determining the applicability of applying a clipping operation to clip the derived prediction refinement sample to a predetermined range [-M, N] according to a rule, where M and N are integers. Method 1440 further includes, at step 1446, performing a transformation between the current video block and the coded representation of the video.

[0185] FIG. 14E shows a flowchart of an exemplary method for video processing. Referring to FIG. 14E, method 1450 includes, at step 1452, determining a coding group size for a current video block of a video, the current video block including a first coding group and a second coding group that are coded using different residual coding modes such that the first coding group and the second coding group are aligned according to a rule. Method 1450 further includes, at step 1454, performing a conversion between the current video block and a coded representation of the video based on the determination.

[0186] FIG. 14F shows a flowchart of an exemplary method for video processing. Referring to FIG. 14F, method 1460 includes, at step 1462, determining the applicability of a predictive refinement optical flow (PROF) tool in which motion information is refined using optical flow based on coded information and / or decoded information associated with a current video block for a conversion between the current video block of a video and a coded representation of the video. Method 1460 further includes, at step 1464, performing the conversion based on the determination.

[0187] FIG. 15 is a block diagram of a video processing apparatus 1500. The apparatus 1500 may be used to implement one or more of the methods described herein. The apparatus 1500 may be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The apparatus 1500 may include one or more processors 1502, one or more memories 1504, and video processing hardware 1506. The processor 1502 may be configured to implement one or more of the methods described herein (including, but not limited to, the methods shown in FIGS. 12 - 14F). The memory (or memories) 1504 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 1506 is a hardware circuit and may be used to implement some of the techniques described herein.

[0188] FIG. 16 is another example of a block diagram of a video processing system in which the disclosed technology may be implemented. FIG. 16 is a block diagram showing an exemplary video processing system 1610 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of the system 1610. The system 1610 may include an input section 1612 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 - or 10 - bit multi - component pixel values, or may be in a compressed or encoded format. The input section 1612 may correspond to a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include priority interfaces such as Ethernet®, Passive Optical Network (PON), etc., and wireless interfaces such as Wi - Fi or cellular interfaces.

[0189] System 1610 may include a coding component 1614 that can implement various coding or encoding methods described herein. The coding component 1614 may reduce the average bit rate of the video from the input section 1612 to the output section of the coding component 1614 to generate a coded representation of the video. Coding techniques are thus sometimes referred to as video compression or video transcoding techniques. The output of the coding component 1614 may either be stored, as represented by component 1616, or transmitted via a communication connection. The stored or communicated bitstream (or coded) representation of the video received at the input section 1612 may be used by component 1618 to generate pixel values or a displayable video that is sent to the display interface 1620. The process of generating a video that a user can view from the bitstream representation is sometimes referred to as video decompression. Further, certain video processing operations are referred to as "coding" operations or tools, but coding tools or operations are used in an encoder, and the corresponding decoding tools or operations that reverse the result of the coding are to be executed in a decoder.

[0190] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB) or a High Definition Multimedia Interface (HDMI (registered trademark)). Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, an IDE interface, and the like. The techniques described herein may be embodied in various electronic devices such as a cellular phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.

[0191] Some embodiments of the disclosed technology include making a determination or judgment to enable a video processing tool or mode. In an example, when a video processing tool or mode is enabled, the encoder will use the tool or mode in processing video blocks, but it is not necessary to change the resulting bitstream based on the utilization of the tool or mode. That is, the conversion from a video block to a video bitstream representation uses the video processing tool or mode when it is enabled based on a determination or judgment. In other examples, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been changed based on the video processing tool or mode. That is, the conversion from a video bitstream representation to a video block is performed using the video processing tool or mode enabled based on a determination or judgment.

[0192] Some embodiments of the disclosed technology include making a determination or judgment to disable a video processing tool or mode. In an example, when a video processing tool or mode is disabled, the encoder does not use the tool or mode in converting a video block to a video bitstream representation. In other examples, when a video processing tool or mode is disabled, the decoder processes the bitstream knowing that the bitstream has not been changed using the video processing tool or mode disabled based on a determination or judgment.

[0193] In this specification, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may correspond to bits that are either at the same position within the bitstream or spread over different locations, as defined by the syntax. For example, a video block may be encoded using the transformed and coded error residual values, and further, bits in the headers and other fields within the bitstream. Here, a video block is a group of pixels corresponding to an operation, such as a coding unit, a transform unit, or a prediction unit, etc.

[0194] Various techniques and embodiments may be described using the following bullet points.

[0195] The first set of bullet points describes specific features and aspects of the techniques disclosed in the previous section.

[0196] 1. A method of video processing, comprising: deriving motion information associated with the video processing unit during conversion between the video processing unit and its bitstream representation; determining whether to refine the motion information for a sub-region of the video processing unit to which a refinement operation is to be applied; applying the refinement operation to the video processing unit based on the determination to refine the motion information. A method having the above steps.

[0197] 2. The method according to item 1, wherein: the refinement operation includes a decoder-side motion vector derivation operation, a decoder-side motion vector refinement (DMVR) operation, a sample refinement operation, or a prediction refinement operation based on optical flow. A method.

[0198] The method according to item 1, wherein the determining step is based on coding information associated with the video processing unit; Method.

[0199] The method according to item 3, wherein the coding information associated with the video processing unit includes at least one of a motion vector, a picture order count (POC) value of a reference picture, a merge index, a motion vector predictor (MVP) index, or whether an SBT mode is used by the video processing unit; Method.

[0200] The method according to item 1, wherein the determining step is based on the relationship between a first motion vector referring to a first reference list and a second motion vector referring to a second reference list; the first motion vector and the second motion vector are obtained before the application of the refinement operation; Method.

[0201] The method according to item 1, wherein the application of the refinement operation is performed in multiple stages, and the refined area has a size determined based on the number of steps of the multiple stages; Method.

[0202] The method according to item 1, wherein the application of the refinement operation includes applying the refinement operation to a first sub-region; and applying the refinement operation to a second sub-region different from the first sub-region; and the refined motion information in the video processing unit is different in the first sub-region and the second sub-region; Method.

[0203] 8. The method according to item 1, further comprising the step of determining whether to apply bidirectional optical flow (BIO) based on the result of the refinement operation, method.

[0204] 9. The method according to item 1, wherein the application of the refinement operation includes the step of selecting N pairs of MVDs (Motion Vector Differences) to generate a final prediction sample, N is a natural number greater than 1, and each pair includes two motion vector differences from different lists, method.

[0205] 10. The method according to item 1, wherein the refinement operation is used together with MMVD (merge with MVD), and the MMVD distance is held as a specific value during the refinement operation, method.

[0206] 11. The method according to item 1, further comprising the step of disabling BDOF for a sample or sub-region of the video processing unit based on at least one of the derived MVD, spatial gradient, or temporal gradient, method.

[0207] 12. The method according to item 1, further comprising the step of calculating a cost for the clipped MVD when the derived MVD in BDOF changes after the clipping operation, method.

[0208] 13. The method according to item 1, wherein the conversion includes generating the bitstream representation from the current video processing unit or generating the current video processing unit from the bitstream representation, method.

[0209] 14. A method for video processing, comprising: obtaining motion information associated with a video processing unit; determining whether to refine the motion information of the video processing unit for a sub-region of the video processing unit, wherein the video processing unit is a unit to which a refinement operation is applied; using the motion information without applying the refinement operation to the video processing unit based on the determination; and a method having the above steps.

[0210] 15. The method according to item 14, wherein the refinement operation includes a decoder-side motion vector derivation operation, a decoder-side motion vector refinement (DMVR) operation, a sample refinement operation, or a prediction refinement operation based on optical flow. A method.

[0211] 16. The method according to item 14, wherein the determining step is based on coding information associated with the video processing unit, the coding information includes at least one of a motion vector, a picture order count (POC) value of a reference picture, a merge index, a motion vector predictor (MVP) index, or whether an SBT mode is used by the video processing unit. A method.

[0212] 17. The method according to item 14, wherein the refinement operation including bidirectional optical flow (BIO) is omitted when the video processing unit is coded in advanced motion vector prediction (AMVP) mode. A method.

[0213] 18. The method according to item 14, wherein the refinement operation is omitted based on the sum of two motion vectors or the absolute value of the MV component. Method

[0214] 19. The method according to clause 14, wherein the refinement operation is omitted when the SBT is used by the video processing unit, Method

[0215] 20. The method according to clause 14, wherein the determining step is based on the relationship between a first motion vector referring to a first reference list and a second motion vector referring to a second reference list, wherein the first motion vector and the second motion vector are obtained before the application of the refinement operation, Method

[0216] 21. The method according to clause 14, further comprising the step of disabling BDOF for a sample or sub-region of the video processing unit based on at least one of the derived MVD, spatial gradient, or temporal gradient, Method

[0217] 22. The method according to clause 14, further comprising the step of calculating a cost for the clipped MVD when the derived MVD in the BDOF changes after the clipping operation, Method

[0218] 23. A method for video processing, during conversion between a current video block and a bitstream representation of the current video block, refining a derived motion vector difference based on one or more vector refinement tools and a candidate motion vector difference (MVD); and executing the conversion using the derived motion vector difference and a motion vector signaled in the bitstream representation, Method having

[0219] 24. The method according to clause 23, The refining step has a two-step process having a first step in which an MVD with a changed y-component and an unchanged x-component is used, and a second step in which an MVD with a changed x-component and an unchanged y-component is used. Method.

[0220] 25. The method according to clause 23, wherein the candidate MVD includes (v x + offsetX, v y ), offsetX is within the range [-ThX1, ThX2], ThX1 and ThX2 are non-negative integer values, and the derived MVD is (v x , v y ). Method.

[0221] 26. The method according to clause 23, wherein the candidate MVD includes (v x , v y + offsetY), offsetY is within the range [-ThY1, ThY2], ThY1 and ThY2 are non-negative integer values, and the derived MVD is (v x , v y ). Method.

[0222] 27. The method according to any one of clauses 23 to 26, wherein the conversion uses a bidirectional optical flow coding tool. Method.

[0223] 28. A method for video processing, comprising restricting a derived motion vector difference (DMVD) to a candidate set during conversion between a current video block and a bitstream representation of the current video block, and performing the conversion using a result of restricting the DMVD to the candidate set. Method.

[0224] 29. The method according to clause 28, wherein the candidate set includes candidates (v x , v y ), v x and / or v y is in the form of K m , where m is an integer,[ method.

[0225] 30. A method for video processing, comprising determining that a motion vector difference is clipped to the range [-N, M] for conversion between a bitstream representation of a current video block of a video and the current video block; the determination of clipping is performed based on rules as further described herein; the method further includes performing the conversion based on the determination, and the rules define using the same range for conversion using a prediction refined optical flow (PROF) and a bidirectional optical flow tool.

[0226] 31. The method according to clause 30, wherein N = M = 31. method.

[0227] 32. The method according to clause 30, wherein N = M = 255. method.

[0228] 33. A method for video processing, comprising determining a motion vector difference of a first sample based on a motion vector difference of a second sample of the current video block for conversion between a bitstream representation of a current video block of a video and the current video block; performing the conversion based on the determination; and the conversion is based on an optical flow coding or decoding tool. method.

[0229] 34. The method according to clause 33, wherein the first sample and the second sample use different optical flow coding tools. Method.

[0230] 35. A method for video processing, comprising: determining whether to enable the use of a predictive refined optical flow (PROF) tool for conversion between the bitstream representation of a current video block and the current video block, wherein the conversion uses an optical flow coding tool; and executing the conversion based on the determination. Method.

[0231] 36. The method according to clause 35, wherein the determining step disables the use of the PROF tool because a reference picture used in the conversion has dimensions different from those of the current picture including the current video block. Method.

[0232] 37. The method according to any one of clauses 35 to 36, wherein the determining step enables the PROF tool because the conversion uses bi-predictive prediction and both reference pictures have the same dimensions. Method.

[0233] 38. The method according to any one of clauses 1 to 37, wherein the conversion includes generating the bitstream representation by encoding the current video block. Method.

[0234] 39. The method according to any one of clauses 1 to 37, wherein the conversion includes generating the current video block by decoding the bitstream representation. Method.

[0235] 40. An apparatus in a video system, comprising a processor and a non - transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 39. Apparatus.

[0236] 41. A computer program product stored in a non - transitory computer - readable medium, comprising program code for performing the method according to any one of Items 1 to 39.

[0237] The second set of items describes specific features and aspects of the technology disclosed in the previous section.

[0238] 1. A video processing method, determining that motion information of a current video block is refined using an optical - flow - based method in which at least one motion - vector offset is derived for a region within the current video block for conversion between the current video block of the video and the coded representation of the video; clipping the at least one motion - vector offset to the range [ - N,M], where N and M are integers based on a rule; and performing the conversion based on the at least one motion - vector offset. A method having the above steps.

[0239] 2. The method according to Item 1, wherein the motion - vector offset is used to refine reference samples for the region within the current video block. Method.

[0240] 3. The method according to Item 1, The method based on the optical flow has at least one of a prediction refined optical flow (PROF) method applied to a block by an affine motion model or a bidirectional optical flow (BDOF) method. Method.

[0241] 4. The method according to item 1, wherein the region in the current video block is the whole of the current video block or a sub-block within the current video block. Method.

[0242] 5. The method according to item 1, wherein the rule determines to use the same range or different ranges for the conversion using a PROF method in which the at least one motion vector offset is derived using optical flow calculation and a BDOF method in which the at least one motion vector offset is derived using a spatial gradient and / or a temporal gradient. Method.

[0243] 6. The method according to item 1, wherein the motion vector offset has a horizontal component and a vertical component, and the clipping of the at least one motion vector offset has clipping the vertical component and / or the horizontal component to the range. Method.

[0244] 7. The method according to item 1, wherein N and M have the same value not equal to 2 K and K is an integer. Method.

[0245] 8. The method according to item 7, wherein the same value is equal to (2 K -1) and K is an integer. Method.

[0246] 9. The method according to item 7, wherein The same value is one of 31, 63, 15, 7, 3, 127, or 255. Method.

[0247] The method according to item 17, M is equal to N, 2×M is not equal to 2 K the value of, and K is an integer. Method.

[0248] The method according to any one of items 1 to 10, The rule determines that the motion vector offset is clipped to one of [-31, 31], [-15, 15], [-63, 63], or [-127, 127] for the conversion using the BDOF. Method.

[0249] The method according to any one of items 1 to 10, The rule determines that the motion vector offset is clipped to one of [-63, 63] or [-31, 31] for the conversion using the PROF. Method.

[0250] The method according to item 1, M is not equal to N, and the sum of M and N is not equal to 2 K the value of, and K is an integer. Method.

[0251] A 14 - video processing method, During the decoder - side motion vector refinement (DMVR) operation used to refine motion information, as the refined motion vector, selecting motion information equal to the result of applying a similarity matching function using one or more motion vector differences associated with the current video block of the video; Executing a conversion between the current video block and the coded representation of the video using the refined motion vector and having a method.

[0252] 15. The method according to item 14, wherein the similarity matching function includes the mean removed sum of absolute differences (MRSAD), the sum of absolute differences (SAD), or the sum of absolute transform differences (SATD). Method.

[0253] 16. The method according to item 14, wherein the motion information is changed by performing one of multiplication, division, addition, or subtraction using a scaling factor. Method.

[0254] 17. The method according to item 16, wherein the scaling factor depends on the motion vector difference (MVD) allowed within the range [-M, N] according to a rule, and M and N are integers greater than 0. Method.

[0255] 18. A video processing method, comprising the steps of making a determination regarding the application of a refinement operation for refining motion information based on the characteristics of a current video block of a video, and performing a conversion between the current video block and the coded representation of the video based on the determination. Method having the above.

[0256] 19. The method according to item 18, wherein the refinement operation includes at least one of a decoder-side motion vector derivation operation, a decoder-side motion vector refinement (DMVR) operation, a sample refinement operation, or a prediction refinement operation based on optical flow. Method.

[0257] 20. The method according to item 18, wherein the characteristics of the current video block correspond to the coding information of the current video block. Method.

[0258] 21. The method according to item 20, wherein the coding information of the current video block includes at least one of a motion vector, a picture order count (POC) value of a reference picture, a merge index, a motion vector predictor (MVP) index, or whether a sub-block transform (SBT) mode is used by the current video block. Method.

[0259] 22. The method according to item 20, wherein the coding information of the current video block includes whether the current video block is coded in advanced motion vector prediction (AMVP) mode or merge mode. Method.

[0260] 23. The method according to item 20, wherein the refinement operation corresponds to a decoder-side motion vector refinement (DMVR) operation and / or a bidirectional optical flow (BDOF) operation, the refinement operation is enabled depending on the sum of two motion vectors or the absolute value of a motion vector. Method.

[0261] 24. The method according to item 20, wherein the refinement operation corresponds to a decoder-side motion vector refinement (DMVR) operation and / or a bidirectional optical flow (BDOF) operation, the refinement operation is permitted according to a rule specifying a merge index or a motion vector predictor (MVP) index. Method.

[0262] 25. The method according to item 20, wherein the refinement operation corresponds to a decoder-side motion vector refinement (DMVR) operation and / or a bidirectional optical flow (BDOF) operation, the refinement operation is not permitted when a sub-block transform (SBT) mode is used by the current video block. Method

[0263] 26. The method according to item 20, wherein the refinement operation corresponds to a decoder-side motion vector refinement (DMVR) operation and / or a bidirectional optical flow (BDOF) operation, the refinement operation is whether or not permitted for a sub-partition including a non-zero residual in a sub-block transform (SBT) mode, Method

[0264] 27. The method according to item 18, wherein the characteristic of the current video block corresponds to a relationship between a first motion vector referring to a first reference list and a second motion vector referring to a second reference list, the first motion vector and the second motion vector are acquired before application of the refinement operation, Method

[0265] 28. The method according to item 27, wherein the refinement operation is invalidated according to a rule according to a degree of symmetry between the first motion vector and the second motion vector being symmetric, Method

[0266] 29. The method according to item 28, wherein the degree of symmetry is i) determined to be symmetric when the sum of MV0 and MV1 has only zero components, ii) determined to be approximately symmetric when abs(MV0 + MV1) < Th1, or iii) determined to be non-symmetric when abs(MV0 + MV1) > Th2 wherein MV0 and MV1 respectively correspond to the first motion vector and the second motion vector, Th1 and Th2 are integers, Method

[0267] 30. The method according to item 28, wherein the degree of symmetry is determined based on the picture order count (POC) value of the reference picture of the current video block; Method.

[0268] 31. The method according to item 18, assuming that N is an integer greater than 0, the determining step determines to apply the refinement operation in a plurality of steps such that the size of the refinement region of the Nth refinement step decreases as the value of N increases; Method.

[0269] 32. The method according to item 18, assuming that N is an integer greater than 1, the determination includes applying the refinement operation by selecting N pairs of motion vector differences to generate a final predicted sample, each pair of the motion vector differences includes two motion vector differences for different lists; Method.

[0270] 33. The method according to item 32, N is 2; Method.

[0271] 34. The method according to item 32, wherein the pairs of the motion vector differences are selected for each sub-region of the current video block; Method.

[0272] 35. A video processing method, comprising deriving motion information associated with a current video block of a video for conversion between the current video block of the video and a coded representation of the video; Applying a refinement operation to the current video block including the first sub-region and the second sub-region according to a rule, the rule allowing the first sub-region and the second sub-region to have different motion information from each other by the refinement operation, the applying step; Executing the conversion using the refined motion information of the current video block; A method comprising:

[0273] 36. The method according to item 35, wherein the first sub-region has decoded motion information; wherein the second sub-region has refined motion vectors derived during the refinement operation applied to the video region covering the second sub-region; A method.

[0274] 37. A video processing method, Deriving motion information associated with the current video block for conversion between the current video block of the video and the coded representation of the video; Determining the applicability of a refinement operation using bidirectional optical flow (BIO) based on the output of decoder-side motion vector refinement (DMVR) used to refine the motion information for sub-regions of the current video block; Executing the conversion based on the determination; A method comprising:

[0275] 38. The method according to item 37, wherein the determining step determines to disable the refinement operation using the BIO when the motion information refined by the DMVR remains unchanged for the sub-region; A method.

[0276] 39. A video processing method, Deriving motion information associated with a current video block coded in a merge mode based on motion vector differences (MMVD) that includes a motion vector representation including a distance table that defines distances between two motion candidates; Applying decoder-side motion vector refinement (DMVR) to the current video block to refine the motion information according to rules specifying how to refine the distances used for the MMVD; Performing a conversion between the current video block and the coded representation of the video; A method having the above steps.

[0277] 40. The method according to clause 39, wherein the rule determines that the distance refined by the DMVR is not included in the distance table; A method.

[0278] 41. The method according to clause 39, wherein the rule determines that the DMVR is not allowed when the distance refined by the DMVR or the coarsely refined distance is included in the distance table, wherein the coarsely refined distance is generated by adding the best integer motion vector difference to the distance; A method.

[0279] 42. The method according to clause 39, wherein the rule determines that the DMVR is allowed for all distances under specific conditions; A method.

[0280] 43. A video processing method, for samples or sub-blocks of a current video block of a video, determining the applicability of bidirectional optical flow (BDOF) that is refined using spatial and / or temporal gradients based on derived motion information according to rules; Based on the determination, performing a conversion between the current video block and the coded representation of the video A method having the above.

[0281] 44. The method according to item 43, wherein the derived motion information includes at least one of a motion vector difference, a spatial gradient, or a temporal gradient A method.

[0282] 45. The method according to item 43, wherein the rule determines that the BDOF is not allowed for the sample by the derived motion information that increases the difference between the predicted sample in the first reference list and the predicted sample in the second reference list A method.

[0283] 46. The method according to item 43, wherein the rule determines that the BDOF is not allowed for the sample or the sub-block by the derived motion information that increases the difference between the predicted sub-block in the first reference list and the predicted sub-block in the second reference list A method.

[0284] 47. The method according to item 43, wherein the derived motion information corresponds to a derived offset, and the derived offset is scaled by a coefficient used for the determining step A method.

[0285] 48. A video processing method, comprising deriving a motion vector difference (MVD) for conversion between a current video block of a video and the coded representation of the video; applying a clipping operation to the derived motion vector difference to generate a clipped motion vector difference Calculating the cost of the clipped motion vector difference using a cost function; Determining that a bidirectional optical flow (BDOF) operation in which the derived motion vector difference is refined using spatial and / or temporal gradients is not allowed based on at least one of the derived motion vector difference, the clipped motion vector difference, or the cost according to a rule; Performing the transformation based on the determination; A method comprising.

[0286] The method according to item 48, The calculating step is performed when the derived motion vector difference (vx, vy) is changed to (clipVx, clipVy) after the clipping operation, (clipVx, clipVy) is different from (vx, vy), A method.

[0287] The method according to item 48, The rule is, 1) costClipMvd > Th × costZeroMvd, or 2) costClipMvd > Th × costDerivedMvd It is determined that the BDOF is not allowed when the following conditions are met, costDerivedMvd and costClipMvd respectively correspond to the costs of the derived motion vector difference and the clipped motion vector difference, A method.

[0288] The method according to item 48, The rule determines that the BDOF is not allowed when at least one of vx and vy is changed after the clipping operation, (vx, vy) is the derived motion vector difference, A method.

[0289] The method according to clause 48, wherein the rule determines that the BDOF is not allowed based on at least one of the absolute differences of the x and y components of the derived motion vector difference (vx, vy) and the clipped motion vector difference (clipVx, clipVy). Method.

[0290] The method according to clause 48, wherein the step of calculating the cost is performed for a specific motion vector difference, and the final motion vector difference is determined as having the minimum cost value. Method.

[0291] The method according to clause 48, wherein the horizontal motion vector difference and the vertical motion vector difference are determined when the motion vector difference changes after the clipping operation. Method.

[0292] The method according to clause 48, wherein for a plurality of motion vector differences having a common part in the cost function, the calculating step is performed so as not to repeat the common part. Method.

[0293] The method according to clause 48, wherein for a plurality of motion vector differences having a common part in the cost function, the common part is removed from the cost function before the calculating step. Method.

[0294] A video processing method, comprising deriving a motion vector difference for conversion between a current video block of a video and a coded representation of the video; and refining the derived motion vector difference based on one or more motion vector refinement tools and candidate motion vector differences (MVD). performing the transformation using the refined motion vector difference; A method comprising the steps of:

[0295] 58. The method according to item 57, wherein the refining step is performed to include a first step in which an MVD with a changed y-component and an unchanged x-component is used, and a second step in which an MVD with a changed x-component and an unchanged y-component is used; A method.

[0296] 59. The method according to item 57, wherein the candidate MVD includes (vx + offsetX, vy), offsetX is within the range [-ThX1, ThX2], ThX1 and ThX2 are non-negative integer values, the derived motion vector difference is (vx, vy); A method.

[0297] 60. The method according to item 57, wherein the candidate MVD includes (vx, vy + offsetY), offsetY is within the range [-ThY1, ThY2], ThY1 and ThY2 are non-negative integer values, the derived motion vector difference is (vx, vy); A method.

[0298] 61. The method according to item 57, wherein the derived motion vector difference is (vx, vy), either vx or vy is refined first; A method.

[0299] 62. The method according to item 57, wherein the derived motion vector difference is (vx, vy), vx and vy are refined together; A method.

[0300] 63. The method according to clause 57, wherein the derived motion vector difference is (vx, vy), and the same number of candidate MVDs are derived from vx and vy, Method.

[0301] 64. The method according to clause 57, wherein the refined motion vector difference results in a minimum cost value obtained using a cost function, Method.

[0302] 65. A video processing method, comprising: limiting the derived motion vector difference associated with the current video block of the video to a candidate set, wherein the derived motion vector difference is used for a refinement operation to refine motion information associated with the current video block, the limiting step; and performing a conversion between the current video block and the coded representation of the video using the derived motion vector difference as a result of the limiting step Method having.

[0303] 66. The method according to clause 65, wherein the candidate set includes the number of permitted horizontal MVD candidates less than the value of (1 + MVDmaxX - MVDminX), and / or the number of permitted vertical MVD candidates less than the value of (1 + MVDmaxY - MVDminY), the maximum and minimum values of the horizontal MVD are represented by MVDmaxX and MVDminX, respectively, the maximum and minimum values of the vertical MVD are represented by MVDmaxY and MVDminY, respectively, Method.

[0304] 67. The method according to clause 65, wherein the candidate set includes candidates (vx, vy) obtained using an equation during the refinement operation, The candidate satisfies: 1) Assuming m is an integer, at least one of vx and vy takes the form of K m or 2) at least one of vx and vy is zero satisfies method.

[0305] 68. The method according to clause 65, wherein the candidate set includes candidates (Δv x (,y), Δv y (x,y)) obtained using an equation during the refinement operation, the candidate satisfies: 1) Assuming m is an integer, at least one of Δv x (x,y) and Δv y (x,y) takes the form of K m or 2) at least one of Δv x (x,y) and Δv y (x,y) is zero satisfies method.

[0306] 69. The method according to clause 65, wherein the candidate set includes candidates (vx, vy) or (Δv x (x,y), Δv y (x,y)) obtained using an equation during the refinement operation, the candidate satisfies: 1) the absolute value of at least one of vx, vy, Δv x (x,y) and Δv y (x,y) is zero, or 2) assuming m is an integer, at least one of vx, vy, Δv x (x,y) and Δv y (x,y) takes the form of K m satisfies method.

[0307] 70. The method according to clause 65,​ The candidate set includes candidates (vx, vy) or (Δv x (x,y), Δv y (x,y)) obtained using an equation during the refinement operation, wherein the candidates 1) satisfy that at least one of vx, vy, Δv x (x,y) and Δv y (x,y) is zero, or 2) assuming m is an integer, at least one of vx, vy, Δv x (x,y) and Δv y (x,y) takes the form of K m or -K m ; A method. Method.

[0308] The method according to item 65, clause 71, wherein the derived motion vector difference is first derived as the initially derived motion vector difference without being restricted to the candidate set, and then restricted to the candidate set; Method.

[0309] The method according to item 71, clause 72, wherein the initially derived motion vector difference and the modified motion vector difference are represented by X and X', X is equal to abs(vx) or abs(vy), and X' is obtained from sign(vx) or sign(vy), where abs(t) corresponds to the absolute value function that returns the absolute value of t, and sign(t) corresponds to the sign function that returns a value according to the sign of t; Method.

[0310] The method according to item 71, clause 73, wherein the candidate set is a power of K and has a value that depends on at least one of vx and vy, and the derived motion vector difference is (vx, vy); Method.

[0311] The method according to item 74, wherein the derived motion vector difference is (vx, vy), and at least one of vx and vy is changed first, Method.

[0312] The method according to item 75, wherein the initially derived motion vector difference is changed by using the logK(x) function and the ceil(y)) function, logK(x) returns the logarithm of x with base K, and ceil(y) returns the smallest integer greater than or equal to y, Method.

[0313] The method according to item 76, wherein the initially derived motion vector difference is changed by using the floor(y) function, floor(y) returns the largest integer less than or equal to y, Method.

[0314] The method according to item 77, wherein the initially derived motion vector difference is changed by using a predefined look-up table, Method.

[0315] The method according to item 78, wherein only candidates within the candidate set are checked to derive a motion vector difference, Method.

[0316] The method according to item 79, wherein the candidate set includes candidate motion vector differences (MVDs) in the form of K N , -K N , or zero, where K is an integer greater than 0 and N is an integer, Method.

[0317] 80. The method according to clause 79, wherein the candidate set includes candidate MVDs in the form of M / K N where M is an integer, M is an integer, Method.

[0318] 81. The method according to clause 79, wherein the candidate set includes candidate MVDs in the form of KM / KN or zero, where M and N are integers, Method.

[0319] 82. The method according to clause 65, wherein the derived motion vector difference is derived to have an x component and / or a y component having values allowed according to a rule, Method.

[0320] 83. The method according to clause 82, wherein the rule is such that the derived motion vector difference (vx, vy) has a value obtained using a floor function, a ceiling function, or a sign function, the floor function floor(t) returns the largest integer less than or equal to t, the ceiling function ceil(t) returns the smallest integer greater than or equal to t, and the sign function sign(t) returns a value according to the sign of t, Method.

[0321] 84. The method according to clause 82, wherein the rule is i) vx is derived by using at least one of a numerator (numX) and a denominator (denoX), and / or ii) vy is derived by using at least one of a numerator (numY) and a denominator (denoY) such that the derived motion vector difference (vx, vy) is obtained, Method.

[0322] 85. The method according to clause 84, wherein The rule determines to use the numerator and / or the denominator as a variable used in a floor function, a ceiling function, or a sign function. The floor function floor(t) returns the largest integer less than or equal to t, the ceiling function ceil(t) returns the smallest integer greater than or equal to t, and the sign function sign(t) returns a value according to the sign of t. Method.

[0323] 86. The method according to clause 84, The rule determines that vx is derived as sign(numX × denoX) × esp2(M). M is an integer that realizes the minimum value of cost(numX, denoX, M). sign(t) returns a value according to the sign of t. Method.

[0324] 87. The method according to clause 84, The rule determines that vx is equal to zero or is obtained as using the formula: sign(numX × denoX) × K F1(logK(abs(numX)+offset1))-F2(logK(abs(denoX)+offset2))2 where offset1 and offset2 are integers. sign(t) returns a value according to the sign of t. sign(t) returns a value according to the sign of t. Method.

[0325] 88. The method according to clause 87, The rule is that i) denoX is equal to zero, ii) numX is equal to zero, iii) numX is equal to zero or denoX is equal to zero, iv) denoX is less than T1, v) numX is less than T1, or vi) numX is less than T1 or denoX is less than T2 in which case vx is equal to 0. T1 and T2 are integers. Method

[0326] 89. The method according to item 87, wherein said offset1 depends on numX Method

[0327] 90. The method according to item 87, wherein said offset1 is equal to K floor(logK(abs(numX))) / 2 + P, where P is an integer, floor(t) returns the largest integer less than or equal to t, logK(x) returns the logarithm of x to the base K, and abs(t) returns the absolute value of t Method

[0328] 91. The method according to item 87, wherein said offset2 depends on denoX Method

[0329] 92. The method according to item 87, wherein said offset1 is equal to K floor(logK(abs(denoX))) / 2 + P, where P is an integer, floor(t) returns the largest integer less than or equal to t, logK(x) returns the logarithm of x to the base K, and abs(t) returns the absolute value of t Method

[0330] 93. The method according to item 84, wherein said rule is that vx depends on the value of delta obtained as delta = floor(log2(abs(numX))) - floor(log2(abs(denoX))) It is determined that floor(t) returns the largest integer less than or equal to t, logK(x) returns the logarithm of x to the base K, and abs(t) returns the absolute value of t floor(t) returns the largest integer less than or equal to t, logK(x) returns the logarithm of x to the base K, and abs(t) returns the absolute value of t Method

[0331] 94. The method according to item 93, wherein said delta is 1) floor(log(abs(numX))) is greater than 0 and abs(numX) & (1 << (floor(log2(abs(numX)))-1)) is not equal to zero, or, 2) floor(log2(abs2(nodeX))) is greater than 0 and abs(denoX) & (1 << (floor(log2(abs(numX)))-1)) is not equal to zero In the case, it is changed, Method.

[0332] 95. The method according to clause 93, vx is set to 0 when delta is less than 0, Method.

[0333] 96. The method according to any one of clauses 83 to 95, At least one of vx and vy is used to derive a parameter indicating a predicted sample value used in BDOF, or a parameter indicating a luma prediction refinement used in PROF, Method.

[0334] 97. The method according to clause 96, The parameter indicating the predicted sample value used in the BDOF is the formula bdofOffset = Round((vx × (gradientHL1[x + 1][y + 1] - gradientHL0[x + 1][y + 1])) >> 1) + Round((vy × (gradientVL1[x + 1][y + 1] - gradientVL0[x + 1][y + 1])) >> 1) It is bdofOffset derived using Method.

[0335] 98. The method according to clause 96, The parameter indicating the luma prediction refinement used in the PROF is the formula ΔI derived using ΔI(posX,posY)=dMvH[posX][posY]×gradientH[posX][posY]+dMvV[posX][posY]×gradientV[posX][posY], dMvH[posX][posY] and dMvV[posX][posY] correspond to vx and vy, Method.

[0336] The method according to item 65, At least one of the horizontal and vertical components of the derived motion vector difference is derived according to the tool used for the refinement operation and is changed to be within the candidate set related to the tool, Method.

[0337] The method according to item 99, At least one of the horizontal and vertical components of the derived motion vector difference is changed by a unified change process regardless of the tool, Method.

[0338] The method according to item 65, The cost is calculated using a cost function for each candidate value, The candidate value corresponding to the minimum cost value is selected as the final motion vector difference (vx, vy), Method.

[0339] The method according to item 65, The candidate set is predefined, Method.

[0340] The method according to item 102, The candidate set includes candidates (vx, vy) such that at least one of vx and vy takes the form of -2 x / 2 X and x is within the range [-(X - 1), (X - 1)], x is within the range [-(X - 1), (X - 1)], Method

[0341] 104. The method according to clause 102, wherein the candidate set is such that vx and vy are i) {-32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32} / 64, ii) {-32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32}, iii) {-32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32} / 32, iv) {-64, -32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32, 64}, v) {-64, -32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32, 64} / 64, vi) {-128, -64, -32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32, 64, 128} / 128, or vii) {-128, -64, -32, -16, -8, -4, -2, -1, 0, 1, 2, 4, 8, 16, 32, 64, 128} and the candidate (vx, vy) is included only from one of the sets containing Method

[0342] 105. The method according to clause 102, wherein the candidate set includes the candidate (vx, vy) such that the horizontal and / or vertical components of the candidate are in the form of S×(2 m / 2 n ), S is 1 or -1, and m and n are integers, Method

[0343] 106. The method according to clause 65, wherein the candidate set depends on the coded information of the current video block, Method

[0344] 107. The method according to clause 65, The candidate set is signaled at the video unit level including a sequence, a picture, a slice, a tile, a brick, or other video regions. Method.

[0345] The method according to item 65, For samples between bidirectional optical flow (BDOF) or predictive refined optical flow (PROF), the refinement operation excludes multiplication operations. Method.

[0346] A video processing method, During the bidirectional optical flow (BDOF) operation used to refine the motion information associated with the current video block of a video unit of a video, applying a clipping operation using a set of clipping parameters determined based on the video unit and / or the use of coding tools in the video unit according to rules; Executing a conversion between the current video block and the coded representation of the video And a method having.

[0347] The method according to item 109, The rule determines to use different sets of clipping parameters for different video units of the video. Method.

[0348] The method according to item 109 or 110, The video unit corresponds to a sequence, a picture, a slice, a tile group, a tile, a coding tree unit, or a coding unit. Method.

[0349] The method according to item 109, The horizontal motion vector difference and the vertical motion vector difference are clipped to different thresholds from each other. Method.

[0350] 113. The method according to item 109, wherein the threshold value used in the clipping operation depends on the dimensions of the picture including the current video block and / or the decoded motion information of the current video block. Method.

[0351] 114. The method according to item 109, wherein the threshold value used in the clipping operation is signaled by a sequence parameter set (SPS), a video parameter set (VPS), an adaptive parameter set (APS), a picture parameter set (PPS), a slice header, and / or a tile group header. Method.

[0352] 115. The method according to item 109, wherein the rule defines determining the set of clipping parameters based on the use of decoder-side motion vector difference (DMVR) or affine inter mode. Method.

[0353] 116. The method according to item 109, wherein the value of the clipping parameter increases or decreases according to whether the coding tool is invalidated for the video unit. Method.

[0354] 117. A video processing method, comprising: applying a clipping operation according to a rule to clip the x component and / or y component of a motion vector difference (vx, vy) during a refinement operation used to refine motion information associated with a current video block of a video; executing a conversion between the current video block and the coded representation of the video using the motion vector difference; and the rule sets the motion vector difference to zero or K before or after the clipping operation.m is defined to be converted to a value taking the form, where m is an integer, Method.

[0355] 118. The method according to item 117, wherein the motion vector difference (vx, vy) is given by the equation bdofOffset = Round((vx × (gradientHL1[x + 1][y + 1] - gradientHL0[x + 1][y + 1])) >> 1) + Round((vy × (gradientVL1[x + 1][y + 1] - gradientVL0[x + 1][y + 1])) >> 1) and has values corresponding to vx and vy used in BDOF to derive bdofOffset according to Method.

[0356] 119. The method according to item 117, wherein the motion vector difference (vx, vy) is given by the equation ΔI(posX, posY) = dMvH[posX][posY] × gradientH[posX][posY] + dMvV[posX][posY] × gradientV[posX][posY] and has values corresponding to dMvH[posX][posY] and dMvV[posX][posY] used in PROF to derive ΔI according to Method.

[0357] 120. The method according to item 117, wherein the rule uses different ranges for the x - component and the y - component depending on the type of the refinement operation, The method according to claim 117.

[0358] 121. A video processing method, selecting a search area used to derive or refine motion information during a decoder - side motion derivation operation or a decoder - side motion refinement operation according to a rule for conversion between a current video block of a video and a coded representation of the video; executing the conversion based on the derived or refined motion information; A method having the above.

[0359] The method according to item 121, wherein the rule defines selecting a non-square motion vector difference (MVD) region; A method.

[0360] The method according to item 122, wherein the non-square MVD region corresponds to a 5×5 rhombus region, a 5×5 octagon region, or a 3×3 rhombus region; A method.

[0361] The method according to item 121, wherein the search region has different search ranges in the horizontal and vertical directions; A method.

[0362] The method according to item 121, wherein the rule defines selecting the search region based on the dimensions and / or shape of the current video block, wherein the dimensions correspond to at least one of the height (H) and width (W) of the current video block; A method.

[0363] The method according to item 125, wherein the rule defines the horizontal MVD search range to be smaller than the vertical MVD search range or to search a rhombus MVD space according to the dimensions of the current video block; A method.

[0364] The method according to item 121, wherein the rule defines selecting the search region based on the motion information of the current video block before refinement; The method according to claim 121.

[0365] 128. The method according to clause 127, wherein the rule determines a horizontal MVD search range to be smaller than a vertical MVD search range or to search a rhombus MVD space according to the motion information including sumAbsHorMv and / or sumAbsVerMv, sumAbsHorMv = abs(MV0[0]) + abs(MV1[0]), sumAbsVerMv = abs(MV0[0]) + abs(MV1[1]), and abs(MVx) refers to the absolute value of MVx, Method.

[0366] 129. The method according to clause 121, wherein the rule determines to select a horizontal search range and / or a vertical search range based on a reference picture. Method.

[0367] 130. The method according to clause 121, wherein the rule determines to select a horizontal search range and / or a vertical search range based on the magnitude of a motion vector. Method.

[0368] 131. A video processing method, comprising: applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with a current video block of the video by using a search area used in a refinement including an integer position of a best match for conversion between the current video block of the video and a coded representation of the video; executing the conversion by using the refined motion vector difference; and applying the DMVR operation includes deriving a sub-pel motion vector difference (MVD) according to a rule. Method.

[0369] 132. The method according to clause 131, wherein The rule determines to derive a horizontal sub-pel MVD when both the positions to the left and right of the integer position of the best match are included in the search area. Method.

[0370] 133. The method according to clause 131, The rule determines to derive a vertical sub-pel MVD when both the positions above and below the integer position of the best match are included in the search area. Method.

[0371] 134. The method according to clause 131, The rule determines to derive the sub-pel MVD for a specific position. Method.

[0372] 135. A video processing method, For a current video block of a video unit of a video, applying a decoder-side motion vector refinement (DMVR) operation to refine a motion vector difference associated with the current video block; Executing a conversion between the current video block and the coded representation of the video using the refined motion vector difference; And having, Applying the DMVR operation includes determining whether to allow or not allow sub-pel motion vector difference (MVD) derivation according to the use of bidirectional optical flow (BDOF) for the video unit. Method.

[0373] 136. The method according to clause 135, The video unit corresponds to a picture, slice, tile group, tile, coding tree unit, or coding unit. Method.

[0374] 137. The method according to clause 135, The sub-pel MVD derivation is not allowed by the use of the BDOF for the video unit, Method.

[0375] 138. The method according to item 135, wherein The sub-pel MVD derivation is allowed by the non-use of the BDOF for the video unit, Method.

[0376] 139. A video processing method, comprising: During the refinement operation using optical flow, deriving a motion vector difference of a first sample of a current video block of the video; Determining a motion vector difference of a second sample based on the derived motion vector difference of the first sample; Performing a conversion between the current video block and the coded representation of the video based on the determination; And a method having the above steps.

[0377] 140. The method according to item 139, wherein The refinement operation is performed using prediction refinement by optical flow (PROF), The motion vector difference of the first sample is derived using an affine model for positions within the current video block and is used to determine the motion vector difference of the second sample for different positions. Method.

[0378] 141. The method according to item 139, wherein The motion vector difference of the first sample located at the upper part is derived using an affine model and is used to determine the motion vector difference of the second sample located at the lower part. The current video block has a width (W) and a height (H). The upper part and the lower part have a size of W×H / 2. Method.

[0379] 142. The method according to item 139, wherein the motion vector difference of the first sample located at the lower part is derived using an affine model and is used to determine the motion vector difference of the second sample located at the upper part, the current video block has a width (W) and a height (H), the upper part and the lower part have a size of W×H / 2, Method.

[0380] 143. The method according to item 139, wherein the motion vector difference of the first sample located at the left part is derived using an affine model and is used to determine the motion vector difference of the second sample located at the right part, the current video block has a width (W) and a height (H), the left part and the right lower part have a size of (W / 2)×H, Method.

[0381] 144. The method according to item 139, wherein the motion vector difference of the first sample located at the right part is derived using an affine model and is used to determine the motion vector difference of the second sample located at the left part, the current video block has a width (W) and a height (H), the left part and the right lower part have a size of (W / 2)×H, Method.

[0382] 145. The method according to item 139, wherein the motion vector difference of the first sample is derived using an affine model and is rounded to a predefined accuracy and / or clipped to a predefined range before being used to determine the motion vector difference of the second sample, Method.

[0383] 146. The method according to item 139, wherein For the first sample or the second sample at the position (x, y), MVDh and MVDv are the horizontal and vertical motion vector differences respectively, and the equations MVDh(x,y)=-MVDh(W-1-x,H-1-y), and MVDv(x,y)=-MVDv(W-1-x,H-1-y) are obtained by using where W and H are the width and height of the current video block, method.

[0384] 147. A video processing method, deriving a prediction refinement sample by applying to a current video block of a refined motion video using a bidirectional optical flow (BDOF) operation; determining the applicability of a clipping operation to clip the derived prediction refinement sample to a predetermined range [-M, N], assuming M and N are integers, according to a rule; performing a conversion between the current video block and the coded representation of the video; and a method having

[0385] 148. The method according to clause 147, wherein the rule determines to apply the clipping operation; method.

[0386] 149. The method according to clause 147, where N is equal to the value of M - 1; method.

[0387] 150. The method according to clause 147, wherein the predetermined range depends on the bit depth of the color components of the current video block; method.

[0388] 151. The method according to clause 150, where M is 2 Kequal to K = Max(K1, BitDepth + K2), and BitDepth is the bit depth of the color component of the current video block, K1 and K2 are integers, Method.

[0389] The method according to clause 150, wherein M is 2 K equal to K = Min(K1, BitDepth + K2), and BitDepth is the bit depth of the color component of the current video block, K1 and K2 are integers, Method.

[0390] The method according to clause 147, wherein the predetermined range is applied to other video blocks using Predictive Refinement Optical Flow (PROF), Method.

[0391] The method according to clause 147, wherein the rule determines not to apply the clipping operation, Method.

[0392] A video processing method, comprising: determining a coding group size for a current video block of a video, the current video block including a first coding group and a second coding group that are coded using different residual coding modes such that the first coding group and the second coding group are aligned according to a rule; performing a conversion between the current video block and a coded representation of the video based on the determination; and a method having the above steps.

[0393] The method according to clause 155, wherein The first coding group is coded using a transform skip mode in which the transform is bypassed or an identity transform is applied, wherein the rule determines that the size of the first coding group is determined based on whether the current video block includes more samples than 2×M and / or based on a residual block size, Method.

[0394] 157. The method according to clause 155, wherein the first coding group is coded without using a transform skip mode in which the transform is bypassed or an identity transform is applied, wherein the rule determines that the size of the first coding group is determined based on whether the width (W) or height (H) of the current video block is equal to K, where K is an integer, Method.

[0395] 158. The method according to clause 155, wherein the first coding group and the second coding group are each coded using a transform skip mode and a regular residual coding mode, wherein the rule determines that the first coding group and the second coding group have sizes of 2×2, 2×8, 2×4, 8×2, or 4×2, Method.

[0396] 159. The method according to clause 158, wherein the size is used for a residual block of N×2 or 2×N, where N is an integer, Method.

[0397] 160. A video processing method, For the conversion between the current video block of the video and the coded representation of the video, based on the coded information and / or decoded information associated with the current video block, determining the applicability of a predictive refinement optical flow (PROF) tool in which motion information is refined using optical flow; executing the conversion based on the determination; A method comprising the steps of.

[0398] 161. The method according to clause 160, wherein the current video block is coded using an affine mode. A method.

[0399] 162. The method according to clause 161, wherein the determining step determines not to apply the PROF tool because the reference picture used in the conversion has dimensions different from those of the current picture including the current video block. A method.

[0400] 163. The method according to clause 161, wherein the determining step determines to apply the PROF tool because the conversion uses dual-predictive prediction and both reference pictures have the same dimensions. A method.

[0401] 164. The method according to clause 161, wherein the applicability of the PROF tool is determined according to the resolution ratio between the reference picture and the current picture including the current video block. A method.

[0402] 165. The method according to clause 160, wherein the determining step determines not to apply the PROF tool according to a rule defining specific conditions. A method.

[0403] 166. The method according to item 165, wherein the specific conditions are i) enabling the generalized double prediction; ii) enabling the weight prediction, or iii) applying a half-pixel interpolation filter including Method.

[0404] 167. The method according to any one of items 1 to 166, wherein the information regarding the refinement operation is signaled at the video unit level including a sequence, a picture, a slice, a tile, a brick, or other video regions. Method.

[0405] 168. The method according to any one of items 1 to 166, wherein the refinement operation is performed according to a rule based on the coded information of the current video block depending on whether bidirectional optical flow (BDOF) or predictive refinement optical flow (PROF) is applied to the current video block. Method.

[0406] 169. The method according to any one of items 1 to 168, wherein the transformation includes generating the coded representation from the video or generating the video from the coded representation. Method.

[0407] 170. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions, wherein when executed by the processor, the instructions cause the processor to perform the method according to any one of items 1 to 169. Apparatus.

[0408] A computer program product stored on a non-transitory computer-readable medium, comprising: program code for performing the method according to any one of clauses 1 to 169.

[0409] From the above, it is recognized that the specific embodiments of the currently disclosed technology have been described herein for purposes of illustration, but various changes may be made without departing from the scope of the invention. Accordingly, the currently disclosed technology is not limited except as by the appended claims.

[0410] The implementation of the subject matter and functional operations described herein can be implemented in various systems and digital electronic circuits, or in computer software, firmware, or hardware, or combinations of one or more of them, including the structures disclosed herein and their structural equivalents. The implementation of the subject matter described herein can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a tangible and non-transitory computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage carrier, a memory device, a composition that provides a machine-readable propagated signal, or combinations of one or more of them. The terms "data processing unit" or "data processing apparatus" include, by way of example, all apparatus, devices, and machines that process data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or combinations of one or more of them.

[0411] A computer program (also known as a program, software, software application, script, or code) can be described in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a single file dedicated to the program in question, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code), in portions of files that hold other programs or data (e.g., one or more scripts stored in a markup language document). A computer program can be deployed to be executed on one computer, or on a single location, or on multiple computers distributed across multiple locations and interconnected by a communication network.

[0412] The processes and logical flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data to produce output. The processes and logical flows can also be executed by, and the apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit).

[0413] Processors suitable for the execution of a computer program include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. In general, a processor will receive instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical disks, or optical disks, or will be operatively coupled for receiving data from or transferring data to or both from such one or more mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices. The processor and the memory may be enhanced by, or incorporated in, dedicated logic circuitry.

[0414] This specification, together with the drawings, is intended to be regarded as merely illustrative, where illustrative means by way of example. As used herein, the use of "or", "or", "or" is intended to include "and / or", "and / or" unless otherwise explicitly stated in the context.

[0415] This specification includes a number of details, but they are to be construed as descriptions of features that may be specific to particular embodiments of a particular invention rather than as limitations on the scope of any invention or of what may be claimed. The particular features described herein in connection with separate embodiments may be implemented in combination with a single embodiment. Conversely, the various features described in connection with a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Further, features may be described and even initially claimed as operating in a particular combination, but in some cases, one or more features from the claimed combination may be excisable from that combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0416] Similarly, operations are represented in the drawings in a particular order, but this should not be understood as requiring that such operations be performed in that particular order or in a sequential order, or that all of the operations shown be performed, to achieve the desired result. Further, the separation of various system components in the embodiments described herein should not be understood as requiring such separation in all embodiments.

[0417] Only a few implementations and examples have been described, and other implementations, enhancements, and variations may be made based on what has been described and illustrated herein.

Claims

1. 1. A method for processing video data, comprising the steps of: determining, for a first conversion between a first bi-predicted video block of a video and a bitstream of the video, that motion information of the first bi-predicted video block is refined using a first optical flow based method, in which at least one first motion vector refinement is derived to refine prediction samples of a region within the first bi-predicted video block; clipping the at least one first motion vector refinement to a first range; performing the first transformation based on the at least one clipped first motion vector refinement; determining that for a second conversion between a second affine video block of the video and the bitstream, motion information of the second affine video block is refined using a second optical flow based method, in which at least one second motion vector refinement is derived to refine prediction samples of an area within the second affine video block; clipping the at least one second motion vector refinement to a second range; performing the second transformation based on the at least one second motion vector refinement clipped; having the first range is different from the second range; A method, wherein the first optical flow based method is a bidirectional optical flow tool for bi-predicted video blocks and the second optical flow based method is a prediction refinement optical flow tool for affine video blocks.

2. the first range is [-N0, M0] and the second range is [-N1, M1], where N0, M0, N1, and M1 are integers; The method of claim 1.

3. [-N0,M0] is [-15,15]. The method of claim 2.

4. [-N1, M1] is [-31, 31]. The method of claim 2.

5. N0 and M0 are 2 K0 and N1 and M1 have the same value not equal to 2 K1 and K0 and K1 are integers. The method of claim 2.

6. N0 and M0 are 2 K0 -1, and N1 and M1 have the same value equal to 2 K1 -1, and K0 and K1 are integers. The method of claim 2.

7. the region within the first bi-predicted video block is the entirety of the first bi-predicted video block or a sub-block within the first bi-predicted video block; the region within the second affine video block is the entirety of the second affine video block or a sub-block within the second affine video block. The method of claim 1.

8. the first motion vector refinement having a horizontal component and a vertical component; Clipping the first motion vector refinement comprises clipping a horizontal component and / or a vertical component of the first motion vector refinement to the first range; the second motion vector refinement having a horizontal component and a vertical component; and clipping the second motion vector refinement comprises clipping a horizontal component and / or a vertical component of the second motion vector refinement to the second range. The method of claim 1.

9. the first transform comprises encoding the first bi-predicted video block into the bitstream, and the second transform comprises encoding the second affine video block into the bitstream.

9. The method according to any one of claims 1 to 8.

10. the first transform comprises decoding the first bi-predicted video block from the bitstream, and the second transform comprises decoding the second affine video block from the bitstream.

9. The method according to any one of claims 1 to 8.

11. 1. An apparatus for processing video data, comprising: a processor and a non-transitory memory having instructions; The instructions, when executed by the processor, cause the processor to: determining, for a first conversion between a first bi-predicted video block of a video and a bitstream of the video, that motion information of the first bi-predicted video block is refined using a first optical flow based method, in which at least one first motion vector refinement is derived to refine prediction samples of a region within the first bi-predicted video block; clipping the at least one first motion vector refinement to a first range; performing said first transformation based on said at least one clipped first motion vector refinement; determining, for a second conversion between a second affine video block of the video and the bitstream, that motion information of the second affine video block is refined using a second optical flow based method, in which at least one second motion vector refinement is derived to refine prediction samples of a region within the second affine video block; clipping the at least one second motion vector refinement to a second range; performing the second transformation based on the clipped refinement of the at least one second motion vector; the first range is different from the second range; the first optical flow based method is a bi-directional optical flow tool for bi-predicted video blocks and the second optical flow based method is a prediction refinement optical flow tool for affine video blocks. Device.

12. A non-transitory computer-readable storage medium storing instructions, comprising: The instructions cause a processor to: determining, for a first conversion between a first bi-predicted video block of a video and a bitstream of the video, that motion information of the first bi-predicted video block is refined using a first optical flow based method, in which at least one first motion vector refinement is derived to refine prediction samples of a region within the first bi-predicted video block; clipping the at least one first motion vector refinement to a first range; performing said first transformation based on said at least one clipped first motion vector refinement; determining, for a second conversion between a second affine video block of the video and the bitstream, that motion information of the second affine video block is refined using a second optical flow based method, in which at least one second motion vector refinement is derived to refine prediction samples of a region within the second affine video block; clipping the at least one second motion vector refinement to a second range; performing the second transformation based on the clipped refinement of the at least one second motion vector; the first range is different from the second range; the first optical flow based method is a bi-directional optical flow tool for bi-predicted video blocks and the second optical flow based method is a prediction refinement optical flow tool for affine video blocks. A non-transitory computer-readable storage medium.

13. A method for storing a bitstream of video, comprising the steps of: determining, for a first bi-predicted video block of the video, that motion information of the first bi-predicted video block is refined using a first optical flow based method, in which at least one first motion vector refinement is derived to refine prediction samples of a region within the first bi-predicted video block; clipping the at least one first motion vector refinement to a first range; generating the bitstream based on the at least one first motion vector refinement clipped; determining, for a second affine video block of the video, that motion information of the second affine video block is refined using a second optical flow based method, in which at least one second motion vector refinement is derived to refine prediction samples of an area within the second affine video block; clipping the at least one second motion vector refinement to a second range; generating the bitstream based on the at least one second motion vector refinement clipped; having the first range is different from the second range; the first optical flow based method is a bi-directional optical flow tool for bi-predicted video blocks and the second optical flow based method is a prediction refinement optical flow tool for affine video blocks. method.

Citation Information

Patent Citations

  • JPP7431253B