Enable DMVR based on information in the picture header

By introducing DMVR and BIO technologies in video encoding and decoding, motion vector management is optimized, the challenges of compression ratio and complexity in existing technologies are solved, and more efficient video encoding and decoding effects are achieved.

CN113545085BActive Publication Date: 2025-09-16DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080018543.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-04
Filing Date
2020-03-03
Publication Date
2025-09-16
Estimated Expiration
2040-03-03

Smart Images

  • Figure CN113545085B_ABST
    Figure CN113545085B_ABST
Patent Text Reader

Abstract

A method of video processing is disclosed, comprising: determining, based on signaled information, whether and / or how to apply decoder-side motion vector refinement (DMVR) for a conversion between a first block of video and a bitstream representation of the first block of video; and performing the conversion based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is intended to claim priority to and the benefit of International Patent Application No. PCT / CN2019 / 076788, filed on March 3, 2019, and International Patent Application No. PCT / CN2019 / 076860, filed on March 4, 2019, in a timely manner, under applicable patent law and / or rules pursuant to the Paris Convention. The entire disclosures of International Patent Application Nos. PCT / CN2019 / 076788 and PCT / CN2019 / 076860 are incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to video encoding and decoding technology, equipment and systems. Background Art

[0004] Currently, efforts are underway to improve the performance of current video codec technologies to provide better compression ratios or to provide video encoding and decoding schemes that allow for lower complexity or parallel implementation. Several new video codec tools have recently been proposed by industry experts and are currently being tested to determine their effectiveness. Summary of the Invention

[0005] Devices, systems, and methods are described that relate to digital video codecs, and in particular, to the management of motion vectors. The methods described can be applied to existing video codec standards (e.g., High Efficiency Video Codec (HEVC) or Universal Video Codec) and future video codec standards or codecs.

[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes performing a conversion between a current video block and a bitstream representation of the current video block, wherein use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that the decoder motion information is enabled during the conversion and a second value of the flag indicates that the decoder motion information is disabled during the conversion.

[0007] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes performing a conversion between a current video block and a bitstream representation of the current video block, wherein use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that the decoder motion information is enabled during the conversion and a second value of the flag indicates that the decoder motion information is disabled during the conversion, and in response to determining that an initial motion vector for the video block has sub-pixel precision, skipping a check associated with a temporary motion vector derived from the initial motion vector and a candidate motion vector difference, wherein the temporary motion vector has integer pixel precision or sub-pixel precision.

[0008] In another representative aspect, the disclosed technology can be used to provide another method for video processing, the method comprising: determining, based on signaled information, whether and / or how to apply decoder-side motion vector refinement (DMVR) for a conversion between a first block of video and a bitstream representation of the first block of video; and performing the conversion based on the determination.

[0009] In another representative aspect, the disclosed technology can be used to provide another method for video processing, the method comprising: determining whether and / or how to apply bidirectional optical flow (BIO) for conversion between a first block of video and a bitstream representation of the first block of video based on signaled information; and performing the conversion based on the determination.

[0010] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes determining whether a decoder-side motion vector refinement (DMVR) process is enabled or disabled for a conversion between a first block of video and a bitstream representation of the first block of video based on at least one of: one or more initial motion vectors associated with the first block and one or more reference pictures associated with the first block, the initial motion vectors comprising motion vectors before applying the DMVR process; and performing the conversion based on the determination.

[0011] In another representative aspect, the disclosed technology can be used to provide another method for video processing. The method includes determining whether a bidirectional optical flow (BIO) process is enabled or disabled for converting between a first block of video and a bitstream representation of the first block of video based on at least one of: one or more initial motion vectors associated with the first block and one or more reference pictures associated with the first block, the initial motion vectors including motion vectors before applying the BIO process and / or decoded motion vectors before applying a decoder-side motion vector refinement (DMVR) process; and performing the conversion based on the determination.

[0012] In another representative aspect, the disclosed technology can be used to provide another method for video processing, the method comprising: determining whether a bidirectional optical flow (BIO) process is enabled or disabled for conversion between a first block of video and a bitstream representation of the first block of video based on one or more motion vector differences between an initial motion vector associated with the first block of video and one or more refined motion vectors, the initial motion vector comprising a motion vector before applying the BIO process and / or applying a DMVR process, the refined motion vector comprising a motion vector after applying the DMVR process; and performing the conversion based on the determination.

[0013] Furthermore, in a representative aspect, an apparatus in a video system is disclosed, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any one or more of the disclosed methods.

[0014] Furthermore, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product comprising program code for performing any one or more of the disclosed methods.

[0015] The above and other aspects and features of the disclosed technology are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 An example of constructing a Merge candidate list is shown.

[0017] Figure 2 Examples of locations of spatial candidates are shown.

[0018] Figure 3 An example of candidate pairs for which redundancy checking of spatial merge candidates is performed is shown.

[0019] Figure 4A and Figure 4B An example of the position of the second prediction unit (PU) based on the size and shape of the current block is shown.

[0020] Figure 5 An example of motion vector scaling for temporal merge candidates is shown.

[0021] Figure 6 An example of candidate positions of time-domain Merge candidates is shown.

[0022] Figure 7 An example of generating a combined bi-predictive Merge candidate is shown.

[0023] Figure 8An example of constructing motion vector prediction candidates is shown.

[0024] Figure 9 An example of motion vector scaling for spatial motion vector candidates is shown.

[0025] Figure 10 An example of neighboring samples used to derive local illumination compensation parameters is shown.

[0026] Figure 11A and Figure 11B Graphs related to a 4-parameter affine model and a 6-parameter affine model are shown respectively.

[0027] Figure 12 An example of the affine motion vector field for each sub-block is shown.

[0028] Figure 13A and Figure 13B Examples of a 4-parameter affine model and a 6-parameter affine model are shown respectively.

[0029] Figure 14 An example of motion vector prediction in affine inter mode for inherited affine candidates is shown.

[0030] Figure 15 An example of motion vector prediction for affine inter mode of constructed affine candidates is shown.

[0031] Figure 16A and Figure 16B A diagram related to the affine Merge mode is shown.

[0032] Figure 17 An example of candidate positions for the affine merge mode is shown.

[0033] Figure 18 An example of a Merge with Motion Vector Differences (MMVD) mode search process is shown.

[0034] Figure 19 Examples of MMVD search points are shown.

[0035] Figure 20 An example of Decoder-side Motion Video Refinement (DMVR) in JEM7 is shown.

[0036] Figure 21 An example of motion vector difference (MVD) related to DMVR is shown.

[0037] Figure 22An example illustrating the checking of a motion vector is shown.

[0038] Figure 23 An example of reference samples in DMVR is shown.

[0039] Figure 24 is a block diagram of an example of a hardware platform for implementing the visual media decoding or visual media encoding techniques described in this document.

[0040] Figure 25 A flow chart illustrating an example method for video encoding and decoding is shown.

[0041] Figure 26 A flow chart illustrating an example method for video encoding and decoding is shown.

[0042] Figure 27 A flow chart illustrating an example method for video encoding and decoding is shown.

[0043] Figure 28 A flow chart illustrating an example method for video encoding and decoding is shown.

[0044] Figure 29 A flow chart illustrating an example method for video encoding and decoding is shown.

[0045] Figure 30 A flow chart illustrating an example method for video encoding and decoding is shown. DETAILED DESCRIPTION

[0046] 1. Video Codec in HEVC / H.265

[0047] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and 264 / MPEG-4 Advanced Video Coding (AVC) standards, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software named the Joint Exploration Model (JEM). In April 2018, a Joint Video Expert Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, targeting a 50% bitrate reduction compared to HEVC.

[0048] The latest version of the VVC draft, General Video Codec (Draft 4), can be found at: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 13_Marrakesh / wg11 / JVET-M1001-v5.zip. The latest reference software for VVC, called VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tree / VTM-4.0.

[0049] 2.1. Inter-frame prediction in HEVC / H.265

[0050] Each inter-prediction PU has motion parameters for one or two reference picture lists. Motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. Motion vectors can be explicitly encoded and decoded as deltas relative to the predicted value.

[0051] When a CU is encoded or decoded in skip mode, one PU is associated with the CU and there are no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified, whereby the motion parameters of the current PU are obtained from neighboring PUs including spatial and temporal candidates. Merge mode can be applied to any inter-predicted PU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector prediction value), the corresponding reference picture index of each reference picture list, and the reference picture list usage are explicitly signaled per PU. Such a mode is named Advanced Motion Vector Prediction (AMVP) in this disclosure.

[0052] When signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a block of samples. This is called "unidirectional prediction." Unidirectional prediction is applicable to both P slices and B slices.

[0053] When signaling indicates that two reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bi-prediction." Bi-prediction is only applicable to B slices.

[0054] The following text provides details about the inter prediction modes specified in HEVC. The description will start with the Merge mode.

[0055] 2.1.1. Reference Image List

[0056] In HEVC, the term inter prediction is used to refer to predictions derived from data elements (e.g., sample values ​​or motion vectors) of reference pictures other than the currently decoded picture. As in H.264 / AVC, a picture can be predicted from multiple reference pictures. Reference pictures used for inter prediction are organized in one or more reference picture lists. A reference index identifies which reference picture in the list should be used to create the prediction signal.

[0057] A single reference picture list (list 0) is used for P slices, and two reference picture lists (list 0 and list 1) are used for B slices. It should be noted that the reference pictures included in lists 0 / 1 can be based on past and future pictures in terms of capture / display order. In addition, the current picture can be on the list of reference pictures in HEVC version 4.

[0058] 2.1.2.Merge mode

[0059] 2.1.2.1. Merge Mode Candidate Derivation

[0060] When predicting a PU using Merge mode, the index pointing to the entry in the Merge candidate list is parsed from the bitstream and used to retrieve the motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:

[0061] Step 1: Initial candidate derivation

[0062] Step 1.1: Spatial Candidate Derivation

[0063] Step 1.2: Redundancy check of spatial candidates

[0064] Step 1.3: Time Domain Candidate Derivation

[0065] Step 2: Additional candidate insertions

[0066] Step 2.1: Create bidirectional prediction candidates

[0067] Step 2.2: Insert zero motion candidates

[0068] exist Figure 1 These steps are also schematically depicted in . For spatial domain Merge candidate derivation, a maximum of four Merge candidates are selected from candidates located at five different positions. For temporal domain Merge candidate derivation, a maximum of one Merge candidate is selected from two candidates. Since the number of candidates for each PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using truncated unary (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of the 2N×2N prediction unit.

[0069] Hereinafter, operations associated with the aforementioned steps are described in detail.

[0070] 2.1.2.2. Spatial Candidate Derivation

[0071] In the derivation of spatial Merge candidates, from Figure 2Up to four Merge candidates are selected from the candidates at the positions depicted in . The order of derivation is A1, B1, B0, A0 and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is not available (for example, because it belongs to another strip or slice) or is intra-coded. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs at Figure 3 The pairs linked by arrows in , and the corresponding candidate is added to the list only if the candidates used for redundancy check do not have the same motion information. Another source of duplicate motion information is a "second PU" associated with a partition different from 2N×2N. As an example, Figure 4 depicts the second PU for the cases of N×2N and 2N×N. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for having only one PU in the codec unit. Similarly, position B1 is not considered when the current PU is partitioned into 2N×N.

[0072] 2.1.2.3. Time Domain Candidate Derivation

[0073] In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on the collocated PU belonging to the picture with the smallest POC difference with the current picture in the given reference picture list. The reference picture list to be used for the derivation of the collocated PU is explicitly signaled in the slice header. Figure 5 The dotted line in the figure shows the scaled motion vector of the temporal merge candidate, which is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors are obtained, one for reference picture list 0 and the other for reference picture list 1, and are combined to form a bidirectional prediction merge candidate.

[0074] like Figure 6As depicted, in a collocated PU (Y) belonging to a reference frame, the position of the temporal candidate is selected between candidates C0 and C1. If the PU at position C0 is not available, is intra-coded, or is outside the current codec tree unit (CTU, also known as LCU, the largest codec unit) row, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0075] 2.1.2.4. Additional candidate insertion

[0076] In addition to the spatiotemporal Merge candidates, there are two additional types of Merge candidates: combined bi-predictive Merge candidates and zero Merge candidates. Combined bi-predictive Merge candidates are generated by utilizing spatiotemporal Merge candidates. Combined bi-predictive Merge candidates are used only for B slices. Combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another. If the two tuples provide different motion hypotheses, they will form a new bi-predictive candidate. As an example, Figure 7 Depicted is the case where two candidates with mvL0 and refIdxL0 or mvL1 and refIdxL1 in the original list (on the left) are used to create combined bi-predictive Merge candidates that are added to the final list (on the right). There are many rules on combining that are considered to generate these additional Merge candidates.

[0077] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.

[0078] 2.1.3.AMVP

[0079] AMVP exploits the spatiotemporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the left, upper temporal neighboring PU positions, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best prediction value from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is truncated unary coding. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, details on the derivation process of motion vector prediction candidates are provided.

[0080] 2.1.3.1. Derivation of AMVP Candidates

[0081] Figure 8 The derivation process of motion vector prediction candidates is summarized.

[0082] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, based on the location of Figure 2 The motion vectors of each PU at the five different positions depicted are used to finally derive two motion vector candidates.

[0083] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different collocated positions. After generating the first list of spatiotemporal candidates, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than two, motion vector candidates with a reference picture index greater than 1 within the list are removed from the associated reference picture list. If the number of spatiotemporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.

[0084] 2.1.3.2. Spatial Motion Vector Candidates

[0085] In the derivation of spatial motion vector candidates, a maximum of two candidates are considered among five potential candidates, which are selected from the candidate located at Figure 2 The positions depicted are derived from the PUs, which are the same as the positions of the motion merge. The derivation order on the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order above the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not use spatial scaling and two use spatial scaling. These four different cases are summarized as follows:

[0086] No airspace scaling

[0087] –(1) Same reference picture list and same reference picture index (same POC)

[0088] –(2) Different reference picture lists but same reference pictures (same POC)

[0089] Airspace scaling

[0090] –(3) Same reference picture list but different reference pictures (different POC)

[0091] –(4) Different reference picture lists and different reference pictures (different POC)

[0092] First check for no spatial scaling, then check for spatial scaling. Regardless of the reference picture list, spatial scaling is considered when the POC between the reference picture of the neighboring PU and the reference picture of the current PU is different. If all PUs of the left candidate are unavailable or intra-coded, scaling of the upper motion vector is allowed to facilitate the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling of the upper motion vector is not allowed.

[0093] like Figure 9 As depicted, in the spatial scaling process, the motion vectors of neighboring PUs are scaled in a similar manner to temporal scaling. The main difference is that the reference picture list and the index of the current PU are given as input; the actual scaling process is the same as the scaling process of temporal scaling.

[0094] 2.1.3.3. Temporal Motion Vector Candidates

[0095] Except for the reference picture index derivation, all the processes for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (see Figure 6 ). The reference picture index is signaled to the decoder.

[0096] 2.2. Local Illumination Compensation in JEM

[0097] Local Illumination Compensation (LIC) is based on a linear model of illumination variations using a scaling factor a and an offset b, and is adaptively enabled or disabled for each inter-mode coded codec unit (CU).

[0098] When LIC is applied to a CU, the least square method is used to derive parameters a and b by using the neighboring samples of the current CU and its corresponding reference samples. More specifically, Figure 10 As shown, the neighboring samples and corresponding samples (identified by the motion information of the current CU or sub-CU) of the subsampling (2:1 subsampling) of the CU in the reference picture are used.

[0099] 2.2.1 Derivation of prediction blocks

[0100] LIC parameters are derived and applied separately for each prediction direction. For each prediction direction, the decoded motion information is used to generate the first prediction block, and then a temporary prediction block is obtained by applying the LIC model. The two temporary prediction blocks are then used to derive the final prediction block.

[0101] When a CU is encoded or decoded in Merge mode, the LIC flag is copied from the neighboring blocks in a manner similar to the motion information copying in Merge mode; otherwise, the LIC flag is signaled to the CU to indicate whether LIC is applicable.

[0102] When LIC is enabled for a picture, an additional CU-level RD check is required to determine whether LIC is applicable to the CU. When LIC is enabled for a CU, the Mean-Removed Sum of Absolute Difference (MR-SAD) and the Mean-Removed Sum of Absolute Hadamard-Transformed Difference (MR-SATD) are used for integer-pixel motion refinement and fractional-pixel motion refinement, respectively, instead of SAD and SATD.

[0103] To reduce the encoding complexity, the following encoding scheme is applied in JEM.

[0104] When there is no significant illumination change between the current picture and its reference pictures, LIC is disabled for the entire picture. To identify this situation, the encoder computes the histogram of the current picture and each of its reference pictures. If the histogram difference between the current picture and each of its reference pictures is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.

[0105] 2.3. Inter-frame prediction method in VVC

[0106] There are several new codec tools for inter-frame prediction improvements, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling MVD, affine prediction mode, Triangular Prediction Mode (TPM), Advanced TMVP (ATMVP, also known as SbTMVP), Generalized Bi-Prediction (GBI), and Bidirectional Optical Flow (BIO).

[0107] 2.3.1. Codec Block Structure in VVC

[0108] In VVC, a quadtree / binary tree / multi-type tree (QT / BT / TT) structure is used to divide the image into square or rectangular blocks.

[0109] In addition to QT / BT / TT, separate trees (also known as dual codec trees) are also used in VVC for I slices / slices. In the case of separate trees, the codec block structure is signaled separately for luma and chroma components.

[0110] 2.3.2. Adaptive Motion Vector Difference Resolution

[0111] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. In VVC, local adaptive motion vector resolution (AMVR) is introduced. In VVC, MVD can be coded or decoded in units of quarter luma samples, integer luma samples, or four luma samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.

[0112] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or quad luma sample MV precision is used.

[0113] When the first MVD resolution flag of the CU is zero, or the CU is not coded (meaning that all MVDs in the CU are zero), the CU uses a quarter-luminance sample MV resolution. When the CU uses integer luminance sample MV precision or four-luminance sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.

[0114] 2.3.3. Affine Motion Compensated Prediction

[0115] In HEVC, only the translation motion model is applied to motion compensation prediction (MCP). Although in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, a 4-parameter affine model and a 6-parameter affine model are used to apply simplified affine transformation motion compensation prediction. As shown in Figure 11, the affine motion field of a block is described by two control point motion vectors (CPMVs) of the 4-parameter affine model and three CPMVs of the 6-parameter affine model.

[0116] The motion vector field (MVF) of the block is described by the following equations using the 4-parameter affine model in equation (1) (where the 4 parameters are defined as variables a, b, e, and f) and the 6-parameter affine model in equation (2) (where the 4 parameters are defined as variables a, b, c, d, e, and f):

[0117]

[0118]

[0119] Among them (mv h 0,mv h 0) is the motion vector of the upper left control point, and (mv h 1, mv h 1) is the motion vector of the upper right control point, and (mv h 2, mv h 2) is the motion vector of the lower left control point. All three motion vectors are called control point motion vectors (CPMVs). (x, y) represents the coordinates of the representative point relative to the upper left sample point in the current block, and (mv h (x, y), mv v (x, y)) is the motion vector derived for the sample at (x, y). The CP motion vector can be signaled (such as in affine AMVP mode) or derived on the fly (such as in affine Merge mode). w and h are the width and height of the current block. In practice, division is implemented by right shift and rounding operations. In VTM, the representative point is defined as the center position of the sub-block, for example, when the coordinates of the upper left corner of the sub-block relative to the upper left sample in the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (i.e., 4×4 in VTM), the motion vector of the entire sub-block is derived using the representative point.

[0120] In order to further simplify the motion compensation prediction, the sub-block based affine transformation prediction is applied. In order to derive the motion vector of each M×N (in the current VVC, M and N are both set to 4) sub-block, the motion vector of the center sample of each sub-block (such as Figure 12 (as shown) can be calculated according to equations (1) and (2) and rounded to 1 / 16 fractional precision. Then, a 1 / 16 pixel motion compensation interpolation filter can be applied to generate a prediction for each sub-block with the derived motion vector. The affine mode introduces a 1 / 16 pixel interpolation filter.

[0121] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the standard motion vector.

[0122] 2.3.3.1. Signaling of Affine Prediction

[0123] Similar to the translational motion model, due to affine prediction, there are also two modes for signaling side information. They are AFFINE_INTER and AFFINE_MERGE modes.

[0124] 2.3.3.2.AF_INTER mode

[0125] AF_INTER mode can be applied to CUs with width and height greater than 8. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used.

[0126] In this mode, for each reference picture list (list 0 or list 1), an affine AMVP candidate list is constructed with three types of affine motion predictors in the following order, where each candidate includes the estimated CPMV of the current block. The best CPMV found at the encoder side (such as Figure 15 The difference between mv0, mv1, mv2) and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.

[0127] 1) Inherited affine motion prediction value

[0128] The checking order is similar to the checking order of spatial MVP in the HEVC AMVP list. First, the inherited affine motion prediction value on the left is derived from the first block that is affine-coded and has the same reference picture as the current block in {A1, A0}. Second, the inherited affine motion prediction value on the top is derived from the first block that is affine-coded and has the same reference picture as the current block in {B1, B0, B2}. Figure 14 The five blocks A1, A0, B1, B0, B2 are depicted in FIG.

[0129] Once a neighboring block is found to be coded in affine mode, the CPMV of the codec covering the neighboring block is used to derive the CPMV prediction value of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the inherited affine MV prediction value on the left will be derived from A0. In this case, the CPMV of the CU covering A0 (as in Figure 16B Zhongyou The upper left CPMV represented by The upper right CPMV represented by is used to derive the estimated CPMV of the current block, which is given by Indicates the upper left (coordinates (x0, y0)), upper right (coordinates (x1, y1)), and lower right positions (coordinates (x2, y2)) of the current block.

[0130] 2) Constructed affine motion prediction value

[0131] like Figure 15As shown, the constructed affine motion prediction value contains the control point motion vector (CPMV) derived from the adjacent inter-frame codec blocks with the same reference picture. If the current affine motion model is 4-parameter affine, the number of CPMVs is 2, otherwise if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. CPMV in the upper left corner Derived from the MV of the first block in group {A, B, C} that is inter-coded and has the same reference picture as the current block. Derived from the MV of the first block in group {D, E} that is inter-coded and has the same reference picture as the current block. Derived from the MV at the first block in group {F, G} that is inter-coded and has the same reference picture as the current block.

[0132] – If the current affine motion model is 4-parameter affine, only if and When both are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMV of the upper left (coordinates (x0, y0)) and upper right (coordinates (x1, y1)) positions of the current block.

[0133] – If the current affine motion model is 6-parameter affine, only if and When all are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMVs for the top left (coordinates (x0, y0)), top right (coordinates (x1, y1)), and bottom right (coordinates (x2, y2)) positions of the current block.

[0134] When inserting the constructed affine motion predictors into the candidate list, no pruning process is applied.

[0135] 3) Normal AMVP motion prediction value

[0136] The following applies until the number of affine motion predictors reaches a maximum value.

[0137] 1) If available, by setting all CPMVs equal to To derive the affine motion prediction value.

[0138] 2) If available, by setting all CPMVs equal to To derive the affine motion prediction value.

[0139] 3) If available, by setting all CPMVs equal to To derive the affine motion prediction value.

[0140] 4) If available, derive affine motion prediction values ​​by setting all CPMVs equal to HEVC TMVPs.

[0141] 5) Derive the affine motion prediction value by setting all CPMVs to zero MV.

[0142] Please note that has been derived in the constructed affine motion predictor.

[0143] In AF_INTER mode, when using 4 / 6 parameter affine mode, 2 / 3 control points are used, so 2 / 3 MVDs need to be encoded and decoded for these control points, as shown in Figure 13. The following derivation of MVs is proposed, i.e., predicting mvd1 and mvd2 from mvd0.

[0144]

[0145]

[0146]

[0147] in, mvd i and mv1 are the predicted motion vector, motion vector difference, and motion vector of the upper left pixel (i=0), upper right pixel (i=1), or lower left pixel (i=2), respectively. Figure 13B Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the separate summation of the two components, i.e., newMV=mvA+mvB, and the two components of newMV are set to (xA+xB) and (yA+yB), respectively.

[0148] 2.3.3.3.AF_Merge Mode

[0149] When the CU is applied to the AF_MERGE mode, it obtains the first block encoded and decoded in the affine mode from the valid neighboring reconstructed blocks. And the selection order of the candidate blocks is from left, top, top right, bottom left to top left, such as Figure 16A As shown (indicated by A, B, C, D, E in order). For example, if the adjacent lower left block is Figure 16B If A0 in the block is encoded and decoded in affine mode, the control point (CP) motion vector mv0 of the upper left corner, upper right corner and lower left corner of the adjacent CU / PU containing block A is obtained N 、mv1 Nand mv2 N . And based on mv0 N 、mv1 N and mv2 N Calculate the motion vector mv0 of the upper left corner / upper right / lower left corner of the current CU / PU C 、mv1 C and mv2 C (For 6-parameter affine models only.) Note that in the current VTM, the upper-left subblock (e.g., a 4×4 block in the VTM) stores mv0, and if the current block is affine-encoded, the upper-right subblock stores mv1. If the current block is affine-encoded with a 6-parameter affine model, the lower-left subblock stores mv2; otherwise (with a 4-parameter affine model), LB stores mv2'. The other submodules store the MVs used for the MC.

[0150] The CPMV mv0 of the current CU is derived from the simplified affine motion model in equations (1) and (2). C 、mv1 C and mv2 C After that, the MVF of the current CU is generated. In order to identify whether the current CU is coded or decoded in AF_MERGE mode, when there is at least one neighboring block coded or decoded in affine mode, an affine flag is signaled in the bitstream.

[0151] The affine merge candidate list is constructed using the following steps:

[0152] 1) Insert inherited affine candidates

[0153] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine codec block. A maximum of two inherited affine candidates are derived from the affine motion models of the neighboring blocks and inserted into the candidate list. For the left prediction, the scan order is {A0, A1}; for the top prediction, the scan order is {B0, B1, B2}.

[0154] 2) Insert the constructed affine candidate

[0155] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining neighboring motion information of each control point.

[0156] a) The motion information of the control point is first obtained from Figure 17 The specified spatial and temporal neighbors are derived as shown. CPk (k = 1, 2, 3, 4) represents the kth control point. A0,

[0157] A1, A2, B0, B1, B2 and B3 are the spatial positions used to predict CPk (k=1, 2, 3); T is the temporal position for predicting CP4.

[0158] The coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H), respectively, where W and H are the width and height of the current block.

[0159] The motion information for each control point is obtained according to the following priority order:

[0160] – For CP1, the priority is B2->B3->A2. If B2 is available, use B2. Otherwise, if B2 is available, use B3. If neither B2 nor B3 is available, use A2. If all three candidates are unavailable, motion information for CP1 cannot be obtained.

[0161] – For CP2, the checking priority is B1->B0.

[0162] – For CP3, the checking priority is A1->A0.

[0163] – For CP4, use T.

[0164] b) Secondly, the combination of control points is used to construct affine merge candidates.

[0165] I. Motion information of three control points is required to construct a 6-parameter affine candidate. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} will be converted into a 6-parameter motion model represented by the top left, top right, and bottom left control points.

[0166] II. Motion information of two control points is required to construct a 4-parameter affine candidate. The two control points can be selected from one of two combinations ({CP1, CP2}, {CP1, CP3}). These two combinations will be converted into a 4-parameter motion model represented by the upper left and upper right control points.

[0167] III. The constructed combinations of affine candidates are inserted into the candidate list in the following order: {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}

[0168] i. For each combination, check the reference index of List X for each CP. If they are all the same, then the combination has a valid CPMV for List X. If the combination does not have a valid CPMV for both List 0 and List 1, then the combination is marked as invalid. Otherwise, it is valid, and the CPMV is placed in the sub-block Merge list.

[0169] 3) Fill with zero motion vectors

[0170] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.

[0171] More specifically, for the sub-block Merge candidate list, MV is set to (0, 0) and the prediction direction is set to the 4-parameter Merge candidate of unidirectional prediction (for P slices) and bidirectional prediction (for B slices) from list 0.

[0172] In VTM4, the CPMV of the affine CU is stored in a separate buffer. The stored CPMV is only used to generate the inherited CPMVP for the most recently encoded CU in affine Merge mode and affine AMVP mode. The sub-block MV inherited from the CPMV is used for motion compensation, MV derivation of the Merge / AMVP list of translation MVs, and deblocking. In order to avoid the picture line buffer of the additional CPMV, the affine motion data inheritance of the CU from the upper CTU is handled differently from the inheritance from the normal neighboring CU. If the candidate CU for affine motion data inheritance is in the upper CTU line, the lower left and lower right sub-block MVs in the line buffer are used for affine MVP derivation instead of the CPMV. In this way, the CPMV is only stored in the local buffer. If the candidate CU is 6-parameter affine encoded and decoded, the affine model is downgraded to a 4-parameter model.

[0173] 2.3.4. Merge with Motion Vector Difference (MMVD)

[0174] In JVET-L0054, the Ultimate Motion Vector Expression (UMVE, also known as MMVD) is given. UMVE and a proposed motion vector expression method are used in Skip or Merge mode.

[0175] Figure 18 An example of the ultimate vector expression (UMVE) search process is shown. Figure 19An example of UMVE search points is shown. In VVC, UMVE reuses the same Merge candidates as those included in the regular Merge candidate list. Among these Merge candidates, a basic candidate can be selected and further extended by the proposed motion vector expression method.

[0176] UMVE provides a new method for representing motion vector difference (MVD), in which the MVD is represented by the starting point, motion amplitude and motion direction.

[0177] The proposed technique uses the Merge candidate list as is, but only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE expansion.

[0178] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.

[0179] Table 1. Basic candidate IDX

[0180] Basic Candidate IDX 0 1 2 3 Nth MVP 1st MVP 2nd MVP 3rd MVP 4th MVP

[0181] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.

[0182] The distance index is the motion magnitude information. The distance index indicates the predefined distance from the starting point information. The predefined distances are as follows:

[0183] Table 2. Distance IDX

[0184]

[0185] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown below.

[0186] Table 3. Direction IDX

[0187] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0188] The UMVE flag is signaled immediately after the Skip or Merge flag is transmitted. If the Skip or Merge flag is true, the UMVE flag is parsed. If the UMVE flag is 1, the UMVE syntax is parsed. If not, the AFFINE flag is parsed. If the AFFINE flag is 1, AFFINE mode is used. If not, the Skip / Merge index is parsed as the VTM's Skip / Merge mode.

[0189] No additional line buffer is required for UMVE candidates because software skip / merge candidates are used directly as base candidates. Using the input UMVE index, MV complementation is determined immediately before motion compensation. There's no need to maintain a long line buffer for this purpose.

[0190] Under the current common test conditions, the first or second merge candidate in the merge candidate list can be selected as the base candidate.

[0191] UMVE is also known as Merge with MV Difference (MMVD).

[0192] In addition, the tile_group_fpel_mmvd_enabled_flag is signaled to the decoder in the slice header to indicate whether fractional distances are used. When fractional distances are disabled, all distances in the default table are multiplied by 4, resulting in a distance table of {1, 2, 4, 8, 16, 32, 64, 128} pixels. Because the size of the distance table remains unchanged, the entropy encoding and decoding of the distance index remains unchanged.

[0193] 2.3.5. Decoder-side Motion Vector Refinement (DMVR)

[0194] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MVs) from list 0 and MVs from list 1, respectively, are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.

[0195] 2.3.5.1.DMVR in JEM

[0196] In the JEM design, motion vectors are refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture to obtain refined MV values ​​without the transmission of additional motion information. Figure 20 An example is depicted in . The bilateral template is generated as a weighted combination (i.e., average) of two prediction blocks from the initial MV0 of list 0 and the initial MV1 of list 1, respectively, as Figure 20 As shown. The template matching operation involves calculating the cost metric between the generated template and the sample area in the reference picture (around the initial prediction block). For each of the two reference pictures, the MV that produces the minimum template cost is considered to be the updated MV of the list to replace the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and 8 surrounding MVs that are offset by one luminance sample from the original MV in the horizontal direction, vertical direction, or both directions. Finally, as Figure 20The two new MVs shown (i.e., MV0′ and MV1′) are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Note that when calculating the cost of a prediction block generated by a surrounding MV, the MV rounded to integer pixels is actually used to obtain the prediction block, rather than the actual MV.

[0197] 2.3.5.2.DMVR in VVC

[0198] For DMVR in VVC, we assume a mirrored MVD between list 0 and list 1, such as Figure 21 As shown, bilateral matching is performed to refine the MV, that is, to find the best MVD among several MVD candidates. The MVs of the two reference picture lists are represented by MVL0 (L0X, L0Y) and MVL1 (L1X, L1Y). The MVD of list 0 represented by (MvdX, MvdY) that can minimize the cost function (e.g., SAD) is defined as the best MVD. As for the SAD function, it is defined as the SAD between the reference block of list 0 and the reference block of list 1, where the reference block of list 0 is derived using the motion vector (L0X+MvdX, L0Y+MvdY) in the list 0 reference image, and the reference block of list 1 is derived using the motion vector (L1X-MvdX, L1Y-MvdY) in the list 1 reference image.

[0199] The motion vector refinement process can be iterated twice. In each iteration, up to 6 MVDs (with integer pixel precision) can be checked in two steps, such as Figure 22 As shown. In the first step, the MVDs (0, 0), (-1, 0), (1, 0), (0, -1), (0, 1) are checked. In the second step, one of the MVDs (-1, -1), (-1, 1), (1, -1), or (1, 1) can be selected and further checked. Assume that the function Sad(x, y) returns the SAD value of MVD(x, y). The MVD represented by (MvdX, MvdY) checked in the second step is determined as follows:

[0200] MvdX=-1;

[0201] MvdY=-1;

[0202] If(Sad(1,0) <Sad(-1,0))

[0203] MvdX=1;

[0204] If(Sad(0,1) <Sad(0,-1))

[0205] MvdY=1;

[0206] In the first iteration, the starting point is the signaled MV, and in the second iteration, the starting point is the signaled MV plus the best MVD selected in the first iteration. DMVR is only applicable when one reference picture is the previous picture and the other reference picture is the next picture and both reference pictures have the same picture order count distance to the current picture.

[0207] As mentioned above, DMVR in VVC first performs integer MVD refinement. This is the first step. Afterwards, fractional MVD refinement is conditionally performed to further refine the motion vectors. This is the second step. Whether or not to execute the second step is conditional on whether the MVD after the current iteration is zero MV. If it is zero MV (the vertical and horizontal components of the MV are 0), the second step is executed.

[0208] The details of fractional MVD refinement are given as follows: It should be noted that MVD represents the difference between the initial motion vector and the final motion vector used in motion compensation.

[0209] A parameter error surface is fitted using integer distance positions and the costs evaluated at those positions, which is then used to determine 1 / 16 th Pixel accurate sub-pixel offset.

[0210] The proposed method is summarized as follows:

[0211] 1. The parameter error surface fit is calculated only if the center position is the best cost position in a given iteration.

[0212] 2. The center position cost and the costs at (-1, 0), (0, -1), (1, 0) and (0, 1) from the center are used to fit the 2-D parabolic error surface equation of the shape

[0213] E(x,y)=A(x-x0) 2 +B(y-y0) 2 +C

[0214] Where (x0, y0) corresponds to the least expensive position, and C corresponds to the minimum cost value. By solving the five equations in the five unknowns, (x0, y0) is calculated as:

[0215] x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0)))

[0216] y0=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))

[0217] (x0, y0) can be calculated to any desired sub-pixel precision by adjusting the precision with which the division is performed (i.e., how many bits of the quotient are calculated). th Pixel precision, only 4 bits of the absolute value of the quotient need to be computed, which makes it suitable for a fast shift-subtract based implementation of 2 divisions associated with each CU.

[0218] 3. Add the calculated (x0, y0) to the integer distance refinement MV to get the sub-pixel accurate refinement increment MV.

[0219] The magnitude of the derived fractional motion vector is constrained to be less than or equal to half a pixel.

[0220] To further simplify the DMVR process, JVET-M0147 proposes several changes to the design in JEM. More specifically, the DMVR design adopted by VTM-4.0 (to be released soon) has the following key features:

[0221] Early termination when the SAD at position (0, 0) between list 0 and list 1 is less than a threshold.

[0222] • Early termination when the SAD between list 0 and list 1 is zero for a certain position.

[0223] DMVR block size: W*H>=64&&H>=8, where W and H are the width and height of the block.

[0224] Split the CU into multiples of 16x16 sub-blocks of DMVR for CU size > 16*16. If only the width or height of the CU is greater than 16, it is split only in the vertical or horizontal direction.

[0225] • Reference block size (W+7)*(H+7) (for luma).

[0226] 25-point SAD-based integer pixel search (i.e. (+-2) refinement search range, single stage)

[0227] DMVR based on bilinear interpolation.

[0228] Sub-pixel refinement based on the “parametric error surface equation”. This process is performed only when the minimum SAD cost is not equal to zero and the best MVD is (0, 0) in the last MV refinement iteration.

[0229] • Luma / Chroma MC w / reference block padding (if needed).

[0230] ●MV used only for refinement of MC and TMVP.

[0231] 2.3.5.2.1 Use of DMVR

[0232] DMVR can be enabled when all of the following conditions are true:

[0233] – The DMVR enabled flag in SPS (i.e., sps_dmvr_enabled_flag) is equal to 1

[0234] –TPM flag, inter-frame affine flag, sub-block Merge flag (ATMVP or affine Merge), and MMVD flag are all equal to 0

[0235] –Merge flag equal to 1

[0236] – The current block is bidirectionally predicted, and the POC distance between the current picture and the reference pictures in list 1 is equal to the POC distance between the reference pictures in list 0 and the current picture

[0237] – The height of the current CU is greater than or equal to 8

[0238] – Luma samples (CU width * height) greater than or equal to 64

[0239] 2.3.5.2.2. Reference samples in DMVR

[0240] For a block of size W*H, assuming the maximum allowed MVD value is + / - offset (e.g., 2 in VVC), and the filter size is filterSize (e.g., 8 for luma and 4 for chroma in VVC), then (W+2*offSet+filterSize–1)*(H+2*offSet+filterSize–1) reference samples can be used. To reduce memory bandwidth, only the center (W+filterSize–1)*(H+filterSize–1) reference samples are extracted, and the other pixels are generated by repeating the boundaries of the extracted samples. An example of an 8x8 block is shown in Figure 1. Figure 23 shown.

[0241] These reference samples are used to perform bilinear motion compensation during motion vector refinement and final motion compensation.

[0242] 2.3.6. Combined Intra- and Inter-frame Prediction

[0243] In JVET-L0100, multi-hypothesis prediction is proposed, where combined intra-frame and inter-frame prediction is a way to generate multiple hypotheses.

[0244] When multi-hypothesis prediction is applied to improve intra mode, multi-hypothesis prediction combines an intra prediction and a Merge index prediction. In a Merge CU, when the flag is true, a flag is signaled for the Merge mode to select the intra mode from the intra candidate list. For the luma component, the intra candidate list is derived from 4 intra prediction modes including DC, planar, horizontal and vertical modes, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. Weighted averaging is used to combine one intra prediction mode selected by the intra mode index and one merge index prediction selected by the merge index. For chroma components, DM is always applied without additional signaling. The weights used for combined prediction are described as follows. When DC or planar mode is selected, or the CB width or height is less than 4, equal weights are applied. For those CBs whose width and height are greater than or equal to 4, when horizontal / vertical mode is selected, one CB is first divided vertically / horizontally into four equal-area regions. Each weight set (denoted as (w_intra i ,w_inter i ), where i is from 1 to 4 and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5) and (w_intra4, w_inter4) = (2, 6)) will be applied to the corresponding areas. (w_intra1, w_inter1) is used for the area closest to the reference sample, while (w_intra4, w_inter4) is used for the area farthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the predicted value can be saved for subsequent reference by neighboring CUs.

[0245] 3. Disadvantages of existing implementation methods

[0246] During motion vector refinement, DMVR and BIO do not involve the original signal, which may result in codec blocks with inaccurate motion information. In addition, DMVR and BIO sometimes use fractional motion vectors after motion refinement, while screen video usually has integer motion vectors, which makes the current motion information even more inaccurate and leads to worse codec performance.

[0247] 4. Example Techniques and Embodiments

[0248] The detailed embodiments described below should be considered as examples for explaining the general concept. These embodiments should not be interpreted narrowly. In addition, these embodiments can be combined in any way.

[0249] In addition to DMVR and BIO mentioned below, the method described below can also be applied to other decoder motion information derivation techniques.

[0250] 1. Whether and / or how DMVR is applied to a prediction unit / codec block / region may depend on a message (e.g., a flag) such as signaled in a sequence (e.g., SPS) / slice (e.g., slice header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0251] a. In one example, a flag may be signaled to indicate whether DMVR is enabled. Alternatively, when such a flag indicates that DMVR is disabled, the DMVR process is skipped and the use of DMVR is inferred to be disabled.

[0252] b. In one example, the signaling message indicating whether and / or how to apply MMVD may also control the use of other techniques such as DMVR.

[0253] i. For example, a flag (eg, tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether and / or how DMVR is applied.

[0254] c. In one example, if tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is not allowed for MMVD, the process of DMVR can be skipped.

[0255] 2. Whether and / or how BIO is applied to a prediction unit / codec block / region may depend on a message (e.g., a flag) such as signaled in a sequence (e.g., SPS) / slice (e.g., slice header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0256] a. In one example, a flag is signaled to indicate that the BIO is enabled. Alternatively, when such a flag indicates that the BIO is disabled, the BIO process is skipped and the BIO is inferred to be disabled.

[0257] b. In one example, the signaling message indicating whether and / or how to apply MMVD may also control the use of other techniques, such as BIO techniques.

[0258] 1. For example, a flag (eg, tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether and / or how BIO is applied.

[0259] c. In one example, if tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is not allowed for MMVD, the BIO process may be skipped.

[0260] 3. Whether to enable or disable the DMVR process for a prediction unit / codec block / region may be determined based on the initial motion vector (ie, the decoded motion vector before applying DMVR) and / or the reference picture.

[0261] a. In one example, when both initial motion vectors are integer motion vectors, the use of DMVR can be disabled.

[0262] b. In one example, DMVR is always disabled when the reference picture does not point to certain pictures, such as when the reference picture index is not equal to 0.

[0263] c. In one example, when the magnitude of the initial motion vector is greater than T, the use of DMVR may be disabled.

[0264] i. In one example, "a motion vector having a magnitude greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0265] ii. In one example, “a motion vector having a magnitude greater than T” is defined as the sum of mv.x and mv.y being greater than T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0266] iii. Additionally, alternatively, the use of DMVR may be disabled when the magnitude of the initial motion vector is greater than T or / and the initial motion vector is an integer motion vector. T may be based on:

[0267] a) Current block size

[0268] b) Current quantization parameter

[0269] c) The magnitude of the initial motion vector

[0270] d) The magnitude of the motion vector of the neighboring block

[0271] e) Alternatively, T may be signaled from the encoder to the decoder.

[0272] 4. Whether to enable or disable the BIO process for a prediction unit / codec block / region may be determined based on the initial motion vector (ie, the motion vector before applying BIO) and / or the reference picture.

[0273] a. In one example, when both initial motion vectors are integer motion vectors, the use of BIO can be disabled.

[0274] b. Alternatively, the BIO process for a prediction unit / codec block / region may be enabled or disabled based on the motion vector difference between the initial motion vector and the refined motion vector. In one example, when such a motion vector difference is sub-pixel, the update of the prediction samples / reconstructed samples is skipped.

[0275] c. In one example, when the reference picture does not point to certain pictures, such as when the reference picture index is not equal to 0, the BIO is always disabled.

[0276] d. In one example, when the magnitude of the initial motion vector is greater than T, the use of BIO may be disabled.

[0277] i. In one example, "a motion vector having a magnitude greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0278] ii. In one example, “a motion vector having a magnitude greater than T” is defined as the sum of mv.x and mv.y being greater than T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0279] iii. Additionally, alternatively, the use of BIO may be disabled when the magnitude of the initial motion vector is greater than T or / and the initial motion vector is an integer motion vector. T may be based on:

[0280] a) Current block size

[0281] b) Current quantization parameter

[0282] c) The magnitude of the initial motion vector

[0283] d) The magnitude of the motion vector of the neighboring block

[0284] e) Alternatively, T may be signaled from the encoder to the decoder.

[0285] 5. The second step in enabling or disabling DMVR for a prediction unit / codec block / region may be determined based on a message (e.g., a flag) such as signaling in a sequence (e.g., SPS) / slice (e.g., slice header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0286] a. In one example, a flag is signaled to indicate whether the second step in DMVR applies.

[0287] i. In one example, when a sign like this

[0288] When (tile_group_subpixel_refinement_enabled_flag, etc.) indicates that the second step in DMVR is disabled, the process of fractional refinement in DMVR is skipped.

[0289] b. In one example, the flag indicating signaling notification of enabling or disabling the second step in DMVR may also indicate information unrelated to the second step in DMVR.

[0290] i. In one example, a flag (eg, tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether the second step in DMVR is applied.

[0291] ii. In one example, if tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is disabled, the second step in DMVR is skipped.

[0292] iii. In one example, a flag (eg, tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether the second step in DMVR is applied.

[0293] 6. Whether to enable or disable the second step in DMVR for a prediction unit / codec block / region may be determined based on the result of integer motion refinement in DMVR.

[0294] c. In one example, when the initial motion vector is not changed after integer motion refinement in DMVR, the use of the second step in DMVR may be disabled.

[0295] d. In one example, the magnitude of a fractional motion vector may be allowed to be greater than half a pixel, where the magnitude of the fractional motion vector may be constrained to be less than or equal to T pixels.

[0296] i. In one example, T >= 0 and T is a floating point number. (For example, T = 1.5).

[0297] ii. In one example, the definition of "the magnitude of a motion vector is less than T" is mv.x < T and / or mv.y < T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0298] iii. In one example, the definition of "the magnitude of a motion vector is less than T" is that the sum of mv.x and mv.y is less than T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0299] iv. Additionally, alternatively, when the magnitude of the fractional motion vector is less than T, the magnitude of the fractional motion vector may be allowed to be greater than half a pixel. T may be based on:

[0300] a) The current block size

[0301] b) The current quantization parameter

[0302] c) The magnitude of the initial motion vector

[0303] d) The magnitude of the motion vector of neighboring blocks

[0304] e) Alternatively, T may be signaled from the encoder to the decoder.

[0305] e. In one example, when the distance between the initial motion vector and the motion vector obtained in the integer motion refinement in DMVR is less than T, the second step in DMVR may be disabled.

[0306] i. In one example, the definition of "the magnitude of a motion vector is less than T" is mv.x < T and / or mv.y < T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0307] ii. In one example, the definition of "the magnitude of a motion vector is less than T" is that the sum of mv.x and mv.y is less than T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0308] iii. Additionally, alternatively, the second step in DMVR may be disabled when the distance between the initial motion vector and the motion vector obtained in integer motion refinement in DMVR is less than T. T may be based on:

[0309] a) Current block size

[0310] b) Current quantization parameter

[0311] c) The magnitude of the initial motion vector

[0312] d) The magnitude of the motion vector of the neighboring block

[0313] e) Alternatively, T may be signaled from the encoder to the decoder.

[0314] 7. Whether to enable or disable the second step in DMVR for a prediction unit / codec block / region can be determined based on the initial motion vector in DMVR.

[0315] a. In one example, when both initial motion vectors are integer motion vectors, the use of the second step in DMVR can be disabled.

[0316] b. In one example, when the reference picture does not point to certain pictures, such as when the reference picture index is not equal to 0, the second step in DMVR is always disabled.

[0317] c. In one example, when the magnitude of the initial motion vector is greater than T, the use of the second step in DMVR may be disabled.

[0318] i. In one example, "a motion vector having a magnitude greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0319] ii. In one example, “a motion vector having a magnitude greater than T” is defined as the sum of mv.x and mv.y being greater than T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0320] iii. Additionally, alternatively, the second step in DMVR may be disabled when the magnitude of the initial motion vector is greater than T or / and the initial motion vector is an integer motion vector. T may be based on:

[0321] a) Current block size

[0322] b) Current quantization parameter

[0323] c) The magnitude of the initial motion vector

[0324] d) The magnitude of the motion vector of the neighboring block

[0325] e) Alternatively, T may be signaled from the encoder to the decoder.

[0326] 8. The accuracy of the initial motion vector in DMVR can be determined based on a signaling message such as signaling in sequence (e.g., SPS) / slice (e.g., slice header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0327] f. In one example, a flag is signaled to indicate whether the initial motion vector in the DMVR is rounded to an integer motion vector.

[0328] g. In one example, the signaled flag indicating whether the initial motion vector in the DMVR is rounded to an integer motion vector may also indicate information unrelated to the initial motion vector in the DMVR.

[0329] i. In one example, a flag (e.g., tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether the initial motion vector in BIO and / or DMVR is rounded to an integer motion vector.

[0330] 9. The accuracy of the initial motion vector in the BIO may depend on signaling messages such as signaling in sequence (e.g., SPS) / slice (e.g., slice header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0331] h. In one example, a flag is signaled to indicate whether the initial motion vector in the BIO is rounded to an integer motion vector.

[0332] i. In one example, the signaled flag indicating whether the initial motion vector in the BIO is rounded to an integer motion vector may also indicate information unrelated to the initial motion vector in the BIO.

[0333] i. In one example, a flag (e.g., tile_group_fpel_mmvd_enabled_flag) indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD, also known as UMVE) may also indicate whether the initial motion vector in BIO and / or DMVR is rounded to an integer motion vector.

[0334] 10. Which MVD candidates (such as those used in the first and / or second steps) are changed adaptively based on the accuracy of the initialized motion vectors.

[0335] j. In one example, when the initialized motion vector is a sub-pixel motion vector, the checking of the provisional motion vector (derived from the initialized motion vector and one MVD candidate) may be skipped if the provisional motion vector is an integer-pixel motion vector.

[0336] k. Optionally, when the initialized motion vector is a sub-pixel motion vector, the checking of the provisional motion vector (derived from the initialized motion vector and one MVD candidate) may be skipped if the provisional motion vector is a sub-pixel motion vector.

[0337] 5. Additional Embodiments

[0338] Text changes in the VVC draft are shown in the table below in underlined bold italic font.

[0339] 7.3.2. Raw Byte Sequence Payload, Trailing Bit, and Byte Alignment Syntax

[0340] 7.3.2.1. Sequence Parameter Set RBSP Syntax

[0341]

[0342]

[0343]

[0344]

[0345]

[0346] sps_subpixel_refinement_enabled_flag equal to 1 specifies that adaptive subpixel MVD refinement may be used in a slice. sps_fpel_mmvd_enabled_flag specifies that adaptive subpixel MVD refinement may not be used in a slice.

[0347] 7.3.4. Slice Group Header Syntax

[0348] 7.3.4.1. Common Slice Group Header Syntax

[0349]

[0350]

[0351]

[0352]

[0353]

[0354]

[0355]

[0356] tile_group_subpixel_refinement_enabled_flag equal to 1 specifies that sub-pixel MVD refinement may be enabled in the current slice group. tile_group_fpel_mmvd_enabled_flag equal to 0 specifies that sub-pixel MVD refinement may not be enabled in the current slice group. When not present, the value of tile_group_fpel_mmvd_enabled_flag is inferred to be 1.

[0357] 8.5.3. Decoder-side Motion Vector Refinement Process

[0358] 8.5.3.1. General

[0359] The inputs to this process are:

[0360] – Luma position (xSb, ySb), specifies the upper left sample of the current codec sub-block relative to the upper left luminance sample of the current picture,

[0361] – The variable sbWidth specifies the width of the current codec sub-block in luminance samples.

[0362] –Variable sbHeight specifies the height of the current codec sub-block in luminance samples.

[0363] – Luma motion vectors mvL0 and mvL1 in 1 / 16 fractional sample accuracy,

[0364] – The selected luma reference picture sample arrays refPicL0L and refPicL1L.

[0365] The output of this process is:

[0366] – Delta luminance motion vectors dMvL0 and dMvL1.

[0367] The variable subPelFlag is set to 0. And the variables srRange, offsetH0, offsetH1, offsetV0, and offsetV1 are all set equal to 2.

[0368] Both components of the delta-luminance motion vectors dMvL0 and dMvL1 are set equal to zero and modified as follows:

[0369] For each X that is 0 or 1, pass the reference picture sample array refPicLX with the luma position (xSb, ySb), the prediction block width set equal to (sbWidth+2*srRange), the prediction block height set equal to (sbHeight+2*srRange), L , motion vector mvLX and refined search range srRange as input, call the fractional sample bilinear interpolation process specified in 8.5.3.2.1 to derive the (sbWidth+2*srRange)x(sbHeight+2*srRange) array predSamplesLX of predicted luminance sample values L .

[0370] – By sbWidth, sbHeight, offsetH0, offsetH1, offsetV0, offsetV1, predSamplesL0 L and predSamplesL1 L As input, the sum of absolute differences calculation procedure specified in 8.5.3.3 is called to derive the list sadList[i] (where i = 0..8).

[0371] – When sadList[4] is greater than or equal to sbHeight*sbWidth, the following applies:

[0372] – The variable bestldx is derived by calling the array entry selection procedure specified in clause 8.5.3.4 with the list sadList[i] (where i = 0..8) as input.

[0373] – If bestIdx is equal to 4, subPelFlag is set equal to 1.

[0374] – Otherwise, the following applies:

[0375] dX = bestIdx%3-1

[0376] dY=bestIdx / 3-1

[0377] dMvL0[0]+=16*dX

[0378] dMvL0[1]+=16*dY

[0379] offsetH0+=dX

[0380] offsetV0+=dY

[0381] offsetH1-=dX

[0382] offsetV1-=dY

[0383] – By sbWidth, sbHeight, offsetH0, offsetH1, offsetV0, offsetV1, predSamplesL0 L and predSamplesL1 L As input, the sum of absolute differences calculation procedure specified in 8.5.3.3 is called to modify the list sadList[i] (where i = 0..8).

[0384] – Modify the variable bestIdx by calling the array entry selection procedure specified in clause 8.5.3.4 with the list sadList[i] (where i = 0..8) as input.

[0385] – If bestIdx is equal to 4, subPelFlag is set equal to 1.

[0386] – Otherwise (bestIdx is not equal to 4), the following applies:

[0387] dMvL0[0]+=16*(bestIdx%3-1)

[0388] dMvL0[1]+=16*(bestIdx / 3-1)

[0389] – If tile_group_subpixel_refinement_enabled_flag is equal to 1

[0390] – When subPelFlag is equal to 1, the parametric motion vector refinement process specified in clause 8.5.3.5 is called with the list sadList[i] (where i=0..8) and the incremental motion vector dMvL0 as input and with the modified dMvL0 as output.

[0391] – The incremental motion vector dMvL1 is derived as follows:

[0392] dMvL1[0]=-dMvL0[0]

[0393] dMvL1[1]=-dMvL0[1]

[0394] 9. Example Implementations of the Disclosed Technology

[0395] Figure 24is a block diagram of a video processing device 2400. The device 2400 can be used to implement one or more of the methods described herein. The device 2400 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. The device 2400 may include one or more processors 2402, one or more memories 2404, and video processing hardware 2406. The processor(s) 2402 can be configured to implement one or more methods described in this document. The memory(s) 2404 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2406 can be used to implement some of the techniques described in this document in hardware circuits, and can be partially or completely part of the processor 2402 (e.g., a graphics processor core GPU or other signal processing circuit).

[0396] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may, for example, correspond to bits that are collocated or dispersed at different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded according to transform and codec error residual values, and also using bits in headers and other fields in the bitstream.

[0397] It will be appreciated that the disclosed methods and techniques will benefit video encoder and / or decoder embodiments incorporated into video processing devices, such as smartphones, laptops, desktop computers, and similar devices, by enabling use of the techniques disclosed in this document.

[0398] Figure 25 is a flow chart of an example method 2500 of video processing. The method 2500 includes, at 2510, performing a conversion between a current video block and a bitstream representation of the current video block, wherein use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion and a second value of the flag indicates that decoder motion information is disabled during the conversion.

[0399] Some embodiments may be described using the following clause-based format.

[0400] 1. A method for visual media processing, comprising:

[0401] Performing a conversion between the current video block and a bitstream representation of the current video block, wherein use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion and a second value of the flag indicates that decoder motion information is disabled during the conversion.

[0402] 2. The method of clause 1, wherein the decoder motion information comprises decoder motion video refinement (DMVR) or bidirectional optical flow (BIO).

[0403] 3. A method according to any of clauses 1-2, wherein the flag is denoted tile_group_fpel_mmvd_enabled_flag.

[0404] 4. A method according to any of clauses 1-3, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD) mode.

[0405] 5. A method according to any of clauses 1-3, wherein the flag indicates whether fractional motion vector difference (MVD) is not allowed in Merge with Motion Vector Difference (MMVD) mode.

[0406] 6. The method of clause 2, wherein the decoder motion information derivation comprises at least one of: one or more initial motion vectors associated with the video block, or one or more reference pictures associated with the video block, wherein the one or more initial motion vectors are decoded motion vectors before using DMVR or BIO.

[0407] 7. The method of clause 6, wherein the one or more initial motion vectors associated with the video block include two initial motion vectors, further comprising:

[0408] In response to determining that both initial motion vectors are integer values, DMVR is skipped.

[0409] 8. The method according to clause 6, further comprising:

[0410] In response to determining that the indices of one or more reference pictures associated with the video block are non-zero, DMVR is skipped.

[0411] 9. The method according to clause 6, further comprising:

[0412] In response to determining that a magnitude of one or more initial motion vectors associated with the video block is greater than a threshold, DMVR is skipped.

[0413] 10. A method according to clause 9, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in the one or more initial motion vectors, wherein mv.x>T and / or mv.y>T, and wherein T represents a threshold value.

[0414] 11. A method according to clause 9, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in the one or more initial motion vectors, where sum(mv.x,mv.y)>T, where T represents a threshold, and where sum(x,y)=x+y.

[0415] 12. A method according to any of clauses 10-11, wherein the motion vector is a motion vector of a current video block.

[0416] 13. A method according to any of clauses 10-12, wherein the threshold value is based on one or more of: the size of the video block, a quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of motion vectors associated with blocks adjacent to the video block, or a value signaled from a visual media processing encoder to a visual media processing decoder.

[0417] 14. The method of clause 6, wherein the one or more initial motion vectors associated with the video block include two initial motion vectors, further comprising:

[0418] In response to determining that both initial motion vectors are integer values, the BIO is skipped.

[0419] 15. The method of clause 6, wherein the one or more initial motion vectors associated with the video block include an initial motion vector and a refinement of the initial motion vector, further comprising:

[0420] Based on the difference between the calculated initial motion vector and the refined one, BIO is enabled or skipped.

[0421] 16. The method according to clause 15, further comprising:

[0422] In response to determining that the difference between the initial motion vector and the refinement is sub-pixel, prediction associated with the video block is skipped.

[0423] 17. The method according to clause 6, further comprising:

[0424] In response to determining that the indices of one or more reference pictures associated with the video block are non-zero, the BIO is skipped.

[0425] 18. The method of clause 6, further comprising:

[0426] In response to determining that a magnitude of one or more initial motion vectors associated with the video block is greater than a threshold, the BIO is skipped.

[0427] 19. A method according to clause 18, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in the one or more initial motion vectors, wherein mv.x>T and / or mv.y>T, and wherein T represents a threshold value.

[0428] 20. A method according to clause 18, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in the one or more initial motion vectors, where sum(mv.x,mv.y)>T, where T represents a threshold, and where sum(x,y)=x+y.

[0429] 21. A method according to any of clauses 19-20, wherein the motion vector is a motion vector of a current video block.

[0430] 22. A method according to any of clauses 19-20, wherein the threshold value is based on one or more of: the size of the video block, a quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of motion vectors associated with blocks adjacent to the video block, or a value signaled from the visual media processing encoder to the visual media processing decoder.

[0431] 23. A method according to any of clauses 1-2, wherein DMVR comprises a second step of score refinement, wherein the flag indicates whether score refinement should be skipped.

[0432] 24. A method according to clause 23, wherein the flag is denoted as tile_group_subpixel_refinement_enabled_flag.

[0433] 25. The method of clause 23, wherein the flag indicates information not relevant to the second step of score refinement.

[0434] 26. The method according to clause 25, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in the Merge with Motion Vector Difference (MMVD) mode.

[0435] 27. The method according to clause 26, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0436] 28. The method according to clause 25, wherein if the flag indicates whether fractional motion vector difference (MVD) is disabled, the second step of fractional refinement is skipped.

[0437] 29. The method according to clause 26, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0438] 30. The method according to any one of clauses 1-2, wherein the DMVR includes a second step of fractional refinement based on the result of integer motion refinement.

[0439] 31. The method according to clause 30, further comprising:

[0440] Skipping the second step of DMVR in response to determining that the initial motion vector of the video block has not changed after integer motion refinement.

[0441] 32. The method according to clause 31, wherein the magnitude of the fractional motion vector is greater than half a pixel, and wherein the magnitude of the fractional motion vector is constrained to be less than or equal to a threshold.

[0442] 33. The method according to clause 31, wherein the threshold is a non-zero floating point number.

[0443] 34. The method according to clause 32, wherein the magnitude of the fractional motion vector associated with the video block includes the magnitude of the horizontal component of the fractional motion vector (represented as mv.x) and the magnitude of the vertical component of the fractional motion vector (represented as mv.y), and wherein mv.x < T and / or mv.y < T, and wherein T represents the threshold.

[0444] 35. The method according to clause 32, wherein the magnitude of the fractional motion vector associated with the video block includes the magnitude of the horizontal component of the fractional motion vector (represented as mv.x) and the magnitude of the vertical component of the fractional motion vector (represented as mv.y), and wherein sum(mv.x, mv.y) < T, where T represents the threshold, and wherein sum(x, y) = x + y.

[0445] 36. The method according to any one of clauses 32 - 35, wherein the threshold is based on one or more of the following: the size of the video block, the quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of the motion vectors associated with blocks adjacent to the video block, or a value signaled from the visual media processing encoder to the visual media processing decoder.

[0446] 37. The method according to clause 30, wherein the difference vector is calculated as the difference between the initial motion vector of the video block and the motion vector obtained after integer motion refinement, further comprising:

[0447] In response to determining that the magnitude of the difference vector is less than the threshold, skipping the second step of DMVR.

[0448] 38. The method according to clause 37, wherein the magnitude of the difference vector associated with the video block includes the magnitude of the horizontal component of the difference vector (denoted as mv.x) and the magnitude of the vertical component of the difference vector (denoted as mv.y), and wherein mv.x < T and / or mv.y < T, and wherein T represents the threshold.

[0449] 39. The method according to clause 37, wherein the magnitude of the difference vector associated with the video block includes the magnitude of the horizontal component of the difference vector (denoted as mv.x) and the magnitude of the vertical component of the difference vector (denoted as mv.y), and wherein sum(mv.x, mv.y) < T, where T represents the threshold, and wherein sum(x, y) = x + y.

[0450] 40. The method according to any one of clauses 37 - 39, wherein the threshold is based on one or more of the following: the size of the video block, the quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of the motion vectors associated with blocks adjacent to the video block, or a value signaled from the visual media processing encoder to the visual media processing decoder.

[0451] 41. The method according to any one of clauses 1 - 2, wherein the flag indicates the precision of the initial motion vector associated with the video block.

[0452] 42. The method according to clause 41, wherein the flag indicates whether the initial motion vector is rounded to an integer value.

[0453] 43. The method according to clause 42, wherein the flag indicates information not related to the initial motion vector.

[0454] 44. The method according to clause 43, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in the Merge with Motion Vector Difference (MMVD) mode.

[0455] 45. A method according to clause 44, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0456] 46. ​​A method of visual media processing, comprising:

[0457] performing a conversion between the current video block and a bitstream representation of the current video block, wherein use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion and a second value of the flag indicates that decoder motion information is disabled during the conversion; and

[0458] In response to determining that the initial motion vector for the video block has sub-pixel precision, checking associated with a temporary motion vector derived from the initial motion vector and the candidate motion vector difference is skipped, wherein the temporary motion vector has integer pixel precision or sub-pixel precision.

[0459] 47. A method according to any one or more of clauses 1-46, wherein the flag is included in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a picture header, a slice group header, a slice header, a codec tree unit row or a region associated with a codec tree unit

[0460] 48. A method as described in any one or more of clauses 1 to 47, wherein visual media processing is implemented on the encoder side.

[0461] 49. A method as described in any one or more of clauses 1 to 47, wherein visual media processing is implemented on the decoder side.

[0462] 50. An apparatus in a video system comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one or more of clauses 1 to 49.

[0463] 51. A computer program product stored on a non-transitory computer readable medium, the computer program product comprising program code for performing the method according to any one or more of clauses 1 to 49.

[0464] Figure 26 26 is a flow chart of an example method 2600 for video processing. The method 2600 includes, at 2602, determining whether and / or how to apply decoder-side motion vector refinement (DMVR) for a conversion between a first block of video and a bitstream representation of the first block of video based on signaled information; and, at 2604, performing the conversion based on the determination.

[0465] In some examples, the signaled information includes a flag present in the bitstream representation.

[0466] In some examples, the signaled information is signaled at at least one of: a sequence level including a sequence parameter set (SPS), a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

[0467] In some examples, during the conversion, DMVR is applied to at least one of a prediction unit, a codec block, and a region.

[0468] In some examples, a flag is signaled to indicate whether DMVR is enabled or disabled.

[0469] In some examples, when the flag indicates that DMVR is disabled, the process of DMVR is skipped and the use of DMVR is inferred to be disabled.

[0470] In some examples, signaling information indicating whether and / or how to apply DMVR is used to control the use of other decoder motion information derivation techniques.

[0471] In some examples, other decoder motion information derivation techniques include bidirectional optical flow (BIO).

[0472] In some examples, a flag indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD) mode is used to determine whether and / or how DMVR is applied.

[0473] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0474] In some examples, if the flag indicates that fractional MVD is not allowed for MMVD, the DMVR process is skipped during the conversion.

[0475] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0476] In some examples, the conversion generates a first block of video from the bitstream representation.

[0477] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0478] Figure 2727 is a flow chart of an example method 2700 for video processing. The method 2700 includes, at 2702, determining whether and / or how to apply bidirectional optical flow (BIO) for conversion between a first block of video and a bitstream representation of the first block of video based on signaled information; and, at 2704, performing the conversion based on the determination.

[0479] In some examples, the signaled information includes a flag present in the bitstream representation.

[0480] In some examples, the signaled information is signaled at at least one of: a sequence level including a sequence parameter set (SPS), a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

[0481] In some examples, during the conversion, a BIO is applied to at least one of a prediction unit, a codec block, and a region.

[0482] In some examples, a flag is signaled to indicate whether the BIO is enabled or disabled.

[0483] In some examples, when the flag indicates that the BIO is disabled, the process of the BIO is skipped and use of the BIO is inferred to be disabled.

[0484] In some examples, signaling information indicating whether and / or how to apply a BIO is used to control the use of other decoder motion information derivation techniques.

[0485] In some examples, other decoder motion information derivation techniques include Decoder Motion Video Refinement (DMVR).

[0486] In some examples, a flag indicating whether fractional motion vector difference (MVD) is allowed in Merge with Motion Vector Difference (MMVD) mode is used to determine whether and / or how BIO is applied.

[0487] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0488] In some examples, if the flag indicates that fractional MVD is not allowed for MMVD, the BIO process is skipped during the conversion.

[0489] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0490] In some examples, the conversion generates a first block of video from the bitstream representation.

[0491] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0492] Figure 28 is a flow chart of an example method 2800 for video processing. The method 2800 includes, at 2802, determining whether a decoder-side motion vector refinement (DMVR) process is enabled or disabled for a conversion between a first block of video and a bitstream representation of the first block of video based on at least one of: one or more initial motion vectors associated with the first block and one or more reference pictures associated with the first block, the initial motion vectors comprising motion vectors before applying the DMVR process; and, at 2804, performing the conversion based on the determination.

[0493] In some examples, during the conversion, a DMVR process is applied to at least one of a prediction unit, a codec block, and a region associated with the first block.

[0494] In some examples, the initial motion vector includes a first initial motion vector in a first direction and / or a second initial motion vector in a second direction.

[0495] In some examples, the DMVR process is disabled when both initial motion vectors are integer motion vectors.

[0496] In some examples, the DMVR process is disabled when the reference picture does not point to certain pictures.

[0497] In some examples, when the reference picture index is not equal to 0, the DMVR process is disabled.

[0498] In some examples, the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a threshold (T).

[0499] In some examples, the magnitude of each initial motion vector includes a magnitude of a horizontal component (mv.x) and a magnitude of a vertical component (mv.y).

[0500] In some examples, when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold T, the DMVR process is disabled.

[0501] In some examples, when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold T, the DMVR process is disabled.

[0502] In some examples, the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a threshold (T) and / or the initial motion vectors are integer motion vectors.

[0503] In some examples, T is determined based on at least one of: a size of the first block, a current quantization parameter, a magnitude of one or more initial motion vectors, and a magnitude of motion vectors of neighboring blocks of the first block.

[0504] In some examples, T is signaled from the encoder to the decoder.

[0505] In some examples, the DMVR process includes a first step in which integer motion vector difference (MVD) refinement is performed with integer precision and a second step in which fractional MVD refinement is performed with fractional precision.

[0506] In some examples, the method further includes determining whether the second step of the DMVR process is enabled or disabled based on the signaled information.

[0507] In some examples, the signaled information includes a flag present in the bitstream representation.

[0508] In some examples, the signaled information is signaled at at least one of: a sequence level including a sequence parameter set (SPS), a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

[0509] In some examples, a flag is signaled to indicate whether the second step of the DMVR process is enabled or disabled.

[0510] In some examples, the flag is tile_group_subpixel_refinement_enabled_flag.

[0511] In some examples, when the flag indicates that the second step of the DMVR process is disabled, the score refinement process in the DMVR process is skipped.

[0512] In some examples, the flag indicating whether the second step in the DMVR process is enabled or disabled also indicates additional information in addition to the second step in the DMVR process.

[0513] In some examples, a flag indicating whether fractional MVD is allowed in Merge with Motion Vector Difference (MMVD) mode is used to indicate whether the second step of the DMVR process is enabled or disabled.

[0514] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0515] In some examples, the method further includes determining how to perform a second step of the DMVR process based on a result of the integer motion refinement in the first step of the DMVR process.

[0516] In some examples, when the initial motion vector is not changed after integer motion refinement in the first step of the DMVR process, the second step in the DMVR process is disabled.

[0517] In some examples, the magnitude of the motion vector obtained after the second step is allowed to be greater than half a pixel and less than or equal to a second threshold T2 pixels.

[0518] In some examples, T2>=0, and T2 is a floating point number.

[0519] In some examples, the magnitude of the motion vector includes the magnitude of a horizontal component (mv.x) and the magnitude of a vertical component (mv.y).

[0520] In some examples, one or more of the horizontal component mv.x and the vertical component mv.y of the motion vector is constrained to be smaller than T2 pixels.

[0521] In some examples, the sum of the horizontal component mv.x and the vertical component mv.y of the motion vector is constrained to be less than T2 pixels.

[0522] In some examples, when the distance between the initial motion vector and the motion vector obtained after integer motion refinement in the first step of the DMVR process is less than a second threshold T2, the second step in the DMVR process is disabled.

[0523] In some examples, the distance includes a magnitude of a horizontal component (mv.x) and a magnitude of a vertical component (mv.y).

[0524] In some examples, when one or more of the horizontal component mv.x and the vertical component mv.y of the distance is less than a threshold value T2, the second step in the DMVR process is disabled.

[0525] In some examples, when the sum of the horizontal component mv.x and the vertical component mv.y is less than a threshold value T2, the second step in the DMVR process is disabled.

[0526] In some examples, T2 is determined based on at least one of: a size of the first block, a current quantization parameter, a magnitude of one or more initial motion vectors, and a magnitude of motion vectors of neighboring blocks of the first block.

[0527] In some examples, T2 is signaled from the encoder to the decoder.

[0528] In some examples, the method further includes determining whether a second step of the DMVR process is enabled or disabled based on the one or more initial motion vectors in the DMVR process.

[0529] In some examples, when both initial motion vectors are integer motion vectors, the second step of the DMVR process is disabled.

[0530] In some examples, the second step of the DMVR process is disabled when the reference picture does not point to certain pictures.

[0531] In some examples, when the reference picture index is not equal to 0, the second step of the DMVR process is disabled.

[0532] In some examples, the second step of the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a third threshold (T3).

[0533] In some examples, the magnitude of each initial motion vector includes a magnitude of a horizontal component (mv.x) and a magnitude of a vertical component (mv.y).

[0534] In some examples, the second step of the DMVR process is disabled when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold value T3.

[0535] In some examples, the second step of the DMVR process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold value T3.

[0536] In some examples, the second step of the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a third threshold (T3) and / or the initial motion vectors are integer motion vectors.

[0537] In some examples, T3 is determined based on at least one of: a size of the first block, a current quantization parameter, a magnitude of one or more initial motion vectors, and a magnitude of motion vectors of neighboring blocks of the first block.

[0538] In some examples, T3 is signaled from the encoder to the decoder.

[0539] In some examples, the accuracy of the initial motion vector in the DMVR process is determined based on signaled information.

[0540] In some examples, the signaled information includes a flag present in the bitstream representation.

[0541] In some examples, the signaled information is signaled at at least one of: a sequence level including a sequence parameter set (SPS), a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

[0542] In some examples, a flag is signaled to indicate whether the initial motion vector in the DMVR process is rounded to an integer motion vector.

[0543] In some examples, the flag indicating whether the initial motion vector in the DMVR process is rounded to an integer motion vector also indicates additional information in addition to the initial motion vector in the DMVR process.

[0544] In some examples, a flag indicating whether fractional MVD is allowed in Merge with Motion Vector Difference (MMVD) mode is used to indicate whether the initial motion vector in the DMVR process is rounded to an integer motion vector.

[0545] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0546] In some examples, whether to use an MVD candidate is adaptively changed based on the accuracy of the initial motion vector.

[0547] In some examples, the MVD candidates include MVD candidates used in the first step and / or the second step of the DMVR process.

[0548] In some examples, when the initial motion vector is a sub-pixel motion vector, checking of the provisional motion vector is skipped if the provisional motion vector is an integer-pixel motion vector, which is derived from the initial motion vector and an MVD candidate.

[0549] In some examples, when the initial motion vector is a sub-pixel motion vector, checking of the provisional motion vector is skipped if the provisional motion vector is a sub-pixel motion vector, the provisional motion vector being derived from the initial motion vector and one of the MVD candidates.

[0550] Figure 29is a flow chart of an example method 2900 for video processing. The method 2900 includes, at 2902, determining whether a bidirectional optical flow (BIO) process is enabled or disabled for a conversion between a first block of video and a bitstream representation of the first block of video based on at least one of: one or more initial motion vectors associated with the first block and one or more reference pictures associated with the first block, the initial motion vectors including motion vectors before applying the BIO process and / or decoded motion vectors before applying a decoder-side motion vector refinement (DMVR) process; and, at 2904, performing the conversion based on the determination.

[0551] In some examples, during the conversion, a BIO process is applied to at least one of a prediction unit, a codec block, and a region associated with the first block.

[0552] In some examples, the initial motion vector includes a first initial motion vector in a first direction and / or a second initial motion vector in a second direction.

[0553] In some examples, the BIO process is disabled when both initial motion vectors are integer motion vectors.

[0554] In some examples, the BIO process is always disabled when the reference picture does not point to certain pictures.

[0555] In some examples, when the reference picture index is not equal to 0, the BIO process is disabled.

[0556] In some examples, the BIO process is disabled when the magnitude of one or more initial motion vectors is greater than a fourth threshold (T4).

[0557] In some examples, the magnitude of each initial motion vector includes a magnitude of a horizontal component (mv.x) and a magnitude of a vertical component (mv.y).

[0558] In some examples, when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold T, the BIO process is disabled.

[0559] In some examples, the BIO process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold value T4.

[0560] In some examples, the BIO process is disabled when the magnitude of one or more initial motion vectors is greater than a fourth threshold (T4) and / or the initial motion vectors are integer motion vectors.

[0561] In some examples, T4 is determined based on at least one of: a size of the first block, a current quantization parameter, a magnitude of one or more initial motion vectors, and a magnitude of motion vectors of neighboring blocks of the first block.

[0562] In some examples, T4 is signaled from the encoder to the decoder.

[0563] In some examples, the accuracy of the initial motion vector in the BIO process is determined based on signaled information.

[0564] In some examples, the signaled information includes a flag present in the bitstream representation.

[0565] In some examples, the signaled information is signaled at at least one of: a sequence level including a sequence parameter set (SPS), a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

[0566] In some examples, a flag is signaled to indicate whether the initial motion vector in the BIO process is rounded to an integer motion vector.

[0567] In some examples, the flag indicating whether the initial motion vector in the BIO process is rounded to an integer motion vector also indicates additional information in addition to the initial motion vector in the BIO process.

[0568] In some examples, a flag indicating whether fractional MVD is allowed in Merge with Motion Vector Difference (MMVD) mode is used to indicate whether the initial motion vector in the BIO process is rounded to an integer motion vector.

[0569] In some examples, the flag is tile_group_fpel_mmvd_enabled_flag.

[0570] Figure 30 3000 is a flow chart of an example method 3000 for video processing. The method 3000 includes, at 3002, determining whether a bidirectional optical flow (BIO) process is enabled or disabled for conversion between a first block of video and a bitstream representation of the first block of video based on one or more motion vector differences between an initial motion vector associated with the first block of video and one or more refined motion vectors, the initial motion vector comprising a motion vector before applying the BIO process and / or applying a DMVR process, the refined motion vector comprising a motion vector after applying the DMVR process; and, at 3004, performing the conversion based on the determination.

[0571] In some examples, during the conversion, a BIO process is applied to at least one of a prediction unit, a codec block, and a region associated with the first block.

[0572] In some examples, when the motion vector difference is sub-pixel, the update of the prediction samples or the reconstructed samples is skipped.

[0573] In some examples, the conversion generates a first block of video from the bitstream representation.

[0574] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0575] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances that affect a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical signal, an optical signal, or an electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0576] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language (including compiled or interpreted languages), and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.

[0577] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0578] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, or be operatively coupled to receive data from or transfer data to, or vice versa. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into dedicated logic circuitry.

[0579] Although this patent document contains many details, these details should not be interpreted as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable subcombination. Furthermore, although features may be described above as working in certain combinations and even initially claimed as such, one or more features from the claimed combination may be excluded from the combination in some cases, and the claimed combination may be directed to a subcombination or variation of the subcombination.

[0580] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0581] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: determining whether and / or how to apply decoder-side motion vector refinement (DMVR) for a first block of video and a conversion between a bitstream of the first block of video based on the signaled information; as well as performing said converting based on said determining, Wherein, if the signaled information indicates that fractional MVD is not allowed to be used for MMVD, the DMVR process is skipped during the conversion.

2. The method according to claim 1, wherein The signaled information includes a flag present in the bitstream.

3. The method according to claim 1 or 2, wherein: The signaled information is signaled in at least one of the following: a sequence level including a sequence parameter set SPS, a slice level including a slice header, a slice group level including a slice group header, a picture level including a picture header, and a block level including a codec tree unit and a codec unit.

4. The method according to claim 1 or 2, wherein: During the conversion, the DMVR is applied to at least one of a prediction unit, a codec block, and a region.

5. The method according to claim 1 or 2, wherein: A flag is signaled to indicate whether DMVR is enabled or disabled.

6. The method according to claim 5, wherein: When the flag indicates that DMVR is disabled, the process of DMVR is skipped and the use of DMVR is inferred to be disabled.

7. The method according to claim 1 or 2, wherein: The signaled information indicating whether and / or how to apply DMVR is used to control the use of other decoder motion information derivation techniques.

8. The method according to claim 7, wherein: The other decoder motion information derivation techniques include bidirectional optical flow (BIO).

9. The method according to claim 1 or 2, wherein: The signaled information indicating whether fractional motion vector difference (MVD) is allowed in the Merge MMVD mode with motion vector difference is used to determine whether and / or how to apply DMVR.

10. The method according to claim 9, wherein: The signaling notification information is tile_group_fpel_mmvd_enabled_flag.

11. The method according to claim 1, wherein The signaling notification information is tile_group_fpel_mmvd_enabled_flag.

12. The method according to claim 1 or 2, wherein: The conversion generates the first block of video from the bitstream.

13. The method according to claim 1 or 2, wherein: The conversion generates the bitstream from the first block of video.

14. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1 to 13.

15. A computer readable medium storing code which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Decoder Side Motion Vector Refinement in Video Coding

    US20190020895A1