Enabling bio based on information in picture header

By managing motion vectors and signaling notifications, the bitstream representation of video blocks is optimized, overcoming the limitations of compression ratio and complexity in existing video encoding and decoding technologies, and achieving more efficient video encoding and decoding.

CN113545076BActive Publication Date: 2026-03-24DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-03
Publication Date
2026-03-24

Smart Images

  • Figure CN113545076B_ABST
    Figure CN113545076B_ABST
Patent Text Reader

Abstract

Enabling BIO based on information in a picture header is disclosed. A method of video processing includes determining, based on signaled information, whether and / or how to apply bi-directional optical flow (BIO) for a conversion between a first block of a video and a bitstream representation of the first block of the video, and performing the conversion based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] In accordance with applicable patent law and / or the rules of the Paris Convention, this application aims to promptly claim priority and interest in International Patent Application No. PCT / CN2019 / 076788, filed March 3, 2019, and International Patent Application No. PCT / CN2019 / 076860, filed March 4, 2019. The entire disclosure of International Patent Application Nos. PCT / CN2019 / 076788 and PCT / CN2019 / 076860 is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to video encoding and decoding technologies, devices, and systems. Background Technology

[0004] Currently, efforts are underway to improve the performance of existing video codec technologies to provide better compression ratios or to offer video encoding and decoding schemes that allow for lower complexity or parallel implementation. Industry experts have recently proposed several new video encoding and decoding tools, which are currently being tested to determine their effectiveness. Summary of the Invention

[0005] This paper describes devices, systems, and methods related to digital video coding and decoding, and specifically, to the management of motion vectors. The described methods can be applied to existing video coding and decoding standards (e.g., High Efficiency Video Codec (HEVC) or Universal Video Codec) and future video coding and decoding standards or codecs.

[0006] In one representative aspect, the disclosed technology can be used to provide a method for video processing. The method includes: performing a conversion between a current video block and a bitstream representation of the current video block, wherein the use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that the decoder motion information is enabled during the conversion, and a second value of the flag indicates that the decoder motion information is disabled during the conversion.

[0007] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: performing a conversion between a current video block and a bitstream representation of the current video block, wherein the use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that the decoder motion information is enabled during the conversion, and a second value of the flag indicates that the decoder motion information is disabled during the conversion; and, in response to determining that an initial motion vector of the video block has subpixel precision, skipping checks associated with a temporary motion vector derived from the difference between the initial motion vector and candidate motion vectors, wherein the temporary motion vector has integer pixel precision or subpixel precision.

[0008] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: determining, based on signaling notification information, whether and / or how to apply decoder-side motion vector refinement (DMVR) for the conversion between a first block of video and a bitstream representation of that first block; and performing the conversion based on that determination.

[0009] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: determining, based on signaling notification information, whether and / or how to apply bidirectional optical flow (BIO) for the conversion between a first block of video and a bitstream representation of that first block; and performing the conversion based on that determination.

[0010] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: determining whether a decoder-side motion vector refinement (DMVR) process is enabled or disabled for the conversion between a first block of video and a bitstream representation of the first block, based on at least one of the following: one or more initial motion vectors associated with the first block and one or more reference images associated with the first block, wherein the initial motion vectors include motion vectors prior to the application of the DMVR process; and performing the conversion based on this determination.

[0011] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: determining whether a bidirectional optical flow (BIO) process is enabled or disabled for the conversion between a first block of video and a bitstream representation of the first block, based on at least one of the following: one or more initial motion vectors associated with the first block and one or more reference images associated with the first block, wherein the initial motion vectors include motion vectors prior to the application of the BIO process and / or decoded motion vectors prior to the application of a decoder-side motion vector refinement (DMVR) process; and performing the conversion based on this determination.

[0012] In another representative aspect, the disclosed technology can be used to provide another method for video processing. This method includes: determining whether a bidirectional optical flow (BIO) process is enabled or disabled for the conversion between the first block of video and its bitstream representation, based on one or more motion vector differences between an initial motion vector associated with a first block of video and one or more refined motion vectors; the initial motion vectors include motion vectors before the application of the BIO process and / or the application of the DMVR process; and the refined motion vectors include motion vectors after the application of the DMVR process; and performing the conversion based on this determination.

[0013] Furthermore, in a representative aspect, an apparatus for a video system is disclosed, including a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform any or more of the disclosed methods.

[0014] In addition, a computer program product stored on a non-transitory computer-readable medium is disclosed, the computer program product including program code for performing any one or more of the disclosed methods.

[0015] The above and other aspects and features of the disclosed technology are described in more detail in the accompanying drawings, description and claims. Attached Figure Description

[0016] Figure 1 An example of constructing a Merge candidate list is shown.

[0017] Figure 2 An example of a candidate location for the airspace is shown.

[0018] Figure 3 An example of a candidate pair for which a spatial merge candidate is performed is shown.

[0019] Figure 4A and Figure 4B An example of the position of the second prediction unit (PU) based on the size and shape of the current block is shown.

[0020] Figure 5 An example of motion vector scaling for temporal Merge candidates is shown.

[0021] Figure 6 An example of candidate locations for temporal Merge candidates is shown.

[0022] Figure 7 An example of generating bidirectional prediction Merge candidates using a combination is shown.

[0023] Figure 8An example of constructing motion vector prediction candidates is shown.

[0024] Figure 9 An example of motion vector scaling for spatial motion vector candidates is shown.

[0025] Figure 10 An example of deriving local illumination compensation parameters for neighboring samples is shown.

[0026] Figure 11A and Figure 11B Illustrations related to the 4-parameter affine model and the 6-parameter affine model are shown respectively.

[0027] Figure 12 An example of the affine motion vector field for each sub-block is shown.

[0028] Figure 13A and Figure 13B Examples of 4-parameter affine models and 6-parameter affine models are shown respectively.

[0029] Figure 14 An example of motion vector prediction for affine inter-frame patterns used for inheritance is shown.

[0030] Figure 15 An example of motion vector prediction for affine inter-frame patterns used to construct affine candidates is shown.

[0031] Figure 16A and Figure 16B An illustration related to the affine Merge pattern is shown.

[0032] Figure 17 An example of candidate positions for the affine Merge pattern is shown.

[0033] Figure 18 An example of the Merge with Motion Vector Differences (MMVD) pattern search process is shown.

[0034] Figure 19 An example of an MMVD search point is shown.

[0035] Figure 20 An example of decoder-side motion video refinement (DMVR) in JEM7 is shown.

[0036] Figure 21 An example of motion vector difference (MVD) related to DMVR is shown.

[0037] Figure 22An example illustrating the inspection of motion vectors is shown.

[0038] Figure 23 An example of a reference sample in a DMVR is shown.

[0039] Figure 24 This is a block diagram of an example hardware platform used to implement the visual media decoding or visual media encoding technologies described in this document.

[0040] Figure 25 A flowchart of an example method for video encoding and decoding is shown.

[0041] Figure 26 A flowchart of an example method for video encoding and decoding is shown.

[0042] Figure 27 A flowchart of an example method for video encoding and decoding is shown.

[0043] Figure 28 A flowchart of an example method for video encoding and decoding is shown.

[0044] Figure 29 A flowchart of an example method for video encoding and decoding is shown.

[0045] Figure 30 A flowchart of an example method for video encoding and decoding is shown. Detailed Implementation

[0046] 1. Video encoding and decoding in HEVC / H.265

[0047] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) standards, as well as the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Expert Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, with the goal of a 50% bitrate reduction compared to HEVC.

[0048] The latest version of the VVC draft (i.e., Universal Video Codec (Draft 4)) can be found at: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 13_Marrakesh / wg11 / JVET-M1001-v5.zip. The latest reference software for VVC (called VTM) can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tree / VTM-4.0.

[0049] 2.1. Inter-frame prediction in HEVC / H.265

[0050] Each inter-frame prediction unit (PU) has motion parameters for one or two lists of reference images. The motion parameters include motion vectors and reference image indices. The use of one of the two reference image lists can also be signaled using `inter_pred_idc`. The motion vectors can be explicitly encoded as increments relative to the predicted values.

[0051] When encoding and decoding a CU in skip mode, a PU is associated with the CU and there are no significant residual coefficients, no encoded motion vector increments, or reference picture indices. A Merge mode is specified, thereby obtaining the motion parameters of the current PU from neighboring PUs, including spatial and temporal candidates. The Merge mode can be applied to any inter-frame prediction PU, not just skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the predicted motion vector value), the corresponding reference picture index for each reference picture list, and the reference picture list are explicitly signaled per PU. Such a mode is named Advanced Motion Vector Prediction (AMVP) in this disclosure.

[0052] When signaling indicates that one of two lists of reference images should be used, a PU is generated from a sample block. This is called "one-way prediction". One-way prediction applies to both P-strips and B-strips.

[0053] When signaling indicates that two reference image lists should be used, a PU is generated from two sample blocks. This is called "bidirectional prediction". Bidirectional prediction is only applicable to B-strips.

[0054] The following text provides details about the inter-frame prediction modes specified in HEVC. The description will begin with the Merge mode.

[0055] 2.1.1. List of Reference Images

[0056] In HEVC, the term inter-frame prediction is used to describe predictions derived from data elements (e.g., sample values ​​or motion vectors) of reference images other than the currently decoded image. As in H.264 / AVC, images can be predicted from multiple reference images. The reference images used for inter-frame prediction are organized into one or more reference image lists. A reference index identifies which reference image in the list should be used to create the predicted signal.

[0057] A single list of reference images (List 0) is used for P-strips, and two lists of reference images (List 0 and List 1) are used for B-strips. It should be noted that the reference images included in List 0 / 1 can be selected based on past and future images in terms of capture / display order. Furthermore, the current image can be on the list of reference images in HEVC version 4.

[0058] 2.1.2. Merge Mode

[0059] 2.1.2.1. Derivation of the candidate Merge pattern

[0060] When predicting a PU using the Merge mode, indices pointing to entries in the Merge candidate list are parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:

[0061] Step 1: Initial Candidate Derivation

[0062] Step 1.1: Spatial Candidate Derivation

[0063] Step 1.2: Redundancy check of airspace candidates

[0064] Step 1.3: Time-domain candidate derivation

[0065] Step 2: Adding candidate insertions

[0066] Step 2.1: Create bidirectional prediction candidates

[0067] Step 2.2: Insert zero-motion candidates

[0068] exist Figure 1 These steps are also illustrated schematically. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since it is assumed at the decoder that the number of candidates per PU is constant, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, the index of the best merge candidate is encoded using truncated unary (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of a 2N×2N prediction unit.

[0069] The operations associated with the foregoing steps are described in detail below.

[0070] 2.1.2.2. Derivation of Airspace Candidates

[0071] In the derivation of the spatial Merge candidate, from the location located Figure 2Up to four merge candidates are selected from the candidates at the positions depicted. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another strip or slice) or if it is intra-frame encoding / decoding. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only those at positions A1, B1, B0, A0, and B2 are considered. Figure 3 The pairs linked by arrows are added to the list only if the candidates used for redundancy checking do not have the same motion information. Another source of duplicate motion information is the "second PU" associated with partitions different from 2N×2N. As an example, Figure 4 depicts the second PU in the cases of N×2N and 2N×N. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for an encoding / decoding unit with only one PU. Similarly, when the current PU is partitioned into 2N×N, position B1 is not considered.

[0072] 2.1.2.3. Derivation of Time-Domain Candidates

[0073] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the juxtaposed PU of the image that has the smallest POC difference with the current image within the given list of reference images. The list of reference images that will be used for the derivation of the juxtaposed PU is displayed in the strip header. Figure 5 The dashed lines in the diagram show the scaled motion vectors for obtaining the temporal merge candidate. These motion vectors are scaled from the motion vectors of the juxtaposed PU using POC distances tb and td, where tb is defined as the POC difference between the current image and its reference image, and td is defined as the POC difference between the juxtaposed image and its reference image. The reference image index for the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For the B-strip, two motion vectors are obtained, one for reference image list 0 and the other for reference image list 1, and they are combined to form a bidirectional prediction merge candidate.

[0074] like Figure 6The description describes the selection of a temporal candidate position between candidate C0 and C1 within the juxtaposed PU(Y) belonging to the reference frame. Position C1 is used if the PU at position C0 is unavailable, intra-frame encoded, or outside the current codec tree unit (CTU, aka LCU, maximum codec unit) row. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0075] 2.1.2.4. Additional Candidate Insertion

[0076] In addition to the spatiotemporal merge candidate, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated by utilizing the spatiotemporal merge candidate. Combined bidirectional prediction merge candidates are only used for B-strips. Combined bidirectional prediction candidates are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of the other. If the two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. As an example, Figure 7 This describes the case where two candidates from the original list (on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create bidirectional predictive merge candidates that are added to the final list (on the right). There are many rules regarding combinations that are considered to generate these additional merge candidates.

[0077] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have zero spatial displacements and a reference image index that starts at zero and increments whenever a new zero-motion candidate is added to the list. Finally, no redundancy checks are performed on these candidates.

[0078] 2.1.3.AMVP

[0079] AMVP utilizes the spatiotemporal correlation between motion vectors and neighboring PUs for explicit transmission of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporally neighboring PU locations on the left and top, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best prediction from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate uses truncated unary coding. In this case, the maximum value to be encoded is 2 (see...). Figure 8 The following sections provide details of the derivation process for the motion vector prediction candidates.

[0080] 2.1.3.1. Derivation of AMVP Candidates

[0081] Figure 8 The derivation process of motion vector prediction candidates is summarized.

[0082] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, based on the location as... Figure 2 The motion vectors of each PU at the five different locations are depicted to ultimately derive two motion vector candidates.

[0083] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different juxtaposition positions. After generating the first spatiotemporal candidate list, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than two, motion vector candidates with reference image indices greater than 1 are removed from the associated reference image list. If the number of spatiotemporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.

[0084] 2.1.3.2. Candidate Spatial Motion Vectors

[0085] In the derivation of the spatial motion vector candidates, at most two candidates are considered from five potential candidates. These five potential candidates are selected from those located in the spatial domain, such as... Figure 2 The positions depicted are derived from the PUs, which are the same as the positions of the motion merge. The derivation order to the left of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order above the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not use spatial scaling and two of which do. These four different cases are summarized as follows:

[0086] • No spatial scaling

[0087] –(1) Same list of reference images and same index of reference images (same POC)

[0088] –(2) Different lists of reference images but the same reference image (same POC)

[0089] • Spatial scaling

[0090] –(3) Same list of reference images but different reference images (different POCs)

[0091] –(4) Different lists of reference images and different reference images (different POCs)

[0092] First, check for cases without spatial scaling, then check for spatial scaling. Regardless of the reference image list, consider spatial scaling when the Proof of Concept (POC) differs between the reference image of a neighboring PU and the reference image of the current PU. If all candidate PUs on the left are unavailable or intra-frame encoded / decoded, scaling of the upper motion vector is allowed to aid in the parallel derivation of the left and upper motion vectors. Otherwise, spatial scaling of the upper motion vector is not allowed.

[0093] like Figure 9 The described process involves scaling the motion vectors of neighboring PUs in a manner similar to temporal scaling during spatial scaling. The main difference is that a list of reference images and the index of the current PU are given as input; the actual scaling process is the same as that of temporal scaling.

[0094] 2.1.3.3. Candidate Motion Vectors in the Temporal Domain

[0095] Except for the derivation of the reference image index, all the procedures for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (see [link]). Figure 6 The reference image index is signaled to the decoder.

[0096] 2.2. Local Illumination Compensation in JEM

[0097] Local illumination compensation (LIC) is based on a linear model of illumination variation, using a scaling factor a and an offset b. It is adaptively enabled or disabled for each inter-frame mode codec's codec unit (CU).

[0098] When LIC is applied to CU, the least squares method is used to derive parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, as... Figure 10 As shown, the neighboring and corresponding sample points (identified by the motion information of the current CU or subCU) of the CU in the reference image are used for subsampling (2:1 subsampling).

[0099] 2.2.1 Derivation of the prediction block

[0100] The LIC parameters are derived and applied for each prediction direction. For each prediction direction, a first prediction block is generated using the decoded motion information, and then a temporary prediction block is obtained by applying the LIC model. Finally, the final prediction block is derived using the two temporary prediction blocks.

[0101] When encoding and decoding the CU in Merge mode, the LIC flag is copied from the neighboring block in a manner similar to motion information copying in Merge mode; otherwise, the LIC flag is signaled to the CU to indicate whether LIC is applicable.

[0102] When LIC is enabled for an image, additional CU-level RD checks are required to determine if LIC is suitable for the CU. When LIC is enabled for the CU, the Mean-Removed Sum of Absolute Difference (MR-SAD) and the Mean-Removed Sum of Absolute Hadamard-Transformed Difference (MR-SATD), instead of SAD and SATD, are used for integer pixel motion thinning and fractional pixel motion thinning, respectively.

[0103] To reduce coding complexity, the following coding scheme is applied in JEM.

[0104] • When there is no significant lighting change between the current image and its reference images, LIC is disabled for the entire image. To identify this situation, histograms of the current image and each reference image of the current image are calculated at the encoder. If the histogram difference between the current image and each reference image of the current image is less than a given threshold, LIC is disabled for the current image; otherwise, LIC is enabled for the current image.

[0105] 2.3. Inter-frame prediction methods in VVC

[0106] Several new codec tools for improving inter-frame prediction exist, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling notification MVD, Affine Prediction Mode, Triangular Prediction Mode (TPM), Advanced TMVP (ATMVP, also known as SbTMVP), Generalized Bi-Prediction (GBI), and Bi-directional Optical Flow (BIO).

[0107] 2.3.1. Encoder / decoder block structure in VVC

[0108] In VVC, images are divided into square or rectangular blocks using a quadtree / binary tree / multi-type tree (QT / BT / TT) structure.

[0109] Besides QT / BT / TT, a separate tree (also known as a dual codec tree) is also used in VVC for I-strips / slices. In the case of a separate tree, the codec block structure is signaled separately for the luma and chroma components.

[0110] 2.3.2. Adaptive Motion Vector Difference Resolution

[0111] In HEVC, when `use_integer_mv_flag` in the stripe header is equal to 0, the motion vector difference (MVD) is signaled in units of quarter-luminance samples (i.e., between the PU's motion vector and the predicted motion vector). In VVC, Local Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, MVD can be encoded and decoded in units of quarter-luminance samples, integer luminance samples, or four luminance samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). MVD resolution is controlled at the codec unit (CU) level, and for each CU with at least one non-zero MVD component, the MVD resolution flag is conditionally signaled.

[0112] For a CU with at least one non-zero MVD component, a signaling flag is used to indicate whether quarter-luminance sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter-luminance sample MV precision is not used, another flag is signaled to indicate whether integer luminance sample MV precision or four-luminance sample MV precision is used.

[0113] When the first MVD resolution flag of the CU is zero, or when no encoding or decoding is performed for the CU (meaning all MVDs in the CU are zero), the CU uses a quarter-lumen sample MV resolution. When the CU uses integer lumen sample MV precision or four-lumen sample MV precision, the MVPs in the CU's AMVP candidate list are rounded to the corresponding precision.

[0114] 2.3.3. Affine Motion Compensation Prediction

[0115] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). While there are many types of motion in the real world, such as zooming in / out, rotation, perspective motion, and other irregular motions, VVC uses simplified affine transformation motion compensation prediction with 4-parameter and 6-parameter affine models. As shown in Figure 11, the affine motion field of a block is described by two control point motion vectors (CPMV) of the 4-parameter affine model and three CPMVs of the 6-parameter affine model.

[0116] The motion vector field (MVF) of the block is described by the following equations using the 4-parameter affine model in equation (1) (where the 4 parameters are defined as variables a, b, e, and f) and the 6-parameter affine model in equation (2) (where the 4 parameters are defined as variables a, b, c, d, e, and f):

[0117]

[0118]

[0119] Where (mv h 0,mv h 0) is the motion vector of the top-left control point, and (mv h 1, MV h 1) is the motion vector of the upper right control point, and (mv h 2, MV h 2) is the motion vector of the lower left control point. All three motion vectors are called the control point motion vectors (CPMV). (x, y) represents the coordinates of the point relative to the upper left sample point within the current block, and (mv) h (x, y), mv v (x, y) is the motion vector derived for the sample point located at (x, y). The CP motion vector can be signaled (e.g., in affine AMVP mode) or derived on the fly (e.g., in affine Merge mode). w and h are the width and height of the current block. In practice, division is implemented by right shift and rounding. In VTM, the representative point is defined as the center position of the sub-block; for example, when the coordinates of the top-left corner of the sub-block relative to the top-left sample point within the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (i.e., 4×4 in VTM), the motion vector of the entire sub-block is derived using the representative point.

[0120] To further simplify motion compensation prediction, a sub-block-based affine transformation prediction was applied. To derive the motion vector for each M×N (M and N are set to 4 in the current VVC) sub-block, the motion vector of the center sample point of each sub-block (e.g., ...) is... Figure 12 The result (as shown) can be calculated according to equations (1) and (2) and rounded to 1 / 16 fractional precision. Then, a 1 / 16 pixel motion compensation interpolation filter can be applied to generate predictions for each sub-block with derived motion vectors. The affine mode introduces a 1 / 16 pixel interpolation filter.

[0121] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the standard motion vector.

[0122] 2.3.3.1. Signaling notification for affine prediction

[0123] Similar to the translational motion model, due to affine prediction, there are also two modes for signaling notification side information: AFFINE_INTER and AFFINE_MERGE modes.

[0124] 2.3.3.2. AF_INTER mode

[0125] For CUs with both width and height greater than 8, the AF_INTER mode can be applied. The affine flag at the CU level is signaled in the bitstream to indicate whether the AF_INTER mode is used.

[0126] In this mode, for each list of reference images (list 0 or list 1), the affine AMVP candidate list is constructed in the following order using three types of affine motion prediction values, where each candidate includes the estimated CPMV of the current block. The best CPMV found on the encoder side (such as...) Figure 15 The difference between mv0, mv1, and mv2 in the estimated CPMV and the CPMV is signaled. Furthermore, a further signaling notification is given for the index of the affine AMVP candidate derived from the estimated CPMV.

[0127] 1) Inherited affine motion prediction values

[0128] The inspection order is similar to that of the spatial MVP in the HEVC AMVP list. First, inherited affine motion prediction values ​​from the left side of the first block derivation, which is affine-coded from {A1, A0} and has the same reference picture as the current block. Second, inherited affine motion prediction values ​​from above the first block derivation, which is affine-coded from {B1, B0, B2} and has the same reference picture as the current block. Figure 14 The five blocks A1, A0, B1, B0, and B2 are depicted in the text.

[0129] Once a neighboring block is found to be encoded in affine mode, the CPMV of the codec unit covering the neighboring block is used to derive the predicted value of the CPMV for the current block. For example, if A1 is encoded in non-affine mode and A0 is encoded in 4-parameter affine mode, the inherited affine MV prediction value on the left will be derived from A0. In this case, the CPMV of the CU covering A0 (as shown in...) Figure 16B Zhongyou The upper left corner CPMV and the symbol formed by... The upper right CPMV (represented by the CPMV) is used to derive the estimated CPMV of the current block, by... This indicates the top left (coordinates (x0, y0)), top right (coordinates (x1, y1)), and bottom right (coordinates (x2, y2)) positions of the current block.

[0130] 2) Constructed affine motion prediction values

[0131] like Figure 15As shown, the constructed affine motion predictions include control-point motion vectors (CPMVs) derived from neighboring inter-frame codec blocks with the same reference image. If the current affine motion model is a 4-parameter affine, the number of CPMVs is 2; otherwise, if the current affine motion model is a 6-parameter affine, the number of CPMVs is 3. (CPMVs are shown in the upper left corner.) The MV is derived from the first block in group {A, B, C} that is inter-coded and has the same reference picture as the current block. (CPMV in the upper right corner) The MV derivation is based on the first block in group {D, E} that is inter-coded and has the same reference image as the current block. (Lower left CPMV) The MV is derived from the first block in group {F, G} that is inter-frame encoded and decoded and has the same reference picture as the current block.

[0132] –If the current affine motion model is a 4-parameter affine, then only if and Only when both are established will the constructed affine motion predictions be inserted into the candidate list, that is, and Used as the top left (coordinates (x0, y0)) and top right (coordinates (x1, y1)) of the current block.

[0133] CPMV of location estimation.

[0134] –If the current affine motion model is a 6-parameter affine, then only if and Only after all parameters are established will the constructed affine motion prediction values ​​be inserted into the candidate list, that is, and CPMV is used to estimate the positions of the current block at the top left (coordinates (x0, y0)), top right (coordinates (x1, y1)), and bottom right (coordinates (x2, y2)).

[0135] When the constructed affine motion predictions are inserted into the candidate list, no pruning process is applied.

[0136] 3) Normal AMVP motion prediction value

[0137] The following applies until the number of affine motion predictions reaches its maximum value.

[0138] 1) If available, by setting all CPMV to equal To derive the predicted value of affine motion.

[0139] 2) If available, by setting all CPMV to equal To derive the predicted value of affine motion.

[0140] 3) If available, by setting all CPMV to equal To derive the predicted value of affine motion.

[0141] 4) If available, derive the affine motion predictions by setting all CPMVs to equal HEVC TMVP.

[0142] 5) The affine motion predictions were derived by setting all CPMVs to zero MV.

[0143] Please note, It has already been derived in the constructed affine motion prediction values.

[0144] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are used. Therefore, 2 / 3 MVDs need to be encoded and decoded for these control points, as shown in Figure 13. The following MV derivation is proposed: predict mvd1 and mvd2 from mvd0.

[0145]

[0146]

[0147]

[0148] in, mvd i mv1 and mv2 are the predicted motion vector, motion vector difference, and motion vector of the top-left pixel (i=0), top-right pixel (i=1), or bottom-left pixel (i=2), respectively. Figure 13B As shown. Note that the sum of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components separately, that is, newMV = mvA + mvB, and the two components of newMV are set as (xA + xB) and (yA + yB) respectively.

[0149] 2.3.3.3.AF_Merge mode

[0150] When the CU is applied in AF_MERGE mode, it obtains the first block encoded and decoded in affine mode from the effective neighboring reconstructed blocks. The selection order of candidate blocks is from left, top, upper right, lower left to upper left, as follows: Figure 16A As shown (represented by A, B, C, D, E in order). For example, if the adjacent lower left block is like... Figure 16B If A0 represents the code that is encoded and decoded in affine mode, then the motion vector mv0 of the control points (CPs) at the top left, top right, and bottom left corners of the neighboring CUs / PUs containing block A is obtained.N mv1 N and mv2 N And based on mv0 N mv1 N and mv2 N Calculate the motion vector mv0 of the top left / top right / bottom left corner on the current CU / PU. C mv1 C and mv2 C (For 6-parameter affine model only). It should be noted that in the current VTM, the top-left sub-block (e.g., a 4×4 block in the VTM) stores mv0, and the top-right sub-block stores mv1 if the current block is affine encoded. If the current block is encoded using a 6-parameter affine model, the bottom-left sub-block stores mv2; otherwise (using a 4-parameter affine model), the LB stores mv2'. Other sub-blocks store the MV used for MCs.

[0151] The CPMV mv0 of the current CU is derived based on the simplified affine motion model in equations (1) and (2). C mv1 C and mv2 C Next, the MVF of the current CU is generated. To identify whether the current CU is encoded or decoded in AF_MERGE mode, the affine flag is signaled in the bitstream when at least one neighboring block is encoded or decoded in affine mode.

[0152] The affine Merge candidate list is constructed using the following steps:

[0153] 1) Insertion of affine candidates for inheritance

[0154] Inherited affine candidates are those derived from the affine motion models of their effective neighboring affine codec blocks. The two largest inherited affine candidates are derived from the affine motion models of neighboring blocks and inserted into the candidate list. For the left-hand prediction, the scan order is {A0, A1}; for the upper-hand prediction, the scan order is {B0, B1, B2}.

[0155] 2) Insertion of constructed affine candidates

[0156] If the number of candidates in the Affine Merge candidate list is less than MaxNumAffineCand

[0157] (For example, 5), then the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to the candidate constructed by combining the neighboring motion information of each control point.

[0158] a) The motion information of the control points first comes from... Figure 17The derivation is performed using the specified spatial and temporal neighbors. CPk (k = 1, 2, 3, 4) represents the k-th control point. A0,

[0159] A1, A2, B0, B1, B2, and B3 are used to predict the spatial location of CPk (k = 1, 2, 3); T is used to predict the temporal location of CP4.

[0160] The coordinates of CP1, CP2, CP3 and CP4 are (0, 0), (W, 0), (H, 0) and (W, H) respectively, where W and H are the width and height of the current block.

[0161] Motion information for each control point is obtained according to the following priority order:

[0162] – For CP1, the check priority is B2->B3->A2. If B2 is available, then B2 is used. Otherwise, if B2 is available, then B3 is used. If neither B2 nor B3 is available, then A2 is used. If all three candidates are unavailable, motion information for CP1 cannot be obtained.

[0163] – For CP2, the check priority is B1->B0.

[0164] – For CP3, the inspection priority is A1->A0.

[0165] – For CP4, use T.

[0166] b) Next, use combinations of control points to construct affine Merge candidates.

[0167] I. Motion information from three control points is needed to construct a 6-parameter affine candidate. The three control points can be selected from one of the following four combinations ({CP1, CP2, CP4}, ...).

[0168] {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}).

[0169] Combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3,

[0170] CP4 will be converted into a 6-parameter motion model represented by control points at the top left, top right, and bottom left.

[0171] II. Motion information from two control points is needed to construct a 4-parameter affine candidate. The two control points can be selected from one of two combinations ({CP1, CP2}, {CP1, CP3}). These two combinations will be converted into a 4-parameter motion model represented by the upper left and upper right control points.

[0172] III. The combinations of constructed affine candidates are inserted into the candidate list in the following order:

[0173] {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4},

[0174] {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}

[0175] i. For each combination, check the reference index of list X for each CP. If they are all the same, then the combination has a valid CPMV for list X. If the combination does not have a valid CPMV for either list 0 or list 1, then the combination is marked as invalid. Otherwise, it is valid, and the CPMV is added to the sub-block Merge list.

[0176] 3) Fill with zero motion vector

[0177] If the number of candidates in the affine Merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.

[0178] More specifically, for the sub-block Merge candidate list, MV is set to (0, 0) and the prediction direction is set to the 4-parameter Merge candidate from the unidirectional prediction (for P-strips) and bidirectional prediction (for B-strips) of list 0.

[0179] In VTM4, the CPMV of an affine CU is stored in a separate buffer. The stored CPMV is used only to generate inherited CPMVPs for the most recently encoded CU in affine Merge and Affixed AMVP modes. Sub-block MVs inherited from the CPMV are used for motion compensation, MV derivation of the Merge / AMVP list for translation MVs, and deblocking. To avoid attaching CPMVs to the picture line buffer, affine motion data inheritance from CUs in the upper CTU is handled differently than inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the upper CTU line, the lower left and lower right sub-block MVs in the line buffer are used for affine MVP derivation instead of the CPMV. Thus, the CPMV is only stored in the local buffer. If the candidate CU is 6-parameter affine encoded, the affine model is downgraded to a 4-parameter model.

[0180] 2.3.4. Merge with Motion Vector Difference (MMVD)

[0181] JVET-L0054 presents the Ultimate Motion Vector Expression (UMVE, also known as MMVD). UMVE, along with a proposed motion vector expression method, is used for skip or merge modes.

[0182] Figure 18 An example of the final vector representation (UMVE) search process is shown. Figure 19 An example of a UMVE search point is shown. In VVC, UMVE reuses the same Merge candidates included in the regular Merge candidate list. Among these Merge candidates, a basic candidate can be selected and further extended using the proposed motion vector representation method.

[0183] UMVE provides a new method for representing motion vector difference (MVD), in which MVD is represented by the starting point, motion amplitude, and motion direction.

[0184] This proposed technique uses the Merge candidate list as is. However, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE extensions.

[0185] The basic candidate index defines the starting point. The basic candidate index indicates the best candidate among the candidates in the list, as shown below.

[0186] Table 1. Basic Candidate IDX

[0187] Basic candidate IDX 0 1 2 3 Nth MVP 1st MVP 2nd MVP 3rd MVP 4th MVP

[0188] If the number of basic candidates is equal to 1, then the basic candidate IDX will not be notified by signaling.

[0189] The distance index is motion amplitude information. The distance index indicates a predefined distance from the starting point. The predefined distances are shown below:

[0190] Table 2. Distance from IDX

[0191]

[0192] The direction index represents the direction of MVD relative to the starting point. The direction index can represent the four directions shown below.

[0193] Table 3. Directional IDX

[0194] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0195] Immediately after transmitting the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, the UMVE flag is resolved. If the UMVE flag is equal to 1, the UMVE syntax is resolved. However, if it is not 1, the AFFINE flag is resolved. If the AFFINE flag is equal to 1, it is in AFFINE mode; otherwise, the skip / merge index is resolved to the VTM's skip / merge mode.

[0196] There is no need for an additional line buffer due to UMVE candidates. The software's skip / merge candidates are used directly as the base candidates. The MV supplement is determined immediately before motion compensation using the input UMVE index. There is no need to reserve a long line buffer for this.

[0197] Under the current general testing conditions, the first or second Merge candidate in the Merge candidate list can be selected as the basic candidate.

[0198] UMVE is also known as Merge with MV difference (MMVD).

[0199] Furthermore, the `tile_group_fpel_mmvd_enabled_flag` flag in the stripe header informs the decoder signaling whether fractional distance is used. When fractional distance is disabled, all distances in the default table are multiplied by 4, i.e., the distance table {1,2,4,8,16,32,64,128} pixels is used. Because the size of the distance table remains unchanged, the entropy encoding and decoding of the distance index remains unchanged.

[0200] 2.3.5. Decoder-Side Motion Vector Refinement (DMVR)

[0201] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MV) from list 0 and MV from list 1 are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.

[0202] 2.3.5.1. DMVR in JEM

[0203] In the JEM design, motion vectors are refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral templates and reconstructed samples in the reference image, in order to obtain refined MV values ​​without the transmission of additional motion information. Figure 20 An example is depicted. The bilateral template is generated as a weighted combination (i.e., average) of two prediction blocks, one from initial MV0 of list 0 and the other from initial MV1 of list 1, as shown below. Figure 20As shown. The template matching operation involves calculating a cost metric between the generated template and the sample region (around the initial prediction block) in the reference image. For each of the two reference images, the MV that produces the minimum template cost is considered the updated MV for that list, replacing the original MV. In JEM, nine MV candidates are searched for each list. The nine MV candidates include the original MV and eight surrounding MVs offset by one brightness sample from the original MV in the horizontal, vertical, or both directions. Finally, as... Figure 20 The two new MVs shown (i.e., MV0′ and MV1′) are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Note that when calculating the cost of a prediction block generated by a surrounding MV, the predicted block is actually obtained using the rounded MV (to integer pixels), not the actual MV.

[0204] 2.3.5.2. DMVR in VVC

[0205] For DMVR in VVC, assuming an MVD mirrored between list 0 and list 1, such as... Figure 21 As shown, bilateral matching is performed to refine the motion vector (MV), i.e., to find the optimal MVD among several MVD candidates. Let MVL0(L0X, L0Y) and MVL1(L1X, L1Y) represent the MVs of two lists of reference images. The MVD of list 0, represented by (MvdX, MvdY), that minimizes the cost function (e.g., SAD) is defined as the optimal MVD. For the SAD function, it is defined as the SAD between the reference block of list 0 and the reference block of list 1, where the reference block of list 0 is derived using the motion vectors (L0X+MvdX, L0Y+MvdY) in the reference image of list 0, and the reference block of list 1 is derived using the motion vectors (L1X-MvdX, L1Y-MvdY) in the reference image of list 1.

[0206] The motion vector thinning process can be iterated twice. In each iteration, up to six MVDs (with integer pixel precision) can be checked in two steps, such as... Figure 22 As shown. In the first step, MVD(0,0), (-1,0), (1,0), (0,-1), and (0,1) are checked. In the second step, one of MVD(-1,-1), (-1,1), (1,-1), or (1,1) can be selected and further checked. Assume the function Sad(x,y) returns the SAD value of MVD(x,y). The MVD checked in the second step, represented by (MvdX, MvdY), is determined as follows:

[0207] MvdX = -1;

[0208] MvdY = -1;

[0209] If (Sad(1,0)) <Sad(-1,0))

[0210] MvdX = 1;

[0211] If (Sad(0,1) <Sad(0,-1))

[0212] MvdY = 1;

[0213] In the first iteration, the starting point is the MV of the signaling notification, and in the second iteration, the starting point is the MV of the signaling notification plus the best MVD selected in the first iteration. DMVR is only applicable when one reference picture is the preceding picture and the other reference picture is the following picture, and both reference pictures have the same picture order count distance to the current picture.

[0214] Furthermore, as mentioned above, the DMVR in VVC first performs integer MVD refinement. This is the first step. Afterward, fractional-precision MVD refinement is conditionally performed to further refine the motion vector. This is the second step. The condition for performing the second step is based on whether the MVD after the current iteration is zero MV. If it is zero MV (the vertical and horizontal components of the MV are 0), then the second step is performed.

[0215] The details of fractional MVD refinement are given below. It should be noted that MVD represents the difference between the initial and final motion vectors used in motion compensation.

[0216] The parameter error surface is fitted using integer distance locations and the costs evaluated at those locations, and then used to determine 1 / 16. th Pixel-precise subpixel offset.

[0217] The proposed methods are summarized as follows:

[0218] 1. Parameter error surface fitting is calculated only if the center position is the optimal cost position in a given iteration.

[0219] 2. The center location cost and the costs at positions (-1, 0), (0, -1), (1, 0), and (0, 1) from the center are used to fit the 2-D parabolic error surface equation of the shape.

[0220] E(x,y)=A(x-x0) 2 +B(y-y0) 2 +C

[0221] Where (x0, y0) corresponds to the position of minimum cost, and C corresponds to the minimum cost value. By solving the five equations with five unknowns, (x0, y0) is calculated as:

[0222] x0=(E(-1,0)-E(1,0)) / 2((E(-1,0)+E(1,0)-2E(0,0)))

[0223] y0=(E(0,-1)-E(0,1)) / ((2((E(0,-1)+E(0,1)-2E(0,0)))

[0224] (x0, y0) can be calculated to any desired sub-pixel precision by adjusting the precision of the division (i.e., how many bits of the quotient are calculated). For 1 / 16 th Pixel precision requires only 4 bits of the absolute value of the quotient, making it suitable for a fast shift-based subtraction implementation with two divisions associated with each CU.

[0225] 3. Add the calculated (x0, y0) to the integer distance thinning MV to obtain the precise thinning increment MV for each subpixel.

[0226] The magnitude of the derived fractional motion vector is constrained to be less than or equal to half a pixel.

[0227] To further simplify the DMVR process, JVET-M0147 proposes several changes to the design in JEM. More specifically, the DMVR design adopted by VTM-4.0 (coming soon) has the following key features:

[0228] • Early termination when the SAD at the (0,0) position between list 0 and list 1 is less than the threshold.

[0229] • Early termination when the SAD between list 0 and list 1 is zero at some position.

[0230] • DMVR block size: W*H>=64&&H>=8, where W and H are the width and height of the block.

[0231] • Divide the CU into multiples of 16×16 sub-blocks of the DMVR with a CU size > 16*16. If only the width or height of the CU is greater than 16, it is divided only in the vertical or horizontal direction.

[0232] • Reference block size (W+7)*(H+7) (for brightness).

[0233] • 25-point SAD-based integer pixel search (i.e., (+-2) refinement of the search range, single stage)

[0234] • DMVR based on bilinear interpolation.

[0235] • Subpixel thinning based on the “parameter error surface equation”. This process is performed only if the minimum SAD cost is not equal to zero and the optimal MVD is (0, 0) in the last MV thinning iteration.

[0236] • Luminosity / Chromaticity MC w / Reference Block Fill (if needed).

[0237] ● MVs for detailed use only in MC and TMVP.

[0238] 2.3.5.2.1 Use of DMVR

[0239] DMVR can be enabled when all of the following conditions are true:

[0240] – The DMVR enable flag in SPS (i.e., sps_dmvr_enabled_flag) is equal to 1.

[0241] – The TPM flag, inter-frame affine flag, sub-block merge flag (ATMVP or affine merge), and MMVD flag are all equal to 0.

[0242] –Merge flag equals 1

[0243] – The current block is predicted bidirectionally, and the POC distance between the current image and the reference images in list 1 is equal to the POC distance between the reference images in list 0 and the current image.

[0244] - The current CU height is greater than or equal to 8

[0245] – The number of luminance samples (CU width * height) is greater than or equal to 64

[0246] 2.3.5.2.2. Reference Samples in DMVR

[0247] For a block of size W*H, assuming the maximum allowed MVD value is + / - offset (e.g., 2 in VVC) and the filter size is filterSize (e.g., 8 for luma and 4 for chroma in VVC), then (W + 2*offset + filterSize – 1) * (H + 2*offset + filterSize – 1) reference samples can be used. To reduce memory bandwidth, only the center (W + filterSize – 1) * (H + filterSize – 1) reference samples are extracted, and other pixels are generated by repeating the boundaries of the extracted samples. An example of an 8x8 block is shown below. Figure 23 As shown.

[0248] During motion vector refinement, these reference samples are used to perform bilinear motion compensation. Simultaneously, these reference samples are also used to perform final motion compensation.

[0249] 2.3.6. Combined Intra-Frame and Inter-Frame Prediction

[0250] In JVET-L0100, multi-hypothesis prediction is proposed, in which combined intra-frame and inter-frame prediction is one way to generate multiple hypotheses.

[0251] When applying multiple hypothesis prediction (MMP) to improve intra-mode, MMP combines an intra-prediction and a Merge index prediction. In the Merge CU, when a flag is true, a flag is signaled for Merge mode to select an intra-mode from the intra-candidate list. For the luma component, the intra-candidate list is derived from four intra-prediction modes: DC, planar, horizontal, and vertical, and the size of the intra-candidate list can be 3 or 4 depending on the block shape. The horizontal mode is not included in the intra-mode list when the CU width is greater than twice the CU height, and the vertical mode is removed from the intra-mode list when the CU height is greater than twice the CU width. A weighted average is used to combine an intra-prediction mode selected by the intra-mode index and a Merge index prediction selected by the Merge index. For the chroma component, DM is always applied without additional signaling. The weights used to combine the predictions are described below. Equal weights are applied when DC or planar modes are selected, or when the CB width or height is less than 4. For CBs with a width and height greater than or equal to 4, when selecting horizontal / vertical mode, a CB is first divided vertically / horizontally into four equal-area regions. Each weight set (represented as (w_intra)) i w_inter i ), where i ranges from 1 to 4 and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5), and (w_intra4, w_inter4) = (2, 6) will be applied to the corresponding regions. (w_intra1, w_inter1) is used for the region closest to the reference sample, while (w_intra4, w_inter4) is used for the region furthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. Furthermore, the intra-frame prediction mode of the intra-frame assumptions of the prediction values ​​can be saved for subsequent reference by neighboring CUs.

[0252] 3. Disadvantages of existing implementation methods

[0253] During the refinement of motion vectors, DMVR and BIO do not involve the original signal, which can result in codec blocks with inaccurate motion information. Furthermore, DMVR and BIO sometimes use fractional motion vectors after motion refinement, while screen video typically has integer motion vectors, making the current motion information even more inaccurate and resulting in poorer codec performance.

[0254] 4. Example technologies and implementation examples

[0255] The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted narrowly. Furthermore, these embodiments can be combined in any way.

[0256] In addition to DMVR and BIO mentioned below, the methods described below can also be applied to other decoder motion information derivation techniques.

[0257] 1. Whether and / or how to apply DMVR to prediction units / code / decode blocks / regions can depend on factors such as sequence (e.g., SPS) / strip (e.g., strip header) / piece group (e.g., piece group header).

[0258] Messages (e.g., flags) that are signaled at the image level (e.g., image header) and the block level (e.g., CTU or CU).

[0259] a. In one example, a signaling notification flag can be used to indicate whether the DMVR is enabled. Alternatively, when such a flag indicates that the DMVR is disabled, the DMVR process is skipped, and the use of the DMVR is inferred to be disabled.

[0260] b. In one example, a signaling notification indicating whether and / or how MMVD is applied can also control the use of other technologies, such as DMVR.

[0261] i. For example, a flag indicating whether fractional motion vector difference (MVD) is allowed in a Merge (MMVD, aka UMVE) with motion vector difference (e.g., tile_group_fpel_mmvd_enabled_flag) can also indicate whether and / or how DMVR is applied.

[0262] c. In one example, if tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is not allowed for MMVD, the DMVR process can be skipped.

[0263] 2. Whether and / or how to apply BIO to prediction units / code / decode blocks / regions may depend on messages (e.g., flags) such as signaling notifications in sequences (e.g., SPS) / strips (e.g., strip headers) / slices (e.g., slice headers) / picture levels (e.g., picture headers) / block levels (e.g., CTUs or CUs).

[0264] a. In one example, a signaling notification flag indicates that BIO is enabled. Alternatively, when such a flag indicates that BIO is disabled, the BIO process is skipped, and BIO is inferred to be disabled.

[0265] b. In one example, a signaling notification message indicating whether and / or how to apply MMVD can also control the use of other technologies, such as BIO technology.

[0266] 1. For example, a flag indicating whether fractional motion vector difference (MVD) is allowed in a Merge with motion vector differences (MMVD, also known as UMVE) (e.g.,

[0267] The tile_group_fpel_mmvd_enabled_flag can also indicate whether and / or how to apply BIO.

[0268] c. In one example, if tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is not allowed for MMVD, the BIO process can be skipped.

[0269] 3. It can be based on the initial motion vector (i.e., the decoded motion vector before applying DMVR) and

[0270] Or refer to the image to determine the DMVR process for enabling or disabling prediction units / codec blocks / regions.

[0271] a. In one example, the use of DMVR can be disabled when both initial motion vectors are integer motion vectors.

[0272] b. In one example, DMVR is always disabled when the reference image does not point to certain images, such as when the reference image index is not equal to 0.

[0273] c. In one example, when the magnitude of the initial motion vector is greater than T, it can be disabled.

[0274] The use of DMVR.

[0275] i. In one example, "the magnitude of a motion vector is greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0276] ii. In one example, "the magnitude of a motion vector is greater than T" is defined as the sum of mv.x and mv.y being greater than T, where mv.x, mv.y, and mv.y are respectively greater than T.

[0277] y is the magnitude of the horizontal component of the motion vector, and mv.y is the magnitude of the vertical component of the current motion vector.

[0278] iii. Alternatively, when the magnitude of the initial motion vector is greater than T or /

[0279] When the initial motion vector is an integer motion vector, the use of DMVR can be disabled. T can be based on:

[0280] a) Current block size

[0281] b) Current quantization parameters

[0282] c) The magnitude of the initial motion vector

[0283] d) The magnitude of the motion vector of the neighboring block

[0284] e) Alternatively, signaling T can be sent from the encoder to the decoder.

[0285] 4. The BIO process for enabling or disabling prediction units / codec blocks / regions can be determined based on the initial motion vector (i.e., the motion vector before applying BIO) and / or a reference image.

[0286] a. In one example, when both initial motion vectors are integer motion vectors, the use of BIO can be disabled.

[0287] b. Alternatively, the enabling or disabling of the BIO process for prediction units / encoder blocks / regions can be determined based on the motion vector difference between the initial motion vector and the refined motion vector. In one example, when such a motion vector difference is a sub-pixel, updates to prediction samples / reconstruction samples are skipped.

[0288] c. In one example, BIO is always disabled when the reference image does not point to certain images, such as when the reference image index is not equal to 0.

[0289] d. In one example, the use of BIO can be disabled when the magnitude of the initial motion vector is greater than T.

[0290] i. In one example, "the magnitude of a motion vector is greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0291] ii. In one example, "the magnitude of a motion vector is greater than T" is defined as the sum of mv.x and mv.y being greater than T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0292] iii. Alternatively, BIO can be disabled when the magnitude of the initial motion vector is greater than T or / and the initial motion vector is an integer motion vector. T can be based on:

[0293] a) Current block size

[0294] b) Current quantization parameters

[0295] c) The magnitude of the initial motion vector

[0296] d) The magnitude of the motion vector of the neighboring block

[0297] e) Alternatively, signaling T can be sent from the encoder to the decoder.

[0298] 5. The second step in enabling or disabling the prediction unit / code-decode block / region in DMVR can be determined based on messages (e.g., flags) such as those signaled in sequence (e.g., SPS) / strip (e.g., strip header) / piece group (e.g., piece group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0299] a. In one example, a signaling notification flag indicates whether the second step in the DMVR is applied.

[0300] i. In one example, when such a flag

[0301] When the second step in the DMVR is disabled (e.g., tile_group_subpixel_refinement_enabled_flag), the fraction refinement process in the DMVR is skipped.

[0302] b. In one example, a flag indicating whether signaling notification for the second step in the DMVR is enabled or disabled may also indicate information unrelated to the second step in the DMVR.

[0303] i. In one example, a flag (e.g., tile_group_fpel_mmvd_enabled_flag) that indicates whether fractional motion vector difference (MVD) is allowed in Merge with motion vector difference (MMVD, also known as UMVE) can also indicate whether to apply the second step in DMVR.

[0304] ii. In one example, if the tile_group_fpel_mmvd_enabled_flag indicates that fractional MVD is disabled, the second step in DMVR is skipped.

[0305] iii. In one example, a flag (e.g., tile_group_fpel_mmvd_enabled_flag) that indicates whether fractional motion vector difference (MVD) is allowed in Merge with motion vector difference (MMVD, also known as UMVE) can also indicate whether to apply the second step in DMVR.

[0306] 6. Whether to enable or disable the second step in DMVR can be determined based on the result of integer motion refinement in DMVR for a prediction unit / coding block / region.

[0307] c. In one example, when the initial motion vector does not change after integer motion refinement in DMVR, the use of the second step in DMVR can be disabled.

[0308] d. In one example, the magnitude of a fractional motion vector can be allowed to be greater than half a pixel, where the magnitude of the fractional motion vector can be constrained to be less than or equal to T pixels.

[0309] i. In one example, T >= 0 and T is a floating-point number. (e.g., T = 1.5).

[0310] ii. In one example, the definition of "the magnitude of a motion vector is less than T" is mv.x < T and / or mv.y < T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0311] iii. In one example, the definition of "the magnitude of a motion vector is less than T" is that the sum of mv.x and mv.y is less than T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0312] iv. Additionally, alternatively, when the magnitude of the fractional motion vector is less than T, the magnitude of the fractional motion vector can be allowed to be greater than half a pixel. T can be based on:

[0313] a) The current block size

[0314] b) Current quantization parameter

[0315] c) Magnitude of the initial motion vector

[0316] d) Magnitude of the motion vector of neighboring blocks

[0317] e) Alternatively, T can be signaled from the encoder to the decoder.

[0318] e. In one example, when the distance between the initial motion vector and the motion vector obtained in the integer motion refinement in DMVR is less than T, the second step in DMVR can be disabled.

[0319] i. In one example, the definition of "the magnitude of a motion vector is less than T" is mv.x < T and / or mv.y < T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0320] ii. In one example, the definition of "the magnitude of a motion vector is less than T" is that the sum of mv.x and mv.y is less than T, where respectively, mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector.

[0321] iii. Additionally, alternatively, when the distance between the initial motion vector and the motion vector obtained in the integer motion refinement in DMVR is less than T, the second step in DMVR can be disabled. T can be based on:

[0322] a) Current block size

[0323] b) Current quantization parameter

[0324] c) Magnitude of the initial motion vector

[0325] d) Magnitude of the motion vector of neighboring blocks

[0326] e) Alternatively, T can be signaled from the encoder to the decoder.

[0327] 7. It is possible to determine whether to enable or disable the second step in DMVR for a prediction unit / coding block / region based on the initial motion vector in DMVR.

[0328] a. In one example, when both initial motion vectors are integer motion vectors, the use of the second step in DMVR can be disabled.

[0329] b. In one example, when the reference picture does not point to certain pictures, such as when the reference picture index is not equal to 0, the second step in DMVR is always disabled.

[0330] c. In one example, when the magnitude of the initial motion vector is greater than T, it can be disabled.

[0331] The use of the second step in DMVR.

[0332] i. In one example, "the magnitude of a motion vector is greater than T" is defined as mv.x>T and / or mv.y>T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0333] ii. In one example, "the magnitude of a motion vector is greater than T" is defined as the sum of mv.x and mv.y being greater than T, where mv.x is the magnitude of the horizontal component of the motion vector and mv.y is the magnitude of the vertical component of the current motion vector, respectively.

[0334] iii. Alternatively, the second step in the DMVR can be disabled when the magnitude of the initial motion vector is greater than T and / or the initial motion vector is an integer motion vector. T can be based on:

[0335] a) Current block size

[0336] b) Current quantization parameters

[0337] c) The magnitude of the initial motion vector

[0338] d) The magnitude of the motion vector of the neighboring block

[0339] e) Alternatively, signaling T can be sent from the encoder to the decoder.

[0340] 8. The accuracy of the initial motion vector in the DMVR can be determined based on signaling notification messages such as those in sequence (e.g., SPS) / strip (e.g., strip header) / piece group (e.g., piece group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0341] f. In one example, a signaling notification flag is used to indicate whether the initial motion vector in the DMVR is rounded to an integer motion vector.

[0342] g. In one example, a signaling notification flag indicating whether the initial motion vector in the DMVR is rounded to an integer motion vector can also indicate information unrelated to the initial motion vector in the DMVR.

[0343] i. In one example, a flag indicating whether fractional motion vector difference (MVD) is allowed in a Merge (MMVD, aka UMVE) with motion vector difference (e.g., tile_group_fpel_mmvd_enabled_flag) can also indicate whether the initial motion vector in the BIO and / or DMVR is rounded to an integer motion vector.

[0344] 9. The accuracy of the initial motion vector in BIO can depend on factors such as the sequence (e.g., SPS) /

[0345] Signaling notification messages in strip (e.g., strip header) / slice group (e.g., slice group header) / picture level (e.g., picture header) / block level (e.g., CTU or CU).

[0346] h. In one example, a signaling notification flag indicates whether the initial motion vector in the BIO is rounded to an integer motion vector.

[0347] i. In one example, a signaling notification flag indicating whether the initial motion vector in the BIO is rounded to an integer motion vector can also indicate information unrelated to the initial motion vector in the BIO.

[0348] i. In one example, a flag indicating whether fractional motion vector difference (MVD) is allowed in a Merge (MMVD, aka UMVE) with motion vector difference (e.g., tile_group_fpel_mmvd_enabled_flag) can also indicate whether the initial motion vector in the BIO and / or DMVR is rounded to an integer motion vector.

[0349] 10. The accuracy of which MVD candidates (such as those used in the first and / or second steps) are adaptively changed based on the initial motion vectors can be changed.

[0350] j. In one example, when the initialized motion vector is a subpixel motion vector, the check on the temporary motion vector can be skipped if the temporary motion vector (derived from the initialized motion vector and an MVD candidate) is an integer pixel motion vector.

[0351] k. Optionally, when the initialized motion vector is a subpixel motion vector, the check of the temporary motion vector can be skipped if the temporary motion vector (derived from the initialized motion vector and an MVD candidate) is a subpixel motion vector.

[0352] 5. Additional Examples

[0353] Text changes in the VVC draft are shown in the table below in bold italic with underline.

[0354] 7.3.2. Raw Byte Sequence Payload, End Bit, and Byte Alignment Syntax

[0355] 7.3.2.1. Sequence Parameter Set (RBSP) Syntax

[0356]

[0357]

[0358] seq_parameter_set_rbsp(){ descriptor sps_seq_parameter_set_id ue(v) … sps_ref_wraparound_enabled_flag u(1) if(sps_ref_wraparound_enabled_flag) sps_ref_wraparound_offset_minus1 ue(v) sps_temporal_mvp_enabled_flag u(1) if(sps_temporal_mvp_enabled_flag) sps_sbtmvp_enabled_flag u(1) sps_amvr_enabled_flag u(1) sps_bdof_enabled_flag u(1) sps_affine_amvr_enabled_flag u(1) sps_dmvr_enabled_flag u(1) sps_subpixel_refinement_enabled_flag u(1) sps_cclm_enabled_flag u(1) if(sps_cclm_enabled_flag&&chroma_format_idc===1) sps_cclm_colocated_chroma_flag u(1) sps_mts_enabled_flag u(1) sps_num_ladf_intervals_minus2 u(2) … }

[0359]

[0360]

[0361] A `sps_subpixel_refinement_enabled_flag` value of 1 indicates that adaptive subpixel MVD thinning can be used in-chip. `sps_fpel_mmvd_enabled_flag` indicates that adaptive subpixel MVD thinning cannot be used in-chip.

[0362] 7.3.4. Fragment Header Syntax

[0363] 7.3.4.1. General Fragment Header Syntax

[0364]

[0365]

[0366]

[0367]

[0368]

[0369]

[0370]

[0371] A `tile_group_subpixel_refinement_enabled_flag` value of 1 indicates that subpixel MVD thinning can be enabled in the current tile group. A `tile_group_fpel_mmvd_enabled_flag` value of 0 indicates that subpixel MVD thinning cannot be enabled in the current tile group. When it does not exist, the value of `tile_group_fpel_mmvd_enabled_flag` is inferred to be 1.

[0372] 8.5.3. Decoder-side motion vector refinement process

[0373] 8.5.3.1. General

[0374] The input to this process is:

[0375] – Luminance position (xSb, ySb), specifies the top-left luminance sample of the current codec sub-block relative to the top-left luminance sample of the current image.

[0376] – The variable sbWidth specifies the width of the current codec sub-block in the luminance sample.

[0377] – The variable sbHeight specifies the height of the current codec sub-block in the luminance sample.

[0378] The brightness motion vectors mvL0 and mvL1 in a fractional sample precision of -1 / 16

[0379] – The selected brightness reference image sample arrays refPicL0L and refPicL1L.

[0380] The output of this process is:

[0381] – Incremental (delta) brightness motion vectors dMvL0 and dMvL1.

[0382] The variable subPelFlag is set to 0. Furthermore, the variables srRange, offsetH0, offsetH1, offsetV0, and offsetV1 are all set to 2.

[0383] Both components of the incremental brightness motion vectors dMvL0 and dMvL1 are set to zero and modified as follows:

[0384] – For each X that is 0 or 1, the prediction block width is set to (sbWidth + 2 * srRange) at the brightness position (xSb, ySb), the prediction block height is set to (sbHeight + 2 * srRange), and the reference image sample array refPicLX is used. L The motion vector mvLX and the refined search range srRange are taken as input, and the fractional sample bilinear interpolation procedure specified in 8.5.3.2.1 is invoked to derive the array predSamplesLX of the predicted brightness sample values ​​(sbWidth+2*srRange)x(sbHeight+2*srRange). L .

[0385] – via sbWidth, sbHeight, offsetH0, offsetH1, offsetV0, offsetV1, predSamplesL0 L and predSamplesL1 L As input, the sum of absolute differences calculation procedure specified in 8.5.3.3 is invoked to derive the list sadList[i] (where i = 0..8).

[0386] – When sadList[4] is greater than or equal to sbHeight*sbWidth, the following applies:

[0387] – The variable bestIdx is derived by calling the array entry selection procedure specified in Section 8.5.3.4 with a list sadList[i] (where i = 0..8) as input.

[0388] – If bestIdx equals 4, then subPelFlag is set to equal 1.

[0389] –Otherwise, the following applies:

[0390] dX = bestIdx%3-1

[0391] dY = bestIdx / 3-1

[0392] dMvL0[0]+=16*dX

[0393] dMvL0[1]+=16*dY

[0394] offsetH0+=dX

[0395] offsetV0+=dY

[0396] offsetH1 - = dX

[0397] offsetV1-=dY

[0398] – By using sbWidth, sbHeight, offsetH0, offsetH1, offsetV0,

[0399] offsetV1, predSamplesL0 L and predSamplesL1 L As input, the sum of absolute differences calculation procedure specified in 8.5.3.3 is invoked to modify the list sadList[i] (where i = 0..8).

[0400] – Modify the variable bestIdx by invoking the array entry selection procedure specified in Section 8.5.3.4 with a list sadList[i] (where i = 0..8) as input.

[0401] – If bestIdx equals 4, then subPelFlag is set to equal 1.

[0402] – Otherwise (bestIdx is not equal to 4), the following applies:

[0403] dMvL0[0]+=16*(bestIdx%3-1)

[0404] dMvL0[1]+=16*(bestIdx / 3-1)

[0405] – If tile_group_subpixel_refinement_enabled_flag equals 1

[0406] – When subPelFlag equals 1, take the list sadList[i] (where i = 0..8) and the incremental motion vector dMvL0 as input and the modified dMvL0 as output, and call the parameterized motion vector refinement process specified in Section 8.5.3.5.

[0407] –The incremental motion vector dMvL1 is derived as follows:

[0408] dMvL1[0]=-dMvL0[0]

[0409] dMvL1[1]=-dMvL0[1]

[0410] 9. Example implementations of the disclosed technology

[0411] Figure 24 This is a block diagram of a video processing apparatus 2400. Apparatus 2400 can be used to implement one or more of the methods described herein. Apparatus 2400 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 2400 may include one or more processors 2402, one or more memories 2404, and video processing hardware 2406. The processors (multiple) 2402 can be configured to implement one or more methods described in this document. The memories (multiple) 2404 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2406 can be used to implement some of the techniques described in this document in hardware circuitry and can be partly or entirely part of the processor 2402 (e.g., a graphics processing unit (GPU) core or other signal processing circuitry).

[0412] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation of the current video block can, for example, correspond to bits juxtaposed or scattered at different locations within the bitstream, as defined in the syntax. For example, a macroblock can be encoded according to the transform and encoding / decoding error residual values ​​and also using bits from other fields in the header and bitstream.

[0413] It should be understood that by allowing the use of the techniques disclosed in this document, the disclosed methods and techniques will be beneficial to video encoder and / or decoder embodiments incorporated into video processing devices such as smartphones, laptops, desktop computers and similar devices.

[0414] Figure 25 This is a flowchart of an example method 2500 for video processing. Method 2500 includes, at 2510, performing a conversion between a current video block and a bitstream representation of the current video block, wherein the use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion, and a second value of the flag indicates that decoder motion information is disabled during the conversion.

[0415] Some embodiments can be described using the following clause-based format.

[0416] 1. A method for visual media processing, comprising:

[0417] Perform a conversion between the current video block and its bitstream representation, wherein the use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion, and a second value of the flag indicates that decoder motion information is disabled during the conversion.

[0418] 2. The method according to Clause 1, wherein the decoder motion information includes decoder motion video refinement (DMVR) or bidirectional optical flow (BIO).

[0419] 3. The method according to any one of Clauses 1-2, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0420] 4. The method according to any one of clauses 1-3, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in Merge (MMVD) mode with motion vector difference.

[0421] 5. The method according to any one of clauses 1-3, wherein the flag indicates whether fractional motion vector difference (MVD) is not allowed in Merge (MMVD) mode with motion vector difference.

[0422] 6. The method according to Clause 2, wherein the decoder motion information derivation includes at least one of the following: one or more initial motion vectors associated with a video block, or one or more reference pictures associated with a video block, wherein the one or more initial motion vectors are decoded motion vectors prior to the use of DMVR or BIO.

[0423] 7. The method according to Clause 6, wherein one or more initial motion vectors associated with the video block include two initial motion vectors, and further includes:

[0424] In response to the determination that both initial motion vectors are integer values, skip DMVR.

[0425] 8. The method described in Clause 6 further includes:

[0426] In response to determining that the index of one or more reference images associated with a video block is non-zero, skip DMVR.

[0427] 9. The method described in Clause 6 further includes:

[0428] In response to determining that the magnitude of one or more initial motion vectors associated with a video block is greater than a threshold, DMVR is skipped.

[0429] 10. The method according to Clause 9, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in one or more initial motion vectors, wherein mv.x>T and / or mv.y>T, and wherein T represents a threshold.

[0430] 11. The method according to Clause 9, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in one or more initial motion vectors, wherein sum(mv.x,mv.y)>T, wherein T represents a threshold, and wherein sum(x,y)=x+y.

[0431] 12. The method according to any one of Clauses 10-11, wherein the motion vector is the motion vector of the current video block.

[0432] 13. The method according to any one of Clauses 10-12, wherein the threshold is based on one or more of the following: the size of the video block, a quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of motion vectors associated with blocks adjacent to the video block, or a value signaled from the visual media processing encoder to the visual media processing decoder.

[0433] 14. The method according to Clause 6, wherein the one or more initial motion vectors associated with the video block include two initial motion vectors, and further includes:

[0434] In response to the determination that both initial motion vectors are integer values, skip the BIO.

[0435] 15. The method according to Clause 6, wherein one or more initial motion vectors associated with the video block include the initial motion vector and a refinement of the initial motion vector, further comprising:

[0436] Based on the difference between the initial motion vector and the refinement, enable or skip BIO.

[0437] 16. The method described under Clause 15 further includes:

[0438] In response to determining the difference between the initial motion vector and the thinning, which is a sub-pixel, predictions associated with video blocks are skipped.

[0439] 17. The method described in Clause 6 further includes:

[0440] In response to determining that the index of one or more reference images associated with a video block is non-zero, skip the BIO.

[0441] 18. The method described under Clause 6 further includes:

[0442] In response to determining that the magnitude of one or more initial motion vectors associated with a video block is greater than a threshold, skip the BIO.

[0443] 19. The method according to Clause 18, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in one or more initial motion vectors, wherein mv.x>T and / or mv.y>T, and wherein T represents a threshold.

[0444] 20. The method according to Clause 18, wherein the magnitude of one or more initial motion vectors associated with the video block includes the magnitude of the horizontal component of the motion vector (denoted as mv.x) and the magnitude of the vertical component of the motion vector (denoted as mv.y), wherein the motion vector is included in one or more initial motion vectors, wherein sum(mv.x,mv.y)>T, wherein T represents a threshold, and wherein sum(x,y)=x+y.

[0445] 21. The method according to any one of Clauses 19-20, wherein the motion vector is the motion vector of the current video block.

[0446] 22. The method according to any one of Clauses 19-20, wherein the threshold is based on one or more of the following: the size of the video block, a quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of motion vectors associated with blocks adjacent to the video block, or a value signaled from the visual media processing encoder to the visual media processing decoder.

[0447] 23. The method according to any one of Clauses 1-2, wherein the DMVR includes a second step of fractional refinement, wherein the flag indicates whether fractional refinement is skipped.

[0448] 24. The method described in Clause 23, wherein the flag is represented as tile_group_subpixel_refinement_enabled_flag.

[0449] 25. The method according to Clause 23, wherein the flag indicates information that is not relevant to the second step of score refinement.

[0450] 26. The method according to Clause 25, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in Merge (MMVD) mode with motion vector difference.

[0451] 27. The method described in accordance with Clause 26, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0452] 28. The method according to Clause 25, wherein if the flag indicates whether fractional motion vector difference (MVD) is disabled, the second step of fractional refinement is skipped.

[0453] 29. The method described in accordance with Clause 26, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0454] 30. The method according to any one of clauses 1-2, wherein the DMVR includes a second step of fractional refinement based on the result of integer motion refinement.

[0455] 31. The method according to clause 30, further comprising:

[0456] In response to determining that the initial motion vector of the video block has not changed after integer motion refinement, skipping the second step of DMVR.

[0457] 32. The method according to clause 31, wherein the magnitude of the fractional motion vector is greater than half a pixel, and wherein the magnitude of the fractional motion vector is constrained to be less than or equal to a threshold.

[0458] 33. The method according to clause 31, wherein the threshold is a non-zero floating-point number.

[0459] 34. The method according to clause 32, wherein the magnitude of the fractional motion vector associated with the video block includes the magnitude of the horizontal component of the fractional motion vector (denoted as mv.x) and the magnitude of the vertical component of the fractional motion vector (denoted as mv.y), and wherein mv.x < T and / or mv.y < T, and wherein T represents the threshold.

[0460] 35. The method according to clause 32, wherein the magnitude of the fractional motion vector associated with the video block includes the magnitude of the horizontal component of the fractional motion vector (denoted as mv.x) and the magnitude of the vertical component of the fractional motion vector (denoted as mv.y), and wherein sum(mv.x, mv.y) < T, where T represents the threshold, and wherein sum(x, y) = x + y.

[0461] 36. The method according to any one of clauses 32-35, wherein the threshold is based on one or more of the following: the size of the video block, the quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of the motion vectors associated with blocks adjacent to the video block, or a value signaled from a visual media processing encoder to a visual media processing decoder.

[0462] 37. The method according to clause 30, wherein the difference vector is calculated as the difference between the initial motion vector of the video block and the motion vector obtained after integer motion refinement, further comprising:

[0463] In response to determining that the magnitude of the difference vector is less than the threshold, skipping the second step of DMVR.

[0464] 38. The method according to clause 37, wherein the magnitude of the difference vector associated with the video block includes the magnitude of the horizontal component of the difference vector (denoted as mv.x) and the magnitude of the vertical component of the difference vector (denoted as mv.y), and wherein mv.x < T and / or mv.y < T, and wherein T represents a threshold value.

[0465] 39. The method according to clause 37, wherein the magnitude of the difference vector associated with the video block includes the magnitude of the horizontal component of the difference vector (denoted as mv.x) and the magnitude of the vertical component of the difference vector (denoted as mv.y), and wherein sum(mv.x, mv.y) < T, where T represents a threshold value, and wherein sum(x, y) = x + y.

[0466] 40. The method according to any one of clauses 37 - 39, wherein the threshold value is based on one or more of the following: the size of the video block, the quantization parameter, the magnitude of one or more initial motion vectors associated with the video block, the magnitude of the motion vectors associated with blocks adjacent to the video block, or a value signaled from a visual media processing encoder to a visual media processing decoder.

[0467] 41. The method according to any one of clauses 1 - 2, wherein the flag indicates the precision of the initial motion vector associated with the video block.

[0468] 42. The method according to clause 41, wherein the flag indicates whether the initial motion vector is rounded to an integer value.

[0469] 43. The method according to clause 42, wherein the flag indicates information not related to the initial motion vector.

[0470] 44. The method according to clause 43, wherein the flag indicates whether fractional motion vector difference (MVD) is allowed in the Merge with Motion Vector Difference (MMVD) mode.

[0471] 45. The method according to clause 44, wherein the flag is represented as tile_group_fpel_mmvd_enabled_flag.

[0472] 46. A method for visual media processing, comprising:

[0473] Performing a conversion between a current video block and a bitstream representation of the current video block, wherein the use of decoder motion information is indicated in a flag in the bitstream representation, such that a first value of the flag indicates that decoder motion information is enabled during the conversion, and a second value of the flag indicates that decoder motion information is disabled during the conversion; and

[0474] In response to determining that the initial motion vector of the video block has subpixel precision, checks associated with temporary motion vectors derived from the difference between the initial motion vector and candidate motion vectors are skipped, where the temporary motion vectors have integer pixel precision or subpixel precision.

[0475] 47. The method according to any one or more of clauses 1-46, wherein the flag is included in the Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture Header, Slice Header, Strip Header, Codec Tree Unit Row, or a region associated with a Codec Tree Unit.

[0476] 48. The method according to any one or more of clauses 1 to 47, wherein visual media processing is an encoder-side implementation.

[0477] 49. The method according to any one or more of clauses 1 to 47, wherein visual media processing is a decoder-side implementation.

[0478] 50. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one or more of claims 1 to 49.

[0479] 51. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the methods described in any one or more of clauses 1 to 49.

[0480] Figure 26 This is a flowchart of an example method 2600 for video processing. Method 2600 includes, at 2602, determining whether and / or how to apply decoder-side motion vector refinement (DMVR) for the conversion between the first block of video and the bitstream representation of the first block of video, based on information from signaling notification; and at 2604, performing the conversion based on the determination.

[0481] In some examples, the signaling notification information includes flags present in the bitstream representation.

[0482] In some examples, the signaling notification information is signaled at at least one of the following: sequence level including sequence parameter set (SPS), strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

[0483] In some examples, during this conversion, DMVR is applied to at least one of the prediction unit, codec block, and region.

[0484] In some examples, flags are signaled to indicate whether the DMVR is enabled or disabled.

[0485] In some examples, when the flag indicates that DMVR is disabled, the DMVR process is skipped, and the use of DMVR is inferred to be disabled.

[0486] In some examples, information indicating whether and / or how to apply DMVR signaling notifications is used to control the use of other decoder motion information derivation techniques.

[0487] In some examples, other decoder motion information derivation techniques include bidirectional optical flow (BIO).

[0488] In some examples, a flag indicating whether fractional motion vector difference (MVD) is allowed in Merge (MMVD) mode with motion vector difference is used to determine whether and / or how DMVR is applied.

[0489] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0490] In some examples, if the flag indicates that fractional MVD is not allowed for MMVD, the DMVR process is skipped during the conversion.

[0491] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0492] In some examples, this conversion generates the first block of the video from the bitstream representation.

[0493] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0494] Figure 27 This is a flowchart of an example method 2700 for video processing. Method 2700 includes, at 2702, determining whether and / or how to apply bidirectional optical flow (BIO) for the conversion between a first block of video and a bitstream representation of the first block of video, based on information from signaling notification; and at 2704, performing the conversion based on the determination.

[0495] In some examples, the signaling notification information includes flags present in the bitstream representation.

[0496] In some examples, the signaling notification information is signaled at at least one of the following: sequence level including sequence parameter set (SPS), strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

[0497] In some examples, during this conversion, BIO is applied to at least one of the prediction unit, codec block, and region.

[0498] In some examples, flags are signaled to indicate whether BIO is enabled or disabled.

[0499] In some examples, when the flag indicates that BIO is disabled, the BIO process is skipped, and the use of BIO is inferred to be disabled.

[0500] In some examples, information indicating whether and / or how to apply BIO signaling notifications is used to control the use of other decoder motion information derivation techniques.

[0501] In some examples, other decoder motion information derivation techniques include decoder motion video refinement (DMVR).

[0502] In some examples, a flag indicating whether fractional motion vector difference (MVD) is allowed in Merge (MMVD) mode with motion vector difference is used to determine whether and / or how BIO is applied.

[0503] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0504] In some examples, if the flag indicates that fractional MVD is not allowed for MMVD, the BIO process is skipped during the transformation.

[0505] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0506] In some examples, this conversion generates the first block of the video from the bitstream representation.

[0507] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0508] Figure 28 This is a flowchart of an example method 2800 for video processing. Method 2800 includes, at 2802, determining whether a decoder-side motion vector refinement (DMVR) process is enabled or disabled for the conversion between a first block of video and a bitstream representation of the first block, based on at least one of the following: one or more initial motion vectors associated with the first block and one or more reference pictures associated with the first block, wherein the initial motion vectors include motion vectors prior to the application of the DMVR process; and at 2804, performing the conversion based on this determination.

[0509] In some examples, during the conversion, the DMVR process is applied to at least one of the prediction unit, codec block, and region associated with the first block.

[0510] In some examples, the initial motion vector includes a first initial motion vector in a first direction and / or a second initial motion vector in a second direction.

[0511] In some examples, the DMVR process is disabled when both initial motion vectors are integer motion vectors.

[0512] In some examples, the DMVR process is disabled when the reference image does not point to certain images.

[0513] In some examples, the DMVR process is disabled when the reference image index is not equal to 0.

[0514] In some examples, the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a threshold (T).

[0515] In some examples, the magnitude of each initial motion vector includes the magnitude of the horizontal component (mv.x) and the magnitude of the vertical component (mv.y).

[0516] In some examples, the DMVR process is disabled when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector are greater than a threshold T.

[0517] In some examples, the DMVR process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than a threshold T.

[0518] In some examples, the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a threshold (T) and / or the initial motion vectors are integer motion vectors.

[0519] In some examples, T is determined based on at least one of the following: the size of the first block, the current quantization parameter, the magnitude of one or more initial motion vectors, and the magnitude of the motion vectors of the neighboring blocks of the first block.

[0520] In some examples, T is the signaling notification from the encoder to the decoder.

[0521] In some examples, the DMVR process includes a first step in which integer motion vector difference (MVD) refinement with integer precision is performed and a second step in which fractional MVD refinement with fractional precision is performed.

[0522] In some examples, the method also includes determining whether the second step of the DMVR process is enabled or disabled based on information from signaling notifications.

[0523] In some examples, the signaling notification information includes flags present in the bitstream representation.

[0524] In some examples, the signaling notification information is signaled at at least one of the following: sequence level including sequence parameter set (SPS), strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

[0525] In some examples, a flag is signaled to indicate whether the second step of the DMVR process is enabled or disabled.

[0526] In some examples, this flag is tile_group_subpixel_refinement_enabled_flag.

[0527] In some examples, the fractional refinement process in the DMVR procedure is skipped when a flag indicates that the second step of the DMVR procedure is disabled.

[0528] In some examples, the flag indicating whether the second step in the DMVR process is enabled or disabled also indicates additional information beyond the second step in the DMVR process.

[0529] In some examples, a flag indicating whether fractional MVD is allowed in Merge (MMVD) mode with motion vector difference is used to indicate whether the second step of the DMVR process is enabled or disabled.

[0530] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0531] In some examples, the method also includes determining how to perform the second step of the DMVR process based on the results of integer motion refinement in the first step of the DMVR process.

[0532] In some examples, the second step in the DMVR process is disabled when the initial motion vector is not changed after integer motion refinement in the first step of the DMVR process.

[0533] In some examples, the magnitude of the motion vector obtained after the second step is allowed to be greater than half a pixel and less than or equal to the second threshold T2 pixels.

[0534] In some examples, T2 >= 0, and T2 is a floating-point number.

[0535] In some examples, the magnitude of the motion vector includes the magnitude of the horizontal component (mv.x) and the magnitude of the vertical component (mv.y).

[0536] In some examples, one or more of the horizontal component mv.x and the vertical component mv.y of the motion vector are constrained to be less than T2 pixels.

[0537] In some examples, the sum of the horizontal component mv.x and the vertical component mv.y of the motion vector is constrained to be less than T2 pixels.

[0538] In some examples, the second step in the DMVR process is disabled when the distance between the initial motion vector and the motion vector obtained after integer motion refinement in the first step of the DMVR process is less than a second threshold T2.

[0539] In some examples, the distance includes the magnitude of the horizontal component (mv.x) and the magnitude of the vertical component (mv.y).

[0540] In some examples, the second step in the DMVR process is disabled when one or more of the horizontal component mv.x and the vertical component mv.y of the distance are less than the threshold T2.

[0541] In some examples, the second step in the DMVR process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y is less than the threshold T2.

[0542] In some examples, T2 is determined based on at least one of the following: the size of the first block, the current quantization parameter, the magnitude of one or more initial motion vectors, and the magnitude of the motion vectors of the neighboring blocks of the first block.

[0543] In some examples, T2 is a signaling notification from the encoder to the decoder.

[0544] In some examples, the method also includes determining whether the second step of the DMVR process is enabled or disabled based on one or more initial motion vectors in the DMVR process.

[0545] In some examples, the second step of the DMVR process is disabled when both initial motion vectors are integer motion vectors.

[0546] In some examples, the second step of the DMVR process is disabled when the reference image does not point to certain images.

[0547] In some examples, the second step of the DMVR process is disabled when the reference image index is not equal to 0.

[0548] In some examples, the second step of the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than a third threshold (T3).

[0549] In some examples, the magnitude of each initial motion vector includes the magnitude of the horizontal component (mv.x) and the magnitude of the vertical component (mv.y).

[0550] In some examples, the second step of the DMVR process is disabled when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector are greater than the threshold T3.

[0551] In some examples, the second step of the DMVR process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than the threshold T3.

[0552] In some examples, the second step of the DMVR process is disabled when the magnitude of one or more initial motion vectors is greater than the third threshold (T3) and / or the initial motion vectors are integer motion vectors.

[0553] In some examples, T3 is determined based on at least one of the following: the size of the first block, the current quantization parameter, the magnitude of one or more initial motion vectors, and the magnitude of the motion vectors of the blocks adjacent to the first block.

[0554] In some examples, T3 is a signaling notification from the encoder to the decoder.

[0555] In some examples, the accuracy of the initial motion vector during the DMVR process is determined based on information from signaling notifications.

[0556] In some examples, the signaling notification information includes flags present in the bitstream representation.

[0557] In some examples, the signaling notification information is signaled at at least one of the following: sequence level including sequence parameter set (SPS), strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

[0558] In some examples, a flag is signaled to indicate whether the initial motion vector during the DMVR process is rounded to an integer motion vector.

[0559] In some examples, the flag indicating whether the initial motion vector during the DMVR process is rounded to an integer motion vector also indicates additional information beyond the initial motion vector during the DMVR process.

[0560] In some examples, a flag indicating whether fractional MVD is allowed in Merge (MMVD) mode with motion vector difference is used to indicate whether the initial motion vector in the DMVR process is rounded to an integer motion vector.

[0561] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0562] In some examples, whether to use MVD candidates is adaptively changed based on the accuracy of the initial motion vectors.

[0563] In some examples, MVD candidates include those used in the first and / or second steps of the DMVR process.

[0564] In some examples, when the initial motion vector is a subpixel motion vector, if the temporary motion vector is an integer pixel motion vector, the check on the temporary motion vector is skipped. The temporary motion vector is derived from the initial motion vector and an MVD candidate.

[0565] In some examples, when the initial motion vector is a subpixel motion vector, if the temporary motion vector is a subpixel motion vector, the check on the temporary motion vector is skipped. The temporary motion vector is derived from the initial motion vector and an MVD candidate.

[0566] Figure 29 This is a flowchart of an example method 2900 for video processing. Method 2900 includes, at 2902, determining whether a bidirectional optical flow (BIO) process is enabled or disabled for the conversion between a first block of video and a bitstream representation of the first block of video, based on at least one of the following: one or more initial motion vectors associated with the first block and one or more reference images associated with the first block, wherein the initial motion vectors include motion vectors prior to the application of the BIO process and / or decoded motion vectors prior to the application of a decoder-side motion vector refinement (DMVR) process; and at 2904, performing the conversion based on this determination.

[0567] In some examples, during the conversion, the BIO process is applied to at least one of the prediction unit, codec block, and region associated with the first block.

[0568] In some examples, the initial motion vector includes a first initial motion vector in a first direction and / or a second initial motion vector in a second direction.

[0569] In some examples, the BIO process is disabled when both initial motion vectors are integer motion vectors.

[0570] In some examples, the BIO process is always disabled when the reference image does not point to certain images.

[0571] In some examples, the BIO process is disabled when the reference image index is not equal to 0.

[0572] In some examples, the BIO process is disabled when the magnitude of one or more initial motion vectors is greater than the fourth threshold (T4).

[0573] In some examples, the magnitude of each initial motion vector includes the magnitude of the horizontal component (mv.x) and the magnitude of the vertical component (mv.y).

[0574] In some examples, the BIO process is disabled when one or more of the horizontal component mv.x and the vertical component mv.y of the initial motion vector are greater than a threshold T.

[0575] In some examples, the BIO process is disabled when the sum of the horizontal component mv.x and the vertical component mv.y of the initial motion vector is greater than the threshold T4.

[0576] In some examples, the BIO process is disabled when the magnitude of one or more initial motion vectors is greater than the fourth threshold (T4) and / or the initial motion vectors are integer motion vectors.

[0577] In some examples, T4 is determined based on at least one of the following: the size of the first block, the current quantization parameter, the magnitude of one or more initial motion vectors, and the magnitude of the motion vectors of the blocks adjacent to the first block.

[0578] In some examples, T4 is a signaling notification from the encoder to the decoder.

[0579] In some examples, the accuracy of the initial motion vector in the BIO process is determined based on information from signaling notifications.

[0580] In some examples, the signaling notification information includes flags present in the bitstream representation.

[0581] In some examples, the signaling notification information is signaled at at least one of the following: sequence level including sequence parameter set (SPS), strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

[0582] In some examples, a flag is signaled to indicate whether the initial motion vector in the BIO process is rounded to an integer motion vector.

[0583] In some examples, the flag indicating whether the initial motion vector in the BIO process is rounded to an integer motion vector also indicates additional information beyond the initial motion vector in the BIO process.

[0584] In some examples, a flag indicating whether fractional MVD is allowed in Merge (MMVD) mode with motion vector difference is used to indicate whether the initial motion vector in the BIO process is rounded to an integer motion vector.

[0585] In some examples, this flag is tile_group_fpel_mmvd_enabled_flag.

[0586] Figure 30 This is a flowchart of an example method 3000 for video processing. Method 3000 includes, at 3002, determining whether a bidirectional optical flow (BIO) process is enabled or disabled for the conversion between the first block of video and a bitstream representation of the first block of video, based on one or more motion vector differences between an initial motion vector and one or more refined motion vectors associated with the first block of video, wherein the initial motion vector includes motion vectors before the application of the BIO process and / or the application of the DMVR process, and the refined motion vectors include motion vectors after the application of the DMVR process; and at 3004, performing the conversion based on this determination.

[0587] In some examples, during the conversion, the BIO process is applied to at least one of the prediction unit, codec block, and region associated with the first block.

[0588] In some examples, when the motion vector difference is a sub-pixel, updates to the predicted or reconstructed samples are skipped.

[0589] In some examples, this conversion generates the first block of the video from the bitstream representation.

[0590] In some examples, the conversion generates a bitstream representation from the first block of the video.

[0591] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their structural equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances influencing machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.

[0592] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.

[0593] The processes and logic described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0594] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0595] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0596] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0597] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.

Claims

1. A video processing method, comprising: Based on the information from the signaling notification, determine whether and / or how to apply bidirectional optical flow (BIO) for the conversion between the first block of the video and the bitstream of the first block of the video. as well as The conversion is performed based on the determination. If the signaling notification indicates that fractional MVD is not allowed for MMVD, then the BIO process is skipped during the conversion.

2. The method according to claim 1, wherein, The signaling notification information includes flags present in the bitstream.

3. The method according to claim 1 or 2, wherein, The signaling notification information is signaled at least one of the following: sequence level including sequence parameter set SPS, strip level including strip header, slice group level including slice group header, picture level including picture header, and block level including codec tree unit and codec unit.

4. The method according to claim 1 or 2, wherein, During the conversion, the BIO is applied to at least one of the prediction unit, codec block, and region.

5. The method according to claim 1 or 2, wherein, The signaling notification information is used to indicate whether BIO is enabled or disabled.

6. The method according to claim 5, wherein, When the signaling notification indicates that BIO is disabled, the BIO process is skipped, and the use of BIO is inferred to be disabled.

7. The method according to claim 1 or 2, wherein, Information indicating whether and / or how to apply the signaling notifications of BIO is used to control the use of other decoder motion information derivation techniques.

8. The method according to claim 7, wherein, Other decoder motion information derivation techniques include decoder motion video refinement (DMVR).

9. The method according to claim 1 or 2, wherein, The signaling notification information indicating whether fractional motion vector difference MVD is allowed in Merge MMVD mode with motion vector difference is used to determine whether and / or how to apply BIO.

10. The method according to claim 9, wherein, The signaling notification information is tile_group_fpel_mmvd_enabled_flag.

11. The method according to claim 1 or 2, wherein, The signaling notification information is tile_group_fpel_mmvd_enabled_flag.

12. The method according to claim 1 or 2, wherein, The conversion generates the first block of video from the bitstream.

13. The method according to claim 1 or 2, wherein, The conversion generates the bitstream from the first block of the video.

14. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 13.

15. A computer-readable medium storing code, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Method and apparatus of motion compensation based on BI-directional optical flow techniques for video coding

    WO2017133661A1