Decoder-side motion inference based on processing parameters
Through the decoder-side motion vector derivation (DMVD) scheme and symmetric motion vector difference (SMVD) mode, the encoding process of video blocks is optimized, the problem of low encoding efficiency in the existing technology is solved, the video encoding efficiency and quality is improved, and the future development of video encoding standards is adapted.
Patent Information
- Application Number
- CN202080000731.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-13
- Filing Date
- 2020-02-14
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-02-14
AI Technical Summary
The existing video encoding technology is low in encoding efficiency when processing motion vector expression and generalized bidirectional prediction, and it is difficult to meet the growing demand for digital video bandwidth.
The decoder-side motion vector derivation (DMVD) scheme is adopted to optimize the encoding process of video blocks by applying weight parameters and refining motion information, and enable or disable the DMVD scheme according to the specific conditions of the current video block, combining the symmetric motion vector difference (SMVD) mode and multiple DMVD schemes to improve the encoding efficiency.
It improves the encoding efficiency of video encoding, optimizes the encoding process of video blocks, improves the video quality and encoding efficiency, and adapts to the development needs of future video encoding standards.
Smart Images

Figure CN111837395B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is intended to claim priority and the benefit of International Patent Application No. PCT / CN2019 / 075068, filed on February 14, 2019, and International Patent Application No. PCT / CN2019 / 082585, filed on April 13, 2019, in a timely manner, under applicable patent laws and / or regulations under the Paris Convention. The entire disclosures of the aforementioned applications are incorporated by reference into this application for all purposes. Technical Field
[0003] This document deals with video and image encoding and decoding. Background Art
[0004] Digital video accounts for the largest use of bandwidth on the Internet and other digital communications networks. As the number of networked user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention
[0005] This document discloses a video coding tool that, in one example aspect, improves coding efficiency over current coding tools related to final motion vector representation or generalized bidirectional prediction.
[0006] A first example video processing method includes obtaining refined motion information for a current video block of a video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least weight parameters, wherein the weight parameters are applied to the prediction block during generation of a final prediction block for the current video block; and performing conversion between the current video block and a bitstream representation of the video using at least the refined motion information and the weight parameters.
[0007] A second example video processing method includes determining that use of a decoder-side motion vector derivation (DMVD) scheme is disabled for conversion between the current video block and an encoded representation of the video due to use of an encoding tool for a current video block of a video, wherein the encoding tool includes applying unequal weighting factors to prediction blocks of the current video block; and performing conversion between the current video block and a bitstream representation of the video based on the determination.
[0008] A third example video processing method includes determining whether to enable or disable one or more decoder-side motion vector derivation (DMVD) schemes for a current video block based on picture order count (POC) values of one or more reference pictures of the current video block of the video and a POC value of a current picture that includes the current video block; and performing conversion between the current video block and a bitstream representation of the video based on the determination.
[0009] A fourth example video processing method includes obtaining refined motion information for a current video block of the video by implementing a decoder-side motion vector derivation (DMVD) scheme for the current video block of the video, wherein a symmetric motion vector difference (SMVD) mode is enabled for the current video block; and performing conversion between the current video block and a bitstream representation of the video using the refined motion information.
[0010] A fifth example video processing method includes: determining, based on a field in a bitstream representation of a video including the current video block, whether a decoder-side motion vector derivation (DMVD) scheme is enabled or disabled for the current video block, wherein a symmetric motion vector difference (SMVD) mode is enabled for the current video block; after determining that the DMVD scheme is enabled, obtaining refined motion information of the current video block by implementing the DMVD scheme on the current video block; and performing conversion between the current video block and the bitstream representation of the video using the refined motion information.
[0011] A sixth example video processing method includes determining, based on rules using block dimensions of a current video block of the video, whether to enable or disable multiple decoder-side motion vector derivation (DMVD) schemes for conversion between a current video block and a bitstream representation of the video; and performing conversion based on the determination.
[0012] A seventh example video processing method includes: determining whether to perform multiple decoder-side motion vector derivation (DMVD) schemes at a sub-block level or a block level of a current video block of a video; after determining to perform the multiple DMVD schemes at the sub-block level, obtaining refined motion information of the current video block by implementing the multiple DMVD schemes at the same sub-block level of the current video block; and performing conversion between the current video block and a bitstream representation of the video using the refined motion information.
[0013] An eighth example video processing method includes: determining whether a decoder-side motion vector derivation (DMVD) scheme is enabled or disabled for a plurality of components of a current video block of a video; after determining that the DMVD scheme is enabled, obtaining refined motion information of the current video block by implementing the DMVD scheme; and performing conversion between the current video block and a bitstream representation of the video during implementation of the DMVD scheme.
[0014] In another example aspect, the methods described above and in this patent document may be implemented by a video encoder device or a video decoder device that includes a processor.
[0015] In another example aspect, the methods described above and in this patent document may be stored on a non-transitory computer-readable program medium in the form of processor-executable instructions.
[0016] These and other aspects are described further throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 An example derivation process for Merge candidate list construction is shown.
[0018] Figure 2 Example locations of spatial merge candidates are shown.
[0019] Figure 3 An example of candidate pairs considered for redundancy checking of spatial merge candidates is shown.
[0020] Figure 4A-4B Example locations of the second PU for Nx2N and 2NxN partitions are shown.
[0021] Figure 5 It is a diagram of the motion vector scaling of the time domain Merge candidate.
[0022] Figure 6 An example of candidate positions of the time-domain merge candidates C0 and C1 is shown.
[0023] Figure 7 An example of a combined bi-predictive Merge candidate is shown.
[0024] Figure 8 The derivation process of motion vector prediction candidates is summarized.
[0025] Figure 9 Diagram showing motion vector scaling of spatial motion vector candidates.
[0026] Figure 10 An example of neighboring samples used to derive illumination compensation (IC) parameters is shown.
[0027] Figures 11A-11B The simplified affine motion models for the 4-parameter affine and 6-parameter affine modes are shown, respectively.
[0028] Figure 12 An example of an affine motion vector field (MVF) for each sub-block is shown.
[0029] Figures 13A-13B Examples of a 4-parameter affine model and a 6-parameter affine model are shown respectively.
[0030] Figure 14 The motion vector predictor (MVP) of AF_INTER inheriting the affine candidate is shown.
[0031] Figure 15 The MVP of AF_INTER of the constructed affine candidate is shown.
[0032] Figures 16A-16B An example of a candidate for AF_MERGE is shown.
[0033] Figure 17 An example of candidate positions for the affine merge mode is shown.
[0034] Figure 18 An example of the ultimate motion vector expression (UMVE) search process is shown.
[0035] Figure 19 An example of a UMVE search point is shown.
[0036] Figure 20 An example of decoder side motion vector refinement (DMVR) based on bilateral template matching is shown.
[0037] Figure 21 An example of a mirrored motion vector difference MVD(0,1) between list 0 and list 1 in DMVR is shown.
[0038] Figure 22 An example of MVs that can be checked in one iteration is shown.
[0039] Figure 23 An example of a hardware platform for implementing the techniques described in this document is shown.
[0040] Figures 24A-24H are eight example flow charts of example methods of video processing.
[0041] Figure 25 An example of a symmetric pattern for decoder-side motion vector derivation is shown.
[0042] Figure 26 is a block diagram illustrating an example video processing system in which the various techniques disclosed herein may be implemented.
[0043] Figure 27 is a block diagram illustrating a video encoding system according to some embodiments of the present disclosure.
[0044] Figure 28 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.
[0045] Figure 29 is a block diagram illustrating a decoder according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0046] This document provides various techniques that can be used by decoders of video bitstreams to improve the quality of decompressed or decoded digital video. In addition, video encoders can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0047] For ease of understanding, section headings are used in this document and do not limit the embodiments and techniques to the corresponding sections. Thus, embodiments from one section can be combined with embodiments from other sections.
[0048] 1. Overview
[0049] The present invention relates to video coding techniques. Specifically, it relates to the interaction between unequal weights applied to prediction blocks and motion vector refinement in video coding. The present invention can be applied to existing video coding standards, such as High Efficiency Video Coding (HEVC), or to standards to be finalized (e.g., Versatile Video Coding (VVC)). It can also be applied to future video coding standards or video codecs.
[0050] 2. Brief discussion
[0051] Video coding standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding architecture that uses temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and put them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of a 50% bitrate reduction compared to HEVC.
[0052] The latest version of the VVC draft, Versatile Video Coding (Draft 2), can be found at:
[0053] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 11_Ljubljana / wg11 / JVET-K1001-v7.zip.
[0054] The latest reference software for VVC is called VTM and can be found at:
[0055] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-2.1.
[0056] 2.1 Inter-frame prediction in HEVC / H.265
[0057] Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly coded as a delta relative to the predicted value.
[0058] When a CU is encoded in skip mode, one PU is associated with the CU and there are no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified whereby the motion parameters of the current PU are obtained from neighboring PUs including spatial and temporal candidates. Merge mode can be applied to any inter-predicted PU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where each PU explicitly signals the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector prediction value), the corresponding reference picture index for each reference picture list and the use of the reference picture list. In this disclosure, this mode is referred to as Advanced Motion Vector Prediction (AMVP).
[0059] When signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a sample block. This is called "unidirectional prediction." Unidirectional prediction is available for P slices and B slices.
[0060] When signaling indicates that both reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bi-prediction." Bi-prediction is only available for B slices.
[0061] The following text provides detailed information about the inter prediction modes specified in HEVC. The description starts with the Merge mode.
[0062] 2.1.1 Reference Image List
[0063] In HEVC, the term inter prediction is used to refer to predictions derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the currently decoded picture. As in H.264 / AVC, a picture can be predicted from multiple reference pictures. Reference pictures used for inter prediction are organized into one or more reference picture lists. A reference index identifies which reference picture in the list should be used to create the prediction signal.
[0064] A single reference picture list (list 0) is used for P slices, and two reference picture lists (list 0 and list 1) are used for B slices. It should be noted that the reference pictures included in lists 0 / 1 can be pictures from the past and the future in terms of capture / display order.
[0065] 2.1.2 Merge Mode
[0066] 2.1.2.1 Merge Mode Candidate Derivation
[0067] When using Merge mode to predict a PU, an index pointing to an entry in the Merge candidate list is parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized in the following sequence of steps:
[0068] Step 1: Initial candidate derivation
[0069] oStep 1.1: Spatial Candidate Derivation
[0070] oStep 1.2: Redundancy check of spatial candidates
[0071] oStep 1.3: Time Domain Candidate Derivation
[0072] Step 2: Additional candidate insertions
[0073] oStep 2.1: Create bidirectional prediction candidates
[0074] oStep 2.2: Insert zero motion candidates
[0075] These steps are also schematically depicted in Figure 1For spatial Merge candidate derivation, a maximum of four Merge candidates are selected from the candidates located at five different positions. For temporal Merge candidate derivation, a maximum of one Merge candidate is selected from the two candidates. Since a constant number of candidates for each PU is assumed at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best Merge candidate is encoded using Truncated Unarybinarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of the 2N×2N prediction unit.
[0076] Hereinafter, operations associated with the above steps will be described in detail.
[0077] Figure 1 An example derivation process for Merge candidate list construction is shown.
[0078] 2.1.2.2 Spatial Candidate Derivation
[0079] In the derivation of spatial Merge candidates, Figure 2 Up to four Merge candidates are selected from the candidates at the positions depicted in . The order of derivation is A1, B1, B0, A0 and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is not available (for example, because it belongs to another strip or slice) or is intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only pairs with Figure 3 The arrows in the list link pairs, and candidates are only added to the list if the corresponding candidates used for redundancy checking do not have the same motion information. Another source of duplicate motion information is a "second PU" associated with a partition other than 2N×2N. As an example, Figure 4A-4B The second PU is depicted for the N×2N and 2N×N cases, respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for having only one PU in the coding unit. Similarly, position B1 is not considered when the current PU is partitioned into 2N×N.
[0080] Figure 2Example locations of spatial merge candidates are shown.
[0081] Figure 3 An example of candidate pairs considered for redundancy checking of spatial merge candidates is shown.
[0082] Figure 4A-4B Example locations of the second PU for Nx2N and 2NxN partitions are shown.
[0083] 2.1.2.3 Time Domain Candidate Derivation
[0084] In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located PU belonging to the picture with the smallest POC (Picture Order Count) difference with the current picture in the given reference picture list. The reference picture list to be used for derivation of the co-located PU is explicitly signaled in the slice header. Figure 5 As shown by the dotted line in the figure, a scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors (one for reference picture list 0 and the other for reference picture list 1) are obtained and combined to generate a bidirectional prediction Merge candidate.
[0085] Figure 5 It is a diagram of the motion vector scaling of the time domain Merge candidate.
[0086] In the collocated PU(Y) belonging to the reference frame, the position of the temporal candidate is selected between candidates C0 and C1, such as Figure 6 If the PU at position C0 is not available, is intra-coded, or is outside the current coding tree unit (CTU, also known as LCU, largest coding unit) row, position C1 is used. Otherwise, position C0 is used in the derivation of the time domain merge candidate.
[0087] Figure 6 An example of candidate positions of the time-domain merge candidates C0 and C1 is shown.
[0088] 2.1.2.4 Additional Candidate Insertion
[0089] In addition to spatial and temporal Merge candidates, there are two additional types of Merge candidates: combined bi-directional prediction Merge candidate and zero Merge candidate. Combined bi-directional prediction Merge candidate is generated by utilizing spatial and temporal Merge candidates. Combined bi-directional prediction Merge candidate is only used for B slices. Combined bi-directional prediction candidate is generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another initial candidate. If these two tuples provide different motion hypotheses, they will form a new bi-directional prediction candidate. As an example, Figure 7 Shown is when two candidates in the original list (on the left) (which have mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create a combined bi-predictive Merge candidate that is added to the final list (on the right). There are many rules for combining that are considered to generate these additional Merge candidates.
[0090] Figure 7 An example of a combined bi-predictive Merge candidate is shown.
[0091] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0092] 2.1.3 AMVP
[0093] AMVP exploits the spatiotemporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the temporally neighboring PU positions to the left and above, removing redundant candidates and adding zero vectors to make the candidate list a constant length. The encoder can then select the best prediction value from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, details about the derivation process of motion vector prediction candidates will be provided.
[0094] 2.1.3.1 Derivation of AMVP Candidates
[0095] Figure 8 The process of deriving motion vector prediction candidates is outlined.
[0096] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, the final Figure 2 The motion vectors of each PU are depicted at five different positions to derive two motion vector candidates.
[0097] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different collocated positions. After generating the first spatiotemporal candidate list, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than two, motion vector candidates whose reference picture index within the associated reference picture list is greater than 1 are removed from the list. If the number of spatiotemporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.
[0098] 2.1.3.2 Spatial Motion Vector Candidates
[0099] In the derivation of spatial motion vector candidates, Figure 2 Consider up to two candidates out of the five potential candidates derived in the PU at the depicted position, which are the same positions as the motion merge positions. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the top side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not require spatial scaling and two use spatial scaling. The four different cases are summarized below:
[0100] No airspace scaling
[0101] -(1) Same reference picture list and same reference picture index (same POC)
[0102] -(2) Different reference picture lists, but same reference pictures (same POC)
[0103] Airspace scaling
[0104] -(3) Same reference picture list, but different reference pictures (different POC)
[0105] -(4) Different reference picture lists, and different reference pictures (different POCs)
[0106] First, non-spatial scaling is checked, followed by spatial scaling. Spatial scaling is considered when the POC differs between the reference pictures of the neighboring PU and the reference picture of the current PU, regardless of the reference picture list. If all PUs of the left candidate are unavailable or are intra-coded, scaling is allowed for the upper motion vector to facilitate parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not allowed for the upper motion vector.
[0107] Figure 9 Diagram showing motion vector scaling of spatial motion vector candidates.
[0108] like Figure 9 As depicted, in the spatial scaling process, the motion vectors of neighboring PUs are scaled in a similar manner to temporal scaling. The main difference is that the reference picture list and the index of the current PU are given as input; the actual scaling process is the same as the temporal scaling process.
[0109] 2.1.3.3 Temporal Motion Vector Candidates
[0110] Except for the reference picture index derivation, all the processes for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (see Figure 6 ). The reference picture index is signaled to the decoder.
[0111] 2.2 Local Illumination Compensation in JEM
[0112] Local Illumination Compensation (LIC) is based on a linear model of illumination variation using a scaling factor a and an offset b, and is adaptively enabled or disabled for each inter-mode coded coding unit (CU).
[0113] Figure 10 An example of adjacent sample points used to derive IC parameters is shown.
[0114] When LIC is applied to a CU, the least square error method is used to derive parameters a and b by using the neighboring samples of the current CU and its corresponding reference samples. More specifically, as Figure 12 As shown, the subsampling (2:1 subsampling) neighboring samples and corresponding samples (identified by the motion information of the current CU or sub-CU) of the reference picture are used.
[0115] 2.2.1 Derivation of prediction blocks
[0116] IC parameters are derived and applied separately for each prediction direction. For each prediction direction, the decoded motion information is used to generate the first prediction block, and then the temporal prediction block is obtained by applying the LIC model. The final prediction block is then derived using the two temporal prediction blocks.
[0117] When the CU is encoded in Merge mode, the LIC flag is copied from the neighboring block in a manner similar to the motion information copying in Merge mode; otherwise, the LIC flag is signaled for the CU to indicate whether LIC is applicable.
[0118] When LIC is enabled for a picture, an additional CU-level RD check is required to determine whether LIC is applied to the CU. When LIC is enabled for a CU, the Mean-Removed Sum of Absolute Difference (MR-SAD) and the Mean-Removed Sum of Absolute Hadamard-Transformed Difference (MR-SATD) (instead of SAD and SATD) are used for integer-pixel motion search and fractional-pixel motion search, respectively.
[0119] In order to reduce the coding complexity, the following coding scheme is applied in JEM.
[0120] When there is no significant illumination change between the current picture and its reference pictures, LIC is disabled for the entire picture. To identify this situation, the encoder computes the histogram of the current picture and each of its reference pictures. If the histogram difference between the current picture and each of its reference pictures is less than a given threshold, LIC is disabled for the current picture; otherwise, LIC is enabled for the current picture.
[0121] 2.3 Inter-frame prediction method in VVC
[0122] There are several new coding tools for inter-frame prediction improvements, such as Adaptive Motion Vector difference Resolution (AMVR) for signaling MVD, affine prediction mode, Triangular Prediction Mode (TPM), Advanced TMVP (ATMVP, also known as SbTMVP), Generalized Bi-Prediction (GBI), and Bi-directional Optical flow (BIO or BDOF).
[0123] 2.3.1 Coding Block Structure in VVC
[0124] In VVC, a quadtree / binarytree / multitree (QT / BT / TT) structure is used to divide the image into square blocks or rectangular blocks.
[0125] In addition to QT / BT / TT, a separate tree (also known as a dual coding tree) is also used for I frames in VVC. For the separate tree, the coding block structure is signaled separately for the luma component and the chroma component.
[0126] 2.3.2 Adaptive Motion Vector Difference Resolution
[0127] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. In VVC, Locally Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, MVD can be encoded in units of quarter luma samples, integer luma samples, or four luma samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the coding unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0128] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or quad luma sample MV precision is used.
[0129] When the first MVD resolution flag of the CU is zero, or the CU is not encoded (meaning all MVDs in the CU are zero), quarter luma sample MV precision is used for the CU. When the CU uses integer luma sample MV precision or four luma sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.
[0130] 2.3.3 Affine Motion Compensated Prediction
[0131] In HEVC, only the translation motion model is used for motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, a simplified affine transformation motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. Figure 13A and Figure 13B As shown, for the 4-parameter affine model, the affine motion field of the block is described by two control point motion vectors (CPMVs), while for the 6-parameter affine model, it is described by 3 CPMVs.
[0132] Figures 11A-11B The simplified affine motion models for the 4-parameter affine mode and the 6-parameter affine mode are shown respectively.
[0133] The motion vector field (MVF) of a block is described by the following equations, which are respectively the 4-parameter affine model in equation (1) (where the 4 parameters are defined as variables a, b, e, and f)) and the 6-parameter affine model in equation (2) (where the 4 parameters are defined as variables a, b, c, d, e, and f):
[0134]
[0135]
[0136] Among them, (mv h 0,mv h 0) is the motion vector of the upper left control point (CP), and (mv h 1,mv h 1) is the motion vector of the upper right control point, and (mv h 2,mv h 2) is the motion vector of the lower left control point. All three motion vectors are called control point motion vectors (CPMVs). (x, y) represents the coordinates of the representative point in the current block relative to the upper left sample point, and (mv h (x,y),mv v(x,y)) is the motion vector derived for the sample point at (x,y). The CP motion vector can be signaled (as in affine AMVP mode) or dynamically derived (as in affine Merge mode). w and h are the width and height of the current block. In practice, division is implemented by right shift and rounding operations. In VTM, the representative point is defined as the center position of the sub-block, for example, when the coordinates of the upper left corner of the sub-block relative to the upper left sample point in the current block are (xs,ys), the coordinates of the representative point are defined as (xs+2,ys+2). For each sub-block (for example, 4×4 in VTM), the motion vector of the entire sub-block is derived using the representative point.
[0137] In order to further simplify the motion compensation prediction, the sub-block based affine transformation prediction is applied. In order to derive the motion vector of each M×N (both M and N are set to 4 in the current VVC) sub-block, the following equations (1) and (2) can be used to calculate the motion vector of each M×N sub-block: Figure 14 The motion vector of the center sample of each sub-block shown is rounded to 1 / 16 fractional precision. Then, a 1 / 16 pixel motion compensation interpolation filter is applied to generate a prediction for each sub-block with the derived motion vector. The affine mode introduces a 1 / 16 pixel interpolation filter.
[0138] Figure 12 An example of the affine MVF for each sub-block is shown.
[0139] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the normal motion vector.
[0140] 2.3.3.1 Signaling of Affine Prediction
[0141] Similar to the translational motion model, there are two modes for signaling the side information generated by affine prediction. They are AFFINE_INTER and AFFINE_MERGE modes.
[0142] 2.3.3.2 AF_INTER mode
[0143] AF_INTER mode can be applied to CUs with width and height greater than 8. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used.
[0144] In this mode, for each reference picture list (list 0 or list 1), an affine AMVP candidate list is constructed with three types of affine motion prediction values in the following order, where each candidate includes the estimated CPMV of the current block. Figure 17The difference between the best CPMV found and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.
[0145] 1) Inheriting affine motion prediction value
[0146] The checking order is similar to the spatial MVP in HEVC AMVP list construction. First, the left-inherited affine motion prediction value is derived from the first block in {A1, A0} that is affine-coded and has the same reference picture as the current block. Second, the top-inherited affine motion prediction value is derived from the first block in {B1, B0, B2} that is affine-coded and has the same reference picture as the current block. Figure 16 depicts five blocks A1, A0, B1, B0, and B2.
[0147] Once a neighboring block is found to be coded in an affine mode, the CPMV of the coding unit covering the neighboring block is used to derive the predicted value of the CPMV of the current block. For example, if A1 is coded in a non-affine mode and A0 is coded in a 4-parameter affine mode, the affine MV predicted value inherited from the left will be derived from A0. In this case, the CPMV of the CU covering A0 (such as Figure 17 As shown, the CPMV in the upper left corner is expressed as The CPMV for the upper right corner is expressed as ) is used to derive the estimated CPMV of the current block (for the top left position (coordinate (x0, y0)), top right position (coordinate (x1, y1)) and bottom right position (coordinate (x2, y2)) of the current block is expressed as ).
[0148] 2) Constructed affine motion prediction value
[0149] like Figure 17 As shown, the constructed affine motion prediction value includes the control point motion vector (CPMV) derived from the adjacent inter-frame coded blocks with the same reference picture. If the current affine motion model is 4-parameter affine, the number of CPMVs is 2, otherwise if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. The CPMV in the upper left corner It is derived from the MV of the first block in group {A, B, C} that is inter-coded and has the same reference picture as the current block. It is derived from the MV of the first block in group {D, E} that is inter-coded and has the same reference picture as the current block. is derived from the MV at the first block in the group {F, G} that is inter-coded and has the same reference picture as the current block.
[0150] If the current affine motion model is 4-parameter affine, only when and When both are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMV of the upper left (coordinate (x0, y0)) and upper right (coordinate (x1, y1)) positions of the current block.
[0151] If the current affine motion model is 6-parameter affine, only when and are established, the constructed affine motion prediction value is inserted into the candidate list, that is, and Used as the estimated CPMVs for the top left (coordinate (x0, y0)), top right (coordinate (x1, y1)), and bottom right (coordinate (x2, y2)) positions of the current block.
[0152] When inserting the constructed affine motion predictors into the candidate list, no pruning process is applied.
[0153] 3) Normal AMVP motion prediction value
[0154] The following conditions apply until the number of affine motion predictors reaches a maximum value.
[0155] 1) By setting all CPMVs equal to (if available) to derive affine motion prediction values.
[0156] 2) By setting all CPMVs equal to (if available) to derive affine motion prediction values.
[0157] 3) By setting all CPMVs equal to (if available) to derive affine motion prediction values.
[0158] 4) Affine motion prediction values are derived by setting all CPMVs equal to HEVC TMVP (if available).
[0159] 5) Derive the affine motion prediction value by setting all CPMVs to zero MV.
[0160] Notice, It has been derived in the constructed affine motion predictor.
[0161] Figures 13A-13B Examples of a 4-parameter affine model and a 6-parameter affine model are shown respectively.
[0162] Figure 14 MVP of AF_INTER of inherited affine candidates is shown.
[0163] Figure 15 The MVP of AF_INTER of the constructed affine candidate is shown.
[0164] Figures 16A-16B Examples of candidates for AF_Merge are shown.
[0165] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVDs need to be encoded for these control points, such as Figure 15 In JVET-K0337, it is proposed to derive MV as follows: predict mvd1 and mvd2 from mvd0.
[0166]
[0167]
[0168]
[0169] in, mvd i and mv1 are the predicted motion vector, motion vector difference, and motion vector of the upper left pixel (i=0), upper right pixel (i=1), or lower left pixel (i=2), respectively. Figure 15 As shown. Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equivalent to the individual summation of the two components. That is, newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB), respectively.
[0170] 2.3.3.3 AF_Merge Mode
[0171] When applying CU in AF_MERGE mode, it obtains the first block coded in affine mode from the valid adjacent reconstructed blocks. And the selection order of candidate blocks is from left, top, top right, bottom left to top left, such as Figure 17 As shown (indicated by A, B, C, D, E in sequence). For example, if the adjacent lower left block is encoded in affine mode, such as Figure 17 As represented by A0 in the figure, the control point (CP) motion vector mv0 of the upper left corner, upper right corner and lower left corner of the adjacent CU / PU containing block A is extracted. N 、mv1N and mv2 N . And based on mv0 N 、mv1 N and mv2 N To calculate the motion vector mv0 of the upper left corner / upper right / lower left corner of the current CU / PU C 、mv1 C and mv2 C (This is only used for the 6-parameter affine model.) It should be noted that in VTM-2.0, if the current block is affine-coded, the sub-block at the top left (e.g., a 4×4 block of VTM) stores mv0, and the sub-block at the top right stores mv1. If the current block is coded with a 6-parameter affine model, the sub-block at the bottom left stores mv2; otherwise (in the case of a 4-parameter affine model), LB stores mv2'. The other sub-blocks store the MV for MC.
[0172] The CPMV mv0 of the current CU is derived from the simplified affine motion model in equations (1) and (2). C and mv1 C and mv2 C After that, the MVF of the current CU is generated. In order to identify whether the current CU is coded in AF_MERGE mode, when there is at least one neighboring block coded in affine mode, the affine flag is signaled in the bitstream.
[0173] In JVET-L0142 and JVET-L0632, the affine merge candidate list is constructed by the following steps:
[0174] 1) Insert inherited affine candidates
[0175] An inherited affine candidate is one derived from the affine motion model of its valid neighboring affine coded blocks. Up to two inherited affine candidates are derived from the affine motion models of the neighboring blocks and inserted into the candidate list. For the left-hand prediction, the scan order is {A0, A1}; for the top prediction, the scan order is {B0, B1, B2}.
[0176] 2) Insert the constructed affine candidate
[0177] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (for example, set to 5), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.
[0178] a) The motion information of the control point is first obtained from Figure 19The derivation of the specified spatial and temporal neighbors is shown. CPk (k = 1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations used to predict CPk (k = 1, 2, 3); T is the temporal location used to predict CP4.
[0179] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.
[0180] Figure 17 An example of candidate positions for the affine merge mode is shown.
[0181] The motion information for each control point is obtained according to the following priority order:
[0182] - For CP1, the priority is B2->B3->A2. If B2 is available, use B2. Otherwise, if B2 is available, use B3. If neither B2 nor B3 is available, use A2. If none of the three candidates are available, motion information for CP1 cannot be obtained.
[0183] - For CP2, the check priority is B1->B0.
[0184] - For CP3, the check priority is A1->A0.
[0185] -For CP4, use T.
[0186] b) Secondly, the combination of control points is used to construct affine merge candidates.
[0187] I. Constructing a 6-parameter affine candidate requires motion information from three control points. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} are converted into a 6-parameter motion model represented by the top-left, top-right, and bottom-left control points.
[0188] II. Constructing a 4-parameter affine candidate requires motion information for two control points. These two control points can be selected from one of the following two combinations: {CP1, CP2} and {CP1, CP3}. These two combinations are converted into a 4-parameter motion model represented by the top left and top right control points.
[0189] III. Insert the constructed combinations of affine candidates into the candidate list in the following order:
[0190] {CP1,CP2,CP3}, {CP1,CP2,CP4}, {CP1,CP3,CP4}, {CP2,CP3,CP4}, {CP1,CP2}, {CP1,CP3}
[0191] i. For each combination, check the reference index of each CP in List X. If they are all the same, then the combination has a valid CPMV for List X. If the combination does not have a valid CPMV for both List 0 and List 1, then the combination is marked as invalid. Otherwise, it is valid, and the CPMV is placed in the sub-block Merge list.
[0192] 3) Fill with zero motion vectors
[0193] If the number of candidates in the affine merge candidate list is less than 5, a zero motion vector with a zero reference index is inserted into the candidate list until the list is full.
[0194] More specifically, for the sub-block Merge candidate list, a 4-parameter Merge candidate, where MV is set to (0,0) and prediction direction is set to unidirectional prediction (for P slices) and bidirectional prediction (for B slices) from list 0.
[0195] 2.3.4 Merge with Motion Vector Difference (MMVD)
[0196] In JVET-L0054, the final motion vector expression (UMVE, also called MMVD) is proposed. UMVE is used for skip mode or merge mode through the proposed motion vector expression method.
[0197] UMVE reuses the same merge candidates as the regular merge candidate list in VVC. Among the merge candidates, a base candidate can be selected and further extended by the proposed motion vector representation method.
[0198] UMVE provides a new method to represent motion vector difference (MVD), which uses the starting point, motion amplitude and motion direction to represent MVD.
[0199] Figure 18 An example of the UMVE search process is shown.
[0200] Figure 19 An example of a UMVE search point is shown.
[0201] The proposed technique uses the Merge candidate list as is, but for the extension of UMVE, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered.
[0202] The base candidate index (IDX) defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.
[0203] Table 1. Basic candidate IDX
[0204]
[0205] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.
[0206] The distance index is the motion magnitude information. The distance index indicates the predefined distance from the start point information. The predefined distances are as follows:
[0207] Table 2. Distance IDX
[0208]
[0209] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown below.
[0210] Table 3. Direction IDX
[0211] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0212] The UMVE flag is signaled immediately after the Skip or Merge flag is sent. If the Skip or Merge flag is true, the UMVE flag is parsed. If the UMVE flag is 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is 1, AFFINE mode is used. However, if it is not 1, the Skip / Merge index is parsed for the VTM's Skip / Merge mode.
[0213] No additional line buffer is needed for UMVE candidates, as software skip / merge candidates are directly used as base candidates. The MV complement is determined before motion compensation using the input UMVE index. There's no need to maintain a long line buffer for this purpose.
[0214] Under the current general test conditions, the first or second merge candidate in the merge candidate list can be selected as the basic candidate.
[0215] UMVE is also known as Merge with MV Difference (MMVD).
[0216] 2.3.5 Decoder-side Motion Vector Refinement (DMVR)
[0217] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MVs) from list 0 and MVs from list 1, respectively, are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.
[0218] 2.3.5.1 DMVR in JEM
[0219] In the JEM design, motion vectors are refined through a bilateral template matching process. Bilateral template matching is applied in the decoder to perform a distortion-based search between the bilateral template and the reconstructed samples in the reference picture to obtain the refined MV without sending additional motion information. Figure 22 An example is described in . From the initial list 0 MV0 and list 1 MV1, bilateral templates are generated as a weighted combination (i.e., average) of the two prediction blocks, as shown in Figure 22 As shown. The template matching operation includes calculating the cost metric between the generated template and the sample area (around the initial prediction block) in the reference picture. For each of the two reference pictures, the MV that produces the minimum template cost is considered to be the updated MV in the list to replace the original MV. In JEM, 9 MV candidates are searched for each list. The 9 MV candidates include the original MV and 8 surrounding MVs, where the 8 surrounding MVs are offset by one luminance sample relative to the original MV in the horizontal or vertical direction or in both directions. Finally, the two new MVs (i.e., Figure 22 MV0′ and MV1′ shown are used to generate the final bidirectional prediction result. The sum of absolute differences (SAD) is used as the cost metric. Note that when calculating the cost of a prediction block generated by a surrounding MV, the rounded MV (rounded to integer pixels) is actually used to obtain the prediction block, not the actual MV.
[0220] Figure 20 An example of DMVR based on bilateral template matching is shown.
[0221] 2.3.5.2 DMVR in VVC
[0222] For DMVR in VVC, such as Figure 21As shown, the MVDs between List 0 and List 1 are assumed to be mirrored, and bilateral matching is performed to refine the MV, i.e., to find the best MVD among several MVD candidates. The MVs of the two reference picture lists are represented by MVL0 (L0X, L0Y) and MVL1 (L1X, L1Y). For List 0, the MVD represented by (MvdX, MvdY) that minimizes a cost function (e.g., SAD) is defined as the best MVD. The SAD function is defined as the SAD between the reference block in List 0 and the reference block in List 1, where the reference block in List 0 is derived using the motion vector (L0X+MvdX, L0Y+MvdY) in the reference picture in List 0, and the reference block in List 1 is derived using the motion vector (L1X-MvdX, L1Y-MvdY) in the reference picture in List 1.
[0223] The motion vector refinement process can be iterated twice. Figure 22 As shown, in each iteration, up to 6 MVDs (with integer pixel precision) can be checked in two steps. In the first step, the MVDs (0,0), (-1,0), (1,0), (0,-1), (0,1) are checked. In the second step, one of the MVDs (-1,-1), (-1,1), (1,-1), or (1,1) can be selected and further checked. Assume that the function Sad(x,y) returns the SAD value of MVD(x,y). The MVD denoted by (MvdX,MvdY) checked in the second step is determined as follows:
[0224] MvdX=-1;
[0225] MvdY=-1;
[0226] If (Sad(1,0) <Sad(-1,0))
[0227] Then MvdX=1;
[0228] If (Sad(0,1) <Sad(0,-1))
[0229] Then MvdY=1;
[0230] In the first iteration, the starting point is the signaled MV, and in the second iteration, the starting point is the signaled MV plus the best MVD selected in the first iteration. DMVR is only applicable when one reference picture is the previous picture and the other reference picture is the next picture, and both reference pictures have the same picture order count distance from the current picture.
[0231] Figure 21 An example of MVD(0,1) mirrored between list 0 and list 1 in DMVR is shown.
[0232] Figure 22 An example of MVs that can be checked in one iteration is shown.
[0233] To further simplify the DMVR process, JVET-M0147 proposes several changes to the JEM design. More specifically, the DMVR design adopted for VTM-4.0 (to be released) has the following key features:
[0234] When the SAD at position (0,0) between list 0 and list 1 is less than a threshold, terminate early.
[0235] When the SAD between list 0 and list 1 is zero for a certain position, terminate early.
[0236] DMVR block size: W*H>=64&&H>=8, where W and H are the width and height of the block.
[0237] For DMVR of CUs with size > 16*16, the CU is divided into multiple 16×16 sub-blocks. If only the width or height of the CU is greater than 16, the division is performed only in the vertical or horizontal direction.
[0238] • Reference block size (W+7)*(H+7) (for luma).
[0239] Integer pixel search based on 25-point SAD (i.e. (+-)2 refinement search range, single stage)
[0240] DMVR based on bilinear interpolation.
[0241] Sub-pixel refinement based on the “parametric error surface equation”. This process is performed only when the minimum SAD cost is not equal to zero and the best MVD is (0,0) in the last MV refinement iteration.
[0242] • Luma / Chroma MC with reference block filling (if needed).
[0243] Only used for MCV and TMVP refinement MV.
[0244] 2.3.5.2.1 Use of DMVR
[0245] DMVR can be enabled when all of the following conditions are true:
[0246] -DMVR enabled flag in SPS (i.e. sps_dmvr_enabled_flag) is equal to 1
[0247] -TPM flag, inter-frame affine flag, sub-block Merge flag (ATMVP or affine Merge), and MMVD flag are all equal to 0
[0248] -Merge flag equal to 1
[0249] - The current block is bi-predicted and the POC distance between the current picture and the reference picture in list 1 is equal to the POC distance between the reference picture in list 0 and the current picture
[0250] -Current CU height is greater than or equal to 8
[0251] - The number of luma samples (CU width * height) is greater than or equal to 64
[0252] 2.3.5.2.2 Sub-pixel Refinement Based on the “Parameter Error Surface Equation”
[0253] The method is outlined as follows:
[0254] 1. The parameter error surface fit is computed only if the center position is the best cost position in a given iteration.
[0255] 2. The center position cost and the costs at the positions (-1, 0), (0, -1), (1, 0), and (0, 1) from the center are used to fit the 2-D parabolic error surface equation of the following form:
[0256] E(x, y) = A(x-x0) 2 +B(y-y0) 2 +C
[0257] Where (x0, y0) corresponds to the least expensive position, and C corresponds to the minimum cost value. By solving the five equations for the five unknowns, (x0, y0) is calculated as follows:
[0258] x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0)))
[0259] y0=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))
[0260] (x0, y0) can be calculated to any desired sub-pixel precision by adjusting the precision with which the division is performed (i.e., how many bits of the quotient are calculated). For 1 / 16 pixel precision, only 4 bits of the absolute value of the quotient need to be calculated, which facilitates the implementation of the 2 divisions required per CU based on fast shift-subtract.
[0261] 3. Add the calculated (x0, y0) to the integer distance refinement MV to get the sub-pixel accurate refinement increment MV.
[0262] 2.3.6 Combined Intra- and Inter-frame Prediction
[0263] In JVET-L0100, multi-hypothesis prediction is proposed, where combined intra-frame and inter-frame prediction is a way to generate multiple hypotheses.
[0264] When multi-hypothesis prediction is applied to improve intra mode, multi-hypothesis prediction combines an intra prediction and a merge index prediction. In a merge CU, when the flag is true, a flag is signaled for the merge mode to select the intra mode from the intra candidate list. For the luma component, the intra candidate list is derived from 4 intra prediction modes including DC, planar, horizontal and vertical modes, and the size of the intra candidate list can be 3 or 4 depending on the block shape. When the CU width is greater than twice the CU height, the horizontal mode is not included in the intra mode list, and when the CU height is greater than twice the CU width, the vertical mode is removed from the intra mode list. A weighted average is used to combine one intra prediction mode selected by the intra mode index and one merge index prediction selected by the merge index. For chroma components, DM is always applied without additional signaling. The weights of the combined prediction are described as follows. When DC or planar mode is selected, or the CB width or height is less than 4, equal weights are applied. For those CBs whose width and height are greater than or equal to 4, when horizontal / vertical mode is selected, one CB is first divided vertically / horizontally into four equal-area regions. i ,w_inter i ) will be applied to the corresponding region, where i is from 1 to 4, and (w_intra1, w_inter1) = (6, 2), (w_intra2, w_inter2) = (5, 3), (w_intra3, w_inter3) = (3, 5), and (w_intra4, w_inter4) = (2, 6). (w_intra1, w_inter1) is used for the region closest to the reference sample, and (w_intra4, w_inter4) is used for the region farthest from the reference sample. The combined prediction can then be calculated by adding the two weighted predictions and shifting them right by 3 bits. In addition, the intra prediction mode of the intra hypothesis of the predicted value can be saved for subsequent reference by neighboring CUs.
[0265] 2.3.7. Symmetrical Motion Vector Difference in JVET-M0481
[0266] In JVET-M0481, Symmetric Motion Vector Difference (SMVD) is proposed for motion information coding in bidirectional prediction.
[0267] First, at the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows:
[0268] – Search for the closest forward reference picture to the current picture in reference picture list 0. If found, set RefIdxSymL0 equal to the reference index of the forward picture.
[0269] – Search for the backward reference picture closest to the current picture in reference picture list 1. If found, set RefIdxSymL1 equal to the reference index of the backward picture.
[0270] – If both the forward and backward pictures are found, BiDirPredFlag is set equal to 1.
[0271] – Otherwise, the following applies:
[0272] – Search for the backward reference picture closest to the current picture in reference picture list 0. If found, set RefIdxSymL0 equal to the reference index of the backward picture.
[0273] – Search for the closest forward reference picture to the current picture in reference picture list 1. If found, set RefIdxSymL1 equal to the reference index of the forward picture.
[0274] – If both the forward and backward pictures are found, BiDirPredFlag is set to 1. Otherwise, BiDirPredFlag is set to 0.
[0275] Second, at the CU level, if the prediction direction of the CU is bidirectional prediction and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled.
[0276] When the flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. For list 0 and list 1, the reference index is set equal to RefIdxSymL0 and RefIdxSymL1, respectively. MVD1 is just set equal to –MVD0. The final motion vector is shown below.
[0277]
[0278] Figure 25 An example diagram of a symmetric pattern is shown.
[0279] The changes to the code unit syntax are shown in Table 4 (in bold italics)
[0280] Table 4 - Modifications to Code Unit Syntax
[0281]
[0282]
[0283] 3. Problems with current video coding technology
[0284] The current decoder-side motion vector derivation (DMVD) may have the following problems:
[0285] 1. Enable DMVR even when weighted prediction is enabled for the current picture.
[0286] 2. When two reference pictures have different POC distances from the current picture, DMVR is disabled.
[0287] 3. For different block sizes, enable DMVR and BIO.
[0288] a. When W*H>=64&&H>=8, enable DMVR
[0289] b. When H>=8&&! (W==4&&H==8), enable BIO
[0290] 4. Perform DMVR and BIO at different sub-block levels.
[0291] a. DMVR can be performed at the sub-block level. When both the width and height of the CU are greater than 16, it is divided into 16×16 sub-blocks. Otherwise, when the width of the CU is greater than 16, it is divided into 16×H sub-blocks in the vertical direction, and when the height of the CU is greater than 16, it is divided into W×16 sub-blocks in the horizontal direction.
[0292] b. Perform BIO at the block level.
[0293] 4. Example Techniques and Embodiments
[0294] The following detailed techniques should be considered as examples to explain the general concept. These techniques should not be interpreted narrowly. In addition, these techniques can be combined in any way.
[0295] In this document, DMVD includes methods that perform motion estimation to derive or refine block / sub-block motion information (such as DMVR and FRUC) and BIO that performs sample-wise motion refinement.
[0296] The unequal weights applied to the prediction blocks may refer to weights used in a GBI process, LIC process, weighted prediction process, or other encoding / decoding processes of a coding tool that require applying additional operations to the prediction blocks other than averaging two prediction blocks, etc.
[0297] Assume that the reference pictures in list 0 and list 1 are Ref0 and Ref1, respectively, that the POC distance between the current picture and Ref0 is PocDist0 (i.e., the POC of the current picture minus the POC of Ref0), and that the POC distance between Ref1 and the current picture is PocDist1 (i.e., the POC of Ref1 minus the POC of the current picture). In this patent document, PocDist1 is the same as PocDis1, and PocDist0 is the same as PocDis0. Denote the width and height of a block as W and H, respectively. Assume that the function abs(x) returns the absolute value of x.
[0298] 1. Parameters (eg, weight information) applied to the prediction block in the final prediction block generation process can be used in the DMVD process.
[0299] a. These parameters can be signaled to the decoder, such as using GBi or weighted prediction. GBi is also called bi-prediction CU Weight (BCW) with coding unit (CU) weights.
[0300] b. These parameters can be derived at the decoder, such as with LIC.
[0301] c. These parameters can be used in the shaping process of mapping a set of sample values to another set of sample values.
[0302] d. In one example, the parameters applied to the prediction block can be applied to the DMVD.
[0303] i. In one example, when calculating a cost function (eg, SAD, MR-SAD, gradient), a weighting factor according to the GBI index is first applied to the predicted block, and then the cost is calculated.
[0304] ii. In one example, when calculating a cost function (eg, SAD, MR-SAD, gradient), a weighting factor and / or offset according to weighted prediction is first applied to the predicted block, and then the cost is calculated.
[0305] iii. In one example, when calculating a cost function (eg, SAD, MR-SAD, gradient), a weighting factor and / or offset according to the LIC parameters is first applied to the predicted block, and then the cost is calculated.
[0306] iv. In one example, when calculating temporal and spatial gradients in a BIO, a weighting factor according to the GBI index is first applied to the prediction block, and then these gradients are calculated.
[0307] v. In one example, when calculating temporal and spatial gradients in a BIO, weighting factors and / or offsets according to weighted prediction are first applied to the predicted block, and then these gradients are calculated.
[0308] vi. In one example, when calculating temporal and spatial gradients in a BIO, weighting factors and / or offsets according to LIC parameters are first applied to the predicted block, and then these gradients are calculated.
[0309] vii. Alternatively, in addition, the cost calculation (eg, SAD, MR-SAD) / gradient calculation is performed in the reshaped domain.
[0310] viii. Alternatively, furthermore, after the motion information is refined, the shaping process is disabled for the prediction block generated using the refined motion information.
[0311] e. In one example, DMVD can be disabled in GBI mode or / and LIC mode or / and weighted prediction or / and multi-hypothesis prediction.
[0312] f. In one example, DMVD can be disabled in weighted prediction when the weighting factors and / or offsets of two reference pictures are different.
[0313] g. In one example, DMVD can be disabled in LIC when the weighting factors and / or offsets of two reference blocks are different.
[0314] 2. Even when the first picture order count distance (PocDis0) is not equal to the second picture order count distance (PocDis1), a DMVD process (eg, DMVR or BIO) may be applied to a bidirectionally predicted block.
[0315] a. In one example, all DMVD processes may be enabled or disabled according to the same rules as for PocDis0 and PocDis1.
[0316] i. For example, when PocDis0 is equal to PocDis1, all DMVD processes may be enabled.
[0317] ii. For example, when PocDis0 is not equal to PocDis1, all DMVD processes may be enabled.
[0318] 1. Alternatively, furthermore, when PocDis0*PocDist1 is less than 0, all DMVD processes may be disabled.
[0319] iii. For example, when PocDis0 is not equal to PocDis1, all DMVD processes may be disabled.
[0320] iv. For example, when PocDis0*PocDist1 is less than 0, all DMVD processes may be disabled.
[0321] b. In one example, for the case where PocDis0 is equal to PocDis1, the current design is enabled.
[0322] i. In one example, the MVD of list 0 can be mirrored to list 1. That is, if (MvdX, MvdY) is used for list 0, then (-MvdX, -MvdY) is used for list 1 to identify the two reference blocks.
[0323] ii. Alternatively, the MVD of list 1 can be mirrored to list 0. That is, if (MvdX, MvdY) is used for list 1, then (-MvdX, -MvdY) is used for list 0 to identify the two reference blocks.
[0324] c. Alternatively, instead of using mirrored MVD for List 0 and List 1 (i.e., (MvdX, MvdY) is used for List 0 and then (-MvdX, -MvdY) can be used for List 1), non-mirrored MVD can be used to identify the two reference blocks.
[0325] i. In one example, the MVD of list 0 can be scaled to list 1 according to PocDist0 and PocDist1.
[0326] 1. Denote the selected MVD for List 0 by (MvdX, MvdY), and then select (-MvdX*PocDist1 / PocDist0, -MvdY*PocDist1 / PocDist0) as the MVD applied to List 1.
[0327] ii. In one example, the MVD of List 1 may be scaled to List 0 based on PocDist0 and PocDist1.
[0328] 1. Denote the selected MVD for List 1 by (MvdX, MvdY), and then select (-MvdX*PocDist0 / PocDist1, -MvdY*PocDist0 / PocDist1) as the MVD applied to List 0.
[0329] iii. The division operation in scaling can be implemented by lookup tables, multiple operations, and right-right operations.
[0330] d. How to define the MVD of two reference pictures (eg, whether to use mirroring or scaling by MVD) may depend on the reference pictures.
[0331] i. In one example, if abs(PocDist0) is less than or equal to abs(PocDist1), the MVD of list 0 may be scaled to list 1 based on PocDist0 and PocDist1.
[0332] ii. In one example, if abs(PocDist0) is greater than or equal to abs(PocDist1), the MVD of list 0 may be scaled to list 1 according to PocDist0 and PocDist1.
[0333] iii. In one example, if abs(PocDist1) is less than or equal to abs(PocDist0), the MVD of list 1 may be scaled to list 0 according to PocDist0 and PocDist1.
[0334] iv. In one example, if abs(PocDist1) is greater than or equal to abs(PocDist0), the MVD of list 1 may be scaled to list 0 based on PocDist0 and PocDist1.
[0335] v. In one example, if one reference picture is the previous picture of the current picture and the other reference picture is the next picture of the current picture, the MVD of list 0 can be mirrored to list 1 and no MVD scaling is performed.
[0336] e. Whether and how to apply a DMVD may depend on the sign of PocDist0 and the sign of PocDist1.
[0337] i. In one example, a DMVD can only be performed when PocDist0*PocDist1<0.
[0338] ii. In one example, a DMVD can only be performed when PocDist0*PocDist1>0.
[0339] f. Alternatively, when PocDist0 is not equal to PocDist1, the DMVD process (eg, DMVR or BIO) may be disabled.
[0340] 3. DMVR or / and other DMVD methods can be enabled in SMVD mode.
[0341] a. In one example, the decoded MVD / MV from the bitstream according to the SMVD mode can be further refined before being used to decode a block.
[0342] b. In one example, in SMVD mode, if the MV / MVD precision is N pixels, DMVR and / or other DMVD methods can be used to refine the MVD via mvdDmvr. mvdDmvr has M pixel precision. N, M = 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, 16, etc.
[0343] i. In one example, M may be less than or equal to N.
[0344] c. In one example, MVD may not be signaled in SMVD mode, but DMVR or / and other DMVD methods may be applied to generate MVD.
[0345] i. Alternatively, furthermore, AMVR information may not be signaled, and MV / MVD accuracy with a predetermined value may be derived (eg, MVD with 1 / 4 pixel accuracy).
[0346] 1. In one example, the indication of the predefined value may be signaled at the sequence / picture / slice group / slice / slice / video data unit level.
[0347] 2. In one example, the predefined value may depend on the mode / motion information, such as affine or non-affine motion.
[0348] d. In one example, for SMVD coded blocks, an indication of whether DMVR or / and other DMVD methods are applied may be signaled.
[0349] i. If DMVR or / and other DMVD methods are applied, MVD may not be signaled.
[0350] ii. In one example, for certain MV / MVD accuracies, such an indication may be signaled. For example, for MV / MVD accuracies of 1 pixel or / and 4 pixels, such an indication may be signaled.
[0351] iii. In one example, such indication may be signaled only when PocDist0 is equal to PocDist1 and Ref0 is the previous picture of the current picture and Ref1 is the subsequent picture of the current picture in display order.
[0352] iv. In one example, such indication may be signaled only when PocDist0 is equal to PocDist1 and Ref0 is the subsequent picture of the current picture and Ref1 is the previous picture of the current picture in display order.
[0353] e. In one example, whether DMVR and / or other DMVD methods are applied to an SMVD-coded block depends on coding information of the current block and / or neighboring blocks.
[0354] i. For example, whether DMVR or / and other DMVD methods are applied to SMVD-coded blocks may depend on the current block dimensions.
[0355] ii. For example, whether DMVR or / and other DMVD methods are applied to an SMVD-coded block may depend on information of a reference picture, such as POC.
[0356] iii. For example, whether DMVR or / and other DMVD methods are applied to SMVD-coded blocks may depend on the signaled MVD information.
[0357] 4. DMVR or / and BIO or / and all DMVD methods can be enabled according to the same rules regarding block dimensions.
[0358] a. In one example, when W*H>=T1&&H>=T2, DMVR and BIO or / and all DMVD methods and / or the proposed method may be enabled. For example, T1=64 and T2=8.
[0359] b. In one example, H>=T1&&!(W==T2&&H==T1), DMVR and BIO or / and all DMVD methods may be enabled. For example, T1=8 and T2=4.
[0360] c. In one example, when the block size contains less than M*H samples, eg, 16 or 32 or 64 luma samples, DMVR and BIO or / and all DMVD methods are not allowed.
[0361] d. In one example, when the block size contains more than M*H samples, for example, 16 or 32 or 64 luma samples, DMVR and BIO or / and all DMVD methods are not allowed.
[0362] e. Alternatively, DMVR and BIO or / and all DMVD methods are not allowed when the minimum size of the width or / and height of the block is less than or not greater than X. In one example, X is set to 8.
[0363] f. Alternatively, when the width of the block is > th1 or >= th1 and / or the height of the block is > th2 or >= th2, DMVR and BIO or / and all DMVD methods are not allowed. In one example, th1 and / or th2 are set to 64.
[0364] i. For example, disable DMVR and BIO or / and all DMVD methods for M×M (e.g., 128×128) blocks.
[0365] ii. For example, disable DMVR and BIO or / and all DMVD methods for N×M / M×N blocks, e.g., where N >= 64 and M = 128.
[0366] iii. For example, disable DMVR and BIO or / and all DMVD methods for N×M / M×N blocks, e.g., where N >= 4 and M = 128.
[0367] g. Alternatively, when the width of the block < th1 or <= th1 and / or the height of the block < th2 or <= th2, DMVR and BIO or / and all DMVD methods are not allowed. In one example, th1 and / or th2 are set to 8.
[0368] 5. DMVR and BIO or / and all DMVD methods can be performed at the same sub - block level.
[0369] a. A motion vector refinement process such as DMVR can be performed at the sub - block level.
[0370] i. Bilateral matching can be performed at the sub - block level instead of the entire block level.
[0371] b. BIO can be performed at the sub - block level.
[0372] i. In one example, the determination of enabling / disabling BIO can be made at the sub - block level.
[0373] ii. In one example, sample - based motion refinement in BIO can be performed at the sub - block level.
[0374] iii. In one example, the determination of enabling / disabling BIO and sample - based motion refinement in BIO can be performed at the sub - block level.
[0375] c. In one example, when the width of the block >= LW or the height >= LH or the width >= LW and the height >= LH, the block can be divided into multiple sub - blocks. Each sub - block is treated in the same way as a normal coding block with a size equal to the sub - block size.
[0376] i. In one example, L is 64, a 64×128 / 128×64 block is divided into two 64×64 sub - blocks, and a 128×128 block is divided into four 64×64 sub - blocks. However, N×128 / 128×N blocks (where N < 64) are not divided into sub - blocks. The L value can refer to LH and / or LW.
[0377] ii. In one example, L is 64, a 64×128 / 128×64 block is divided into two 64×64 sub-blocks, and a 128×128 block is divided into four 64×64 sub-blocks. Meanwhile, an N×128 / 128×N block (where N<64) is divided into two N×64 / 64×N sub-blocks. The L value can refer to LH and / or LW.
[0378] iii. In one example, when the width (or height) is greater than L, it is divided vertically (or horizontally), and the width and / or height of the sub-block is not greater than L. The L value may refer to LH and / or LW.
[0379] d. In one example, when the size of a block (ie, width*height) is greater than a threshold L1, it can be divided into multiple sub-blocks. Each sub-block is treated in the same way as a normal coding block with a size equal to the sub-block size.
[0380] i. In one example, a block is divided into sub-blocks of equal size and no larger than L1.
[0381] ii. In one example, if the width (or height) of a block is not greater than the threshold L2, it is not split vertically (or horizontally).
[0382] iii. In one example, L1 is 1024, and L2 is 32. For example, a 16×128 block is divided into two 16×64 sub-blocks.
[0383] e. The threshold L can be predefined or signaled at the SPS / PPS / picture / slice / slice group / slice level.
[0384] f. Alternatively, the threshold may depend on certain coding information, such as block size, picture type, temporal layer index, etc.
[0385] 6. The decision of whether and how to apply a DMVD can be made once and shared by all color components, or can be made multiple times for different color components.
[0386] a. In one example, the decision for DMVD is made based on information of the Y (or G) component, and then the other color components.
[0387] b. In one example, the DMVD applied to the Y (or G) component is determined based on information of the Y (or G) component, and the DMVD applied to the Cb (or Cb or B or R) component is determined based on information of the Cb (or Cb or B or R) component.
[0388] Figure 23is a block diagram of a video processing device 2300. The device 2300 can be used to implement one or more methods described herein. The device 2300 can be embodied in a smartphone, a tablet, a computer, an Internet of Things (IoT) receiver, etc. The device 2300 may include one or more processors 2302, one or more memories 2304, and video processing hardware 2306. The processor(s) 2302 can be configured to implement one or more methods described in this document. The memory(s) 2304 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2306 can be used to implement some of the techniques described in this document in hardware circuitry and can be partially or completely part of the processor 2302 (e.g., a graphics processor core GPU or other signal processing circuitry).
[0389] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of a current video block may, for example, correspond to bits that are collocated or distributed across different locations within the bitstream, as defined by the syntax. For example, a macroblock may be encoded based on a transformed and coded error residual value, and may also be encoded using bits in the header and other fields in the bitstream.
[0390] It should be appreciated that several techniques have been disclosed that will benefit video encoder and decoder embodiments incorporated into video processing devices such as smartphones, laptops, desktop computers, and similar devices by enabling use of the techniques disclosed in this document.
[0391] Figure 24A 24 is a flow chart of an example method 2400 for video processing. The method 2400 includes, at 2402, obtaining refined motion information for a current video block of a video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least weight parameters, wherein the weight parameters are applied to the prediction block during generation of a final prediction block for the current video block. The method 2400 includes, at 2404, performing conversion between the current video block and a bitstream representation of the video using at least the refined motion information and the weight parameters.
[0392] In some embodiments of method 2400, a field in the bitstream representation indicates a weight parameter. In some embodiments of method 2400, the indication of the weight parameter is signaled using a bidirectional prediction (BCW) technique with coding unit weights. In some embodiments of method 2400, the indication of the weight parameter is signaled using a weighted prediction technique. In some embodiments of method 2400, the weight parameter is derived. In some embodiments of method 2400, the weight parameter is derived using a local illumination compensation (LIC) technique. In some embodiments of method 2400, the weight parameter is associated with a shaping process that maps a set of sample values to another set of sample values. In some embodiments of method 2400, the DMVD scheme is implemented by applying the weight parameter to a prediction block of the current video block. In some embodiments of method 2400, the conversion includes calculating a prediction cost function for the current video block by first applying the weight parameter indexed by bidirectional prediction (BCW) with coding unit weights to the prediction block and then calculating the prediction cost function.
[0393] In some embodiments of method 2400, converting includes calculating a prediction cost function for the current video block by first applying weight parameters according to a weighted prediction scheme to the prediction block and then calculating the prediction cost function. In some embodiments of method 2400, converting includes calculating a prediction cost function for the current video block by first applying weight parameters according to a local illumination compensation (LIC) scheme to the prediction block and then calculating the prediction cost function. In some embodiments of method 2400, the prediction cost function is a gradient function. In some embodiments of method 2400, the prediction cost function is a sum of absolute differences (SAD) cost function. In some embodiments of method 2400, the prediction cost function is a mean removed sum of absolute differences (MR-SAD) cost function.
[0394] In some embodiments of the method 2400, the converting includes calculating temporal gradients and spatial gradients of a bidirectional optical flow (BIO) scheme for the current video block by first applying weight parameters according to a bidirectional prediction (BCW) index with coding unit weights to the prediction block, and then calculating the temporal gradients and the spatial gradients. In some embodiments of the method 2400, the converting includes calculating temporal gradients and spatial gradients of a bidirectional optical flow (BIO) scheme for the current video block by first applying weight parameters according to a weighted prediction scheme to the prediction block, and then calculating the temporal gradients and the spatial gradients.
[0395] In some embodiments of method 2400, the conversion includes computing temporal and spatial gradients of a bidirectional optical flow (BIO) scheme for the current video block by first applying weight parameters according to a local illumination compensation (LIC) scheme to the prediction block and then computing the temporal and spatial gradients. In some embodiments of method 2400, the computation of the prediction cost function or the temporal or spatial gradients is performed in a shaping domain. In some embodiments of method 2400, the shaping process is disabled for the prediction block generated using refined motion information of the current video block.
[0396] Figure 24B 24 is a flow chart of an example method 2410 for video processing. The method 2410 includes, at 2412, determining that due to use of an encoding tool for a current video block of a video, use of a decoder-side motion vector derivation (DMVD) scheme is disabled for conversion between the current video block of the video and an encoded representation of the video. The method 2410 includes, at 2414, performing conversion between the current video block and a bitstream representation of the video based on the determination, wherein the encoding tool includes applying unequal weighting factors to prediction blocks for the current video block. In some embodiments of the method 2410, the encoding tool is configured to use the weighting factors in a sample prediction process.
[0397] In some embodiments of method 2410, the encoding tool includes a bidirectional prediction (BCW) mode with coding unit weights. In some embodiments of method 2410, the two weighting factors used for two prediction blocks in the BCW mode are not equal. In some embodiments of method 2410, the weighting factors are indicated in a field in the bitstream representation of the current video block. In some embodiments, the DMVD scheme includes a decoder-side motion vector refinement (DMVR) encoding mode that derives refined motion information based on a prediction cost function. In some embodiments, the DMVD scheme includes a bidirectional optical flow (BDOF) encoding mode encoding tool that derives refined predictions based on gradient calculations. In some embodiments of method 2410, the BCW mode used by the current video block includes a field using an index representing a BCW index and a weighting factor, and the BCW index is not equal to 0.
[0398] In some embodiments of method 2410, the encoding tool includes a weighted prediction mode. In some embodiments of method 2410, the weighted prediction mode used by the current video block includes applying weighted prediction to at least one prediction block of the current video block. In some embodiments of method 2410, the encoding tool includes a local illumination compensation (LIC) mode. In some embodiments of method 2410, the encoding tool includes a multi-hypothesis prediction mode. In some embodiments of method 2410, a first weight parameter for a first reference picture and a second weight parameter for a second reference picture are associated with the weighted prediction mode for the current video block, and in response to the first weight parameter being different from the second weight parameter, it is determined that the DMVD scheme is disabled for the current video block.
[0399] In some embodiments of method 2410, the first weight parameter and / or the second weight parameter are indicated in a field in a bitstream representation of a video unit comprising the current video block, the video unit comprising at least one of a picture or a slice. In some embodiments of method 2410, the first linear model parameter is used for a first reference picture of the current video block, and the second linear model parameter is used for a second reference picture of the current video block, and in response to the first linear model parameter being different from the second linear model parameter, it is determined that the DMVD scheme is disabled for the current video block.
[0400] Figure 24C 24 is a flow chart of an example method 2420 for video processing. The method 2420 includes, at 2422, determining whether to enable or disable one or more decoder-side motion vector derivation (DMVD) schemes for the current video block based on picture order count (POC) values of one or more reference pictures of the current video block of the video and a POC value of a current picture that includes the current video block. The method 2420 includes, at 2424, performing conversion between the current video block and a bitstream representation of the video based on the determination.
[0401] In some embodiments of the method 2420, determining whether to enable or disable one or more DMVD schemes is based on a relationship between a first POC distance (PocDis0) representing a first distance from a first reference picture of the current video block to the current picture and a second POC distance (PocDis1) representing a second distance from the current picture to a second reference picture of the current video block. In some embodiments of the method 2420, the first reference picture is a reference picture list 0 of the current video block, and the second reference picture is a reference picture list 1 of the current video block.
[0402] In some embodiments of the method 2420, PocDist0 is set to a first POC value of the current picture minus a second POC value of the first reference picture, and PocDist1 is set to a third POC value of the second reference picture minus the first POC value of the current picture. In some embodiments of the method 2420, in response to PocDis0 not being equal to PocDis1, one or more DMVD schemes are enabled. In some embodiments of the method 2420, determining whether to enable or disable one or more of the one or more DMVD schemes is based on applying the same rules to PocDis0 and PocDis1. In some embodiments of the method 2420, in response to PocDis0 being equal to PocDis1, the one or more DMVD schemes are enabled.
[0403] In some embodiments of the method 2420, in response to PocDis0 multiplied by PocDis1 being less than zero, one or more DMVD schemes are disabled. In some embodiments of the method 2420, in response to PocDis0 not being equal to PocDis1, one or more DMVD schemes are disabled. In some embodiments of the method 2420, during conversion, the one or more DMVD schemes use a first set of motion vector differences (MVDs) for a first reference picture list and a second set of MVDs for a second reference picture list to identify two reference blocks, the first MVD set being a mirrored version of the second MVD set. In some embodiments of the method 2420, during conversion, the one or more DMVD schemes use a first set of motion vector differences (MVDs) for a first reference picture list and a second set of MVDs for a second reference picture list to identify two reference blocks, the second MVD set being a mirrored version of the first MVD set.
[0404] In some embodiments of method 2420, during conversion, one or more DMVD schemes use a first motion vector difference (MVD) set for a first reference picture list and a second MVD set for a second reference picture list to identify two reference blocks, the first MVD set being a non-mirrored version of the second MVD set. In some embodiments of method 2420, the first MVD set is scaled to the second MVD set based on PocDis0 and PocDis1. In some embodiments of method 2420, the first MVD set comprising (MvdX, MvdY) is scaled to the second MVD set calculated as follows: (-MvdX*PocDis1 / PocDis0, -MvdY*PocDis1 / PocDis0). In some embodiments of method 2420, the second MVD set is scaled to the first MVD set based on PocDis0 and PocDis1. In some embodiments of method 2420, the second set of MVDs comprising (MvdX, MvdY) is scaled to the first set of MVDs calculated as follows: (-MvdX*PocDis0 / PocDis1, -MvdY*PocDis0 / PocDis1).
[0405] In some embodiments of method 2420, a division operation of the scaling operation is implemented using a lookup table, a multiplication operation, or a right-to-right operation. In some embodiments of method 2420, one or more DMVD schemes determine, during a DMVD process, a first set of motion vector differences (MVDs) for a first reference picture list and a second set of MVDs for a second reference picture list for a current video block of the video based on a POC value of a reference picture for the current video block and a POC value of a current picture that includes the current video block. In some embodiments of method 2420, in response to a first absolute value of PocDis0 being less than or equal to a second absolute value of PocDis1, the first set of MVDs is scaled according to PocDis0 and PocDis1 to generate a second set of MVDs. In some embodiments of method 2420, in response to the first absolute value of PocDis0 being greater than or equal to the second absolute value of PocDis1, the first set of MVDs is scaled according to PocDis0 and PocDis1 to generate a second set of MVDs.
[0406] In some embodiments of method 2420, in response to the second absolute value of PocDis1 being less than or equal to the first absolute value of PocDis0, the second MVD set is scaled according to PocDis0 and PocDis1 to generate the first MVD set. In some embodiments of method 2420, in response to the second absolute value of PocDis1 being greater than or equal to the first absolute value of PocDis0, the second MVD set is scaled according to PocDis0 and PocDis1 to generate the first MVD set. In some embodiments of method 2420, in response to the two reference pictures including a first reference picture preceding a current picture and a second reference picture following the current picture, the first MVD set is mirrored to generate the second MVD set, and no scaling is performed to obtain the first MVD set or the second MVD set. In some embodiments of method 2420, determining whether to enable or disable one or more DMVD schemes is based on a first sign of a first picture order count distance (PocDis0) representing a first distance from a first reference picture of the current video block to the current picture and a second sign of a second picture order count distance (PocDis1) representing a second distance from the current picture to a second reference picture of the current video block.
[0407] In some embodiments of the method 2420, in response to a result of multiplying PocDis0 having a first sign by PocDis1 having a second sign being less than zero, one or more DMVD schemes are enabled. In some embodiments of the method 2420, in response to a result of multiplying PocDis0 having a first sign by PocDis1 having a second sign being greater than zero, one or more DMVD schemes are enabled. In some embodiments of the method 2420, in response to a first picture order count distance (PocDis0) representing a first distance from a first reference picture of the current video block to the current picture being not equal to a second picture order count distance (PocDis1) representing a second distance from the current picture to a second reference picture of the current video block, one or more DMVD schemes are disabled.
[0408] In some embodiments of method 2420, a first MVD set is used to refine motion information for a first reference picture list, and a second MVD set is used to refine motion information for a second reference picture list. In some embodiments of method 2420, the first reference picture list is reference picture list 0, and the second reference picture list is reference picture list 1.
[0409] Figure 24D24 is a flow chart of an example method 2430 for video processing. The method 2430 includes, at 2432, obtaining refined motion information for a current video block of the video by performing a decoder-side motion vector derivation (DMVD) scheme on the current video block of the video, wherein a symmetric motion vector difference (SMVD) mode is enabled for the current video block. The method 2430 includes, at 2434, performing a conversion between the current video block and a bitstream representation of the video using the refined motion information.
[0410] In some embodiments of the method 2430, the bitstream representation includes motion vector differences (MVDs) for refined motion information, and the MVDs are decoded according to an SMVD mode and further refined before being used to decode the current video block. In some embodiments of the method 2430, in the SMVD mode, a DMVD scheme is used to refine the motion vector differences (MVDs) for the refined motion information by changing the motion vector (MV) precision or the MVD precision from N pixel precision to M pixel precision, where N and M are equal to 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16. In some embodiments of the method 2430, M is less than or equal to N. In some embodiments of the method 2430, the bitstream representation does not include signaling of motion vector differences (MVDs) for refined motion information in the SMVD mode, and the MVDs are generated using the DMVD scheme.
[0411] In some embodiments of method 2430, adaptive motion vector difference resolution (AMVR) information is not signaled in the bitstream representation of video blocks encoded in SMVD mode, and motion vector (MV) accuracy or motion vector difference (MVD) accuracy of the refined motion information is derived from a predefined value. In some embodiments of method 2430, the MV accuracy or MVD accuracy is 1 / 4 pixel accuracy. In some embodiments of method 2430, the predefined value is signaled in the bitstream representation at the sequence, picture, slice group, slice, or video data unit level. In some embodiments of method 2430, the predefined value depends on mode information or motion information. In some embodiments of method 2430, the mode information or motion information includes affine motion information or non-affine motion information.
[0412] Figure 24E24 is a flow chart of an example method 2440 for video processing. The method 2440 includes, at 2442, determining whether a decoder-side motion vector derivation (DMVD) scheme is enabled or disabled for the current video block based on a field in a bitstream representation of a video including the current video block, and enabling a symmetric motion vector difference (SMVD) mode for the current video block. The method 2440 includes, at 2444, obtaining refined motion information for the current video block by applying the DMVD scheme to the current video block after determining that the DMVD scheme is enabled. The method 2440 includes, at 2446, performing a conversion between the current video block and the bitstream representation of the video using the refined motion information.
[0413] In some embodiments of the method 2440, in response to the DMVD scheme being enabled, a motion vector difference (MVD) is not signaled in the bitstream representation. In some embodiments of the method 2440, for one or more motion vector (MV) precisions or motion vector difference (MVD) precisions, a field is present in the bitstream representation indicating whether the DMVD scheme is enabled or disabled. In some embodiments of the method 2440, the one or more MV precisions or MVD precisions include 1 pixel precision and / or 4 pixel precision.
[0414] In some embodiments of method 2440, in response to a first picture order count distance (PocDis0) representing a first distance from a first reference picture (Ref0) of a current video block to the current picture being equal to a second picture order count distance (PocDis1) representing a second distance from the current picture to a second reference picture (Ref1) of the current video block, there is a field in the bitstream representation indicating whether the DMVD scheme is enabled or disabled, and the first reference picture (Ref0) precedes the current picture and the second reference picture (Ref1) follows the current picture in display order.
[0415] In some embodiments of method 2440, in response to a first picture order count distance (PocDis0) representing a first distance from a first reference picture (Ref0) of a current video block to the current picture being equal to a second picture order count distance (PocDis1) representing a second distance from the current picture to a second reference picture (Ref1) of the current video block, there is a field in the bitstream representation indicating whether the DMVD scheme is enabled or disabled, and the second reference picture (Ref1) precedes the current picture and the first reference picture (Ref0) follows the current picture in display order.
[0416] In some embodiments of method 2440, a DMVD scheme is enabled in SMVD mode based on coding information for a current video block and / or one or more neighboring blocks. In some embodiments of method 2440, the DMVD scheme is enabled in SMVD mode based on the block dimensions of the current video block. In some embodiments of method 2440, the DMVD scheme is enabled in SMVD mode based on information related to reference pictures for the current video block. In some embodiments of method 2440, the information related to reference pictures includes picture order count (POC) information. In some embodiments of method 2440, the DMVD scheme is enabled in SMVD mode based on signaling of motion vector difference (MVD) information in a bitstream representation. In some embodiments of method 2420, one or more DMVD schemes include a decoder-side motion vector refinement (DMVR) scheme. In some embodiments of methods 2430 and 2440, the DMVD scheme includes a decoder-side motion vector refinement (DMVR) scheme. In some embodiments of method 2430, one or more DMVD schemes include a bidirectional optical flow (BDOF) scheme. In some embodiments of methods 2430 and 2440 , the DMVD scheme includes a bidirectional optical flow (BDOF) scheme.
[0417] Figure 24F 24 is a flow chart of an example method 2450 for video processing. The method 2450 includes, at 2452, determining whether to enable or disable multiple decoder-side motion vector derivation (DMVD) schemes for conversion between a current video block and a bitstream representation of the video based on rules using block dimensions of the current video block of the video. The method 2450 includes, at 2454, performing the conversion based on the determination.
[0418] In some embodiments of method 2450, in response to (W*H)>=T1 and H>=T2, enabling multiple DMVD schemes is determined, where W and H are the width and height of the current video block, respectively, and T1 and T2 are rational numbers. In some embodiments of method 2450, T1 is 64 and T2 is 8. In some embodiments of method 2450, in response to H>=T1 and W is not equal to T2 or H is not equal to T1, enabling multiple DMVD schemes is determined, where W and H are the width and height of the current video block, respectively, and T1 and T2 are rational numbers. In some embodiments of method 2450, T1 is 8 and T2 is 4.
[0419] In some embodiments of method 2450, disabling the multiple DMVD schemes is determined in response to a first sample number of the current video block being less than a second sample number. In some embodiments of method 2450, disabling the multiple DMVD schemes is determined in response to a first sample number of the current video block being greater than a second sample number. In some embodiments of method 2450, the second sample number is 16 luma samples, 32 luma samples, 64 luma samples, or 128 luma samples. In some embodiments of method 2450, disabling the multiple DMVD schemes is determined in response to a width of the current video block being less than a value.
[0420] In some embodiments of method 2450, disabling the multiple DMVD schemes is determined in response to the height of the current video block being less than a value. In some embodiments of method 2450, the value is 8. In some embodiments of method 2450, disabling the multiple DMVD schemes is determined in response to the width of the current video block being greater than or equal to a first threshold and / or in response to the height of the current video block being greater than or equal to a second threshold. In some embodiments of method 2450, the width is 128 and the height is 128. In some embodiments of method 2450, the width is greater than or equal to 64 and the height is 128, or the width is 128 and the height is greater than or equal to 64. In some embodiments of method 2450, the width is greater than or equal to 4 and the height is 128, or the width is 128 and the height is greater than or equal to 4. In some embodiments of method 2450, the first threshold and the second threshold are 64.
[0421] In some embodiments of method 2450, disabling the plurality of DMVD schemes is determined in response to a width of the current video block being less than or equal to a first threshold and / or in response to a height of the current video block being less than or equal to a second threshold. In some embodiments of method 2450, the first threshold and the second threshold are both 8. In some embodiments, the plurality of DMVD schemes include a decoder-side motion vector refinement (DMVR) scheme that refines motion information based on a cost function. In some embodiments, the plurality of DMVD schemes include a bidirectional optical flow (BDOF) scheme that refines motion information based on a gradient calculation.
[0422] Figure 24G 2460 is a flow chart of an example method 2460 for video processing. The method 2460 includes, at 2462, determining whether to perform multiple decoder-side motion vector derivation (DMVD) schemes at a sub-block level or a block level for a current video block of a video. The method 2460 includes, at 2464, after determining to perform the multiple DMVD schemes at the sub-block level, obtaining refined motion information for the current video block by implementing the multiple DMVD schemes at the same sub-block level of the current video block. The method 2460 includes, at 2466, performing conversion between the current video block and a bitstream representation of the video using the refined motion information.
[0423] In some embodiments of method 2460, the plurality of DMVD schemes include a decoder-side motion vector refinement (DMVR) scheme. In some embodiments of method 2460, the refined motion information is obtained by applying bilateral matching in the DMVR scheme at a sub-block level of the current video block. In some embodiments of method 2460, the plurality of DMVD schemes include a bidirectional optical flow (BDOF) encoding scheme. In some embodiments of method 2460, a determination is made to enable or disable the BDOF encoding scheme at a sub-block level of the current video block. In some embodiments of method 2460, a determination is made to enable the BDOF encoding scheme, and the refined motion information is obtained by performing a sample-wise refinement of the motion information in the BDOF encoding scheme, which is performed at a sub-block level of the current video block.
[0424] In some embodiments of method 2460, determining whether to enable or disable a BDOF encoding scheme at a sub-block level of the current video block, and determining to perform a sample-based motion information refinement process in the BDOF encoding scheme at a sub-block level of the current video block. In some embodiments of method 2460, a width and a height of the sub-block are both equal to 16. In some embodiments of method 2460, dividing the current video block into the plurality of sub-blocks is responsive to: a first width of the current video block being greater than or equal to a value, or a first height of the current video block being greater than or equal to the value, or both the first width being greater than or equal to the value and the first height being greater than or equal to the value.
[0425] In some embodiments of method 2460, each of the plurality of sub-blocks is processed by one or more DMVD schemes in the same manner as a coding block having a size equal to the sub-block size. In some embodiments of method 2460, the value is 64, and in response to the current video block having a first width of 64 and a first height of 128 or a first width of 128 and a first height of 64, the current video block is divided into two sub-blocks, wherein each of the two sub-blocks has a second width and a second height of 64. In some embodiments of method 2460, the value is 64, and in response to the current video block having a first width and a first height of 128, the current video block is divided into four sub-blocks, wherein each of the two sub-blocks has a second width and a second height of 64.
[0426] In some embodiments of method 2460, in response to the current video block having a first width of N and a first height of 128 or a first width of 128 and a first height of N, the current video block is not divided into sub-blocks, where N is less than 64. In some embodiments of method 2460, the value is 64, and in response to the current video block having a first width of N and a first height of 128 or a first width of 128 and a first height of N, where N is less than 64, the current video block is divided into two sub-blocks, where each of the two sub-blocks has a second width of N and a second height of 64 or a second width of 64 and a second height of N.
[0427] In some embodiments of method 2460, in response to a first width of the current video block being greater than a value, the current video block is split vertically, and the second width of the sub-blocks of the current video block is less than or equal to the value. In some embodiments of method 2460, in response to a first height of the current video block being greater than a value, the current video block is split horizontally, and the second height of the sub-blocks of the current video block is less than or equal to the value. In some embodiments of method 2460, the value is 16. In some embodiments of method 2460, the second width of the sub-blocks of the current video block is 16. In some embodiments of method 2460, the second height of the sub-blocks of the current video block is 16. In some embodiments of method 2460, in response to a first size of the current video block being greater than a first threshold, the current video block is split into a plurality of sub-blocks. In some embodiments of method 2460, each of the plurality of sub-blocks is processed using one or more DMVD schemes in the same manner as a coding block having a second size equal to the sub-block size.
[0428] In some embodiments of method 2460, each of the plurality of sub-blocks has the same size that is less than or equal to a first threshold. In some embodiments of methods 2450 and 2460, the current video block is a luma video block. In some embodiments of method 2450, the determination of whether to enable or disable multiple DMVD schemes is performed for the luma video block and is shared by the associated chroma video blocks. In some embodiments of method 2460, the determination of whether to implement multiple DMVD schemes at the sub-block level is performed for the luma video block and is shared by the associated chroma video blocks. In some embodiments of method 2460, the determination not to split the current video block horizontally or vertically into the plurality of sub-blocks is made in response to the height or width of the current video block being less than or equal to a second threshold. In some embodiments of method 2460, the first threshold is 1024 and the second threshold is 32.
[0429] In some embodiments of method 2460, the value is predefined or signaled at the sequence parameter set (SPS), picture parameter set (PPS), picture, slice, slice group, or slice level for the current video block. In some embodiments of method 2460, the value, the first threshold, or the second threshold depends on coding information for the current video block. In some embodiments of method 2460, the determination of the sub-block size is the same for multiple DMVD schemes. In some embodiments of method 2460, the coding information for the current video block includes a block size, a picture type, or a temporal layer index for the current video block. In some embodiments of methods 2450 and 2460, the multiple DMVDs for the current video block include all DMVD schemes for the current video block.
[0430] Figure 24H 24 is a flow chart of an example method 2470 for video processing. The method 2470 includes, at 2472, determining whether a decoder-side motion vector derivation (DMVD) scheme is enabled or disabled for a plurality of components of a current video block of a video. The method 2470 includes, at 2474, after determining that the DMVD scheme is enabled, obtaining refined motion information for the current video block by implementing the DMVD scheme. The method 2470 includes, at 2476, performing conversion between the current video block and a bitstream representation of the video during implementation of the DMVD scheme.
[0431] In some embodiments of method 2470, the determination of whether to enable or disable the DMVD scheme is performed once and shared by multiple components. In some embodiments of method 2470, the determination of whether to enable or disable DMVD is performed multiple times for multiple components. In some embodiments of method 2470, the determination of whether to enable or disable DMVD is performed first for one of the multiple components, and then performed or shared with one or more remaining components of the multiple components. In some embodiments of method 2470, the one component is a luma component or a green component. In some embodiments of method 2470, the determination of whether to enable or disable DMVD is performed for the one component based on information about the one component of the multiple components. In some embodiments of method 2470, the one component is a luma component, a chroma component, a green component, a blue component, or a red component.
[0432] Some embodiments may be described using the following clause-based format.
[0433] Clause 1. A video processing method, comprising: implementing, by a processor, a decoder-side motion vector derivation (DMVD) scheme for performing motion vector refinement by deriving parameters based on derivation rules during conversion between a current video block and a bitstream representation of the current video block.
[0434] Clause 2. The technique of clause 1, wherein the parameter is derived from parameters of a final prediction block applied to the current video block.
[0435] Clause 3. The technique of clause 1, wherein the parameter is signaled in the bitstream representation.
[0436] Clause 4. The technique of clause 1, wherein the parameter is derived by a processor.
[0437] Clause 5. The technique of any of clauses 1-4, wherein the derivation rule specifies using parameters for deriving a final prediction block for a DMVD scheme.
[0438] Clause 6. The technology of clause 5, wherein the conversion comprises calculating a prediction cost function for the current video block by first applying one of generalized bidirectional coding weights or weights of a weighted prediction scheme or weights of a local illumination compensation scheme, or a temporal or spatial gradient of a bidirectional optical flow scheme, and then calculating the prediction cost function.
[0439] Clause 7. The technique of clause 6, wherein the prediction cost function is a gradient function or a sum of absolute difference (SAD) cost function.
[0440] Clause 8. The technique of clause 2, wherein the parameter is a parameter of local illumination compensation for the final prediction block.
[0441] Clause 9. A video processing method, comprising: selectively using a decoder-side motion vector derivation (DMVD) scheme for motion vector refinement during conversion between a current video block and a bitstream representation of the current video block based on an activation rule.
[0442] Clause 10. The technique of clause 9, wherein the enabling rule specifies disabling the DMVD scheme if the conversion uses a generalized bidirectional coding mode, or a local illumination compensation mode, or a weighted prediction mode, or a multi-hypothesis prediction mode.
[0443] Clause 11. The technique of clause 9, wherein the enabling rule specifies use of a DMVD scheme for a current video block that is a bidirectionally predicted block using unequal picture order count distances.
[0444] Clause 12. The technique of clause 9, wherein the enabling rule specifies use of the DMVD scheme based on a relationship between picture order count distances PocDis0 and PocDis1 representing two directions of bidirectional prediction of the current video block.
[0445] Clause 13. The technique of clause 12, wherein the enabling rule specifies to use the DMVD scheme if PocDis0 = PocDis1.
[0446] Clause 14. The technique of clause 12, wherein the activation rule specifies use of the DMVD scheme if PocDis0 is not equal to PocDis1.
[0447] Clause 15. The technique of clause 12, wherein the activation rule specifies use of the DMVD scheme if PocDis0 multiplied by PocDis1 is less than zero.
[0448] Clause 16. The technique of any of clauses 9-14, wherein during conversion, the DMVD scheme uses List 0 and List 1 as two reference picture lists, and wherein List 0 is a mirrored version of List 1.
[0449] Clause 17. The technique of clause 15, wherein the DMVD scheme comprises using motion vector differences of list 0 and list 1 according to a scaling based on a distance of PocDis0 and PocDis1.
[0450] Clause 18. The technique of clause 17, wherein the DMVD scheme comprises using motion vector differences of list 0 scaled to motion vector differences of list 1.
[0451] Clause 19. The technique of clause 17, wherein the DMVD scheme comprises using motion vector differences of list 1 scaled to motion vector differences of list 0.
[0452] Clause 20. The technique of any of clauses 9-14, wherein the DMVD scheme comprises using reference pictures according to their picture order counts.
[0453] Clause 21. The technique of clause 9, wherein the enabling rule is based on a dimension of the current video block.
[0454] Clause 22. The technology of clause 21, wherein the DMVD scheme includes decoder-side motion vector refinement (DMVR) enabled if W*H>=T1&&H>=T2, where W and H are the width and height of the current video block, and T1 and T2 are rational numbers.
[0455] Clause 23. The technology of clause 21, wherein the DMVD scheme includes a bidirectional optical flow (BIO) encoding method enabled when W*H>=T1&&H>=T2, where W and H are the width and height of the current video block, and T1 and T2 are rational numbers.
[0456] Clause 24. The technique of clause 21, wherein the DMVD scheme includes decoder-side motion vector refinement (DMVR) enabled if H>=T1 &&!(W==T2 &&H==T1), where W and H are the width and height of the current video block, and T1 and T2 are rational numbers.
[0457] Clause 25. The technique of clause 21, wherein the DMVD scheme comprises a bidirectional optical flow (BIO) coding scheme enabled if H>=T1 &&!(W==T2 &&H==T1), where W and H are the width and height of the current video block, and T1 and T2 are rational numbers.
[0458] Clause 26. The technology of any one of clauses 9 to 21, wherein the DMVD scheme is a decoder-side motion vector refinement (DMVR) scheme or a bidirectional optical flow (BIO) encoding scheme, and wherein the DMVD scheme is disabled if the width of the current video block is > th1 or the height is > th2.
[0459] Clause 27. A video processing technique comprising: selectively using a decoder-side motion vector derivation (DMVD) scheme for motion vector refinement during conversion between a current video block and a bitstream representation of the current video block based on a rule by applying the decoder-side motion vector derivation (DMVD) scheme at a sub-block level.
[0460] Clause 28. The technique of clause 27, wherein the DMVD scheme is a decoder-side motion vector refinement (DMVR) scheme or a bidirectional optical flow (BIO) scheme.
[0461] Clause 29. The technique of clause 28, wherein the DMVD scheme is a BIO scheme, and wherein the rule specifies that the DMVD scheme is based on applicability on a sub-block-by-sub-block basis.
[0462] Clause 30. The technology of clause 29, wherein the width of the current video block is greater than or equal to LW or the height is greater than or equal to LH, or the width*height is greater than a threshold L1, wherein L1, L, W, and H are integers, and wherein the conversion is performed by dividing the current video block into a plurality of sub-blocks, which are further processed using a DMVD scheme.
[0463] Clause 31. The technique of clause 30, wherein partitioning comprises partitioning the current video block horizontally.
[0464] Clause 32. The technique of clause 30, wherein partitioning comprises vertically partitioning the current video block.
[0465] Clause 33. A technique according to any of clauses 30-32, wherein L is signaled in the bitstream representation at a sequence parameter set level, a picture parameter set level, a picture level, a slice level, a slice group level, or a slice level, or wherein L is implicitly signaled based on the size of the current video block or the type of the picture containing the current video block or the temporal layer index of the current video block.
[0466] Clause 34. The technique of any of clauses 1-33, wherein DMVD is applied to the current video block depending on the luma or chroma type of the current video block.
[0467] Clause 35. The technique of any of clauses 1-34, wherein converting uses a DMVD scheme determined based on a decision to use DMVD for the block or a different luma or chroma type corresponding to the current video block.
[0468] Clause 36. The technique of any of clauses 1-35, wherein the DMVD scheme comprises a decoder-side motion vector refinement scheme or a bidirectional optical flow scheme.
[0469] Clause 37. A video processing technique comprising: during a conversion between a current video block and a bitstream representation of the current video block, wherein the current video block uses a symmetric motion vector difference codec technique, using a decoder-side motion vector derivation technique by which a motion vector of the current video block is refined during the conversion, wherein the symmetric motion vector difference codec technique uses symmetric motion vector difference derivation; and performing the conversion using the decoder-side motion vector derivation technique.
[0470] Clause 38. The technique of clause 37, wherein the decoder-side motion vector derivation technique comprises decoder-side motion vector refinement.
[0471] Clause 39. A technique according to any of clauses 37-38, wherein the decoder-side motion vector derivation technique changes the motion vector precision from N pixels used for a symmetric motion vector difference codec technique to M pixel precision, wherein N and M are fractional or integer numbers, and wherein N and M are equal to 1 / 16, 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, or 16.
[0472] Clause 40. The technology of clause 39, wherein M is less than or equal to N.
[0473] Clause 41. The technique of any of clauses 37-41, wherein the bitstream representation does not include a motion vector difference indication for the current video block, and wherein a decoder-side motion vector derivation technique is used to derive the motion vector difference.
[0474] Clause 42. The technique of any of clauses 37-42, wherein the bitstream representation indicates whether a decoder-side motion vector derivation technique and a symmetric motion vector derivation technique are used for transformation of the current video block.
[0475] Clause 43. The technique of any of clauses 1-42, wherein converting comprises generating a bitstream representation from the current video block or generating the current video block from the bitstream representation.
[0476] Clause 44. A video encoding apparatus comprising a processor configured to implement the method of one or more of clauses 1 to 43.
[0477] Clause 45. A video decoding device comprising a processor configured to implement the method of one or more of clauses 1 to 43.
[0478] Clause 46. A computer-readable medium having code stored thereon, the code, when executed, causing a processor to perform the method of any one or more of clauses 1 to 43.
[0479] Figure 26 is a block diagram illustrating an example video processing system 2100 in which the various techniques disclosed herein may be implemented. Various embodiments may include some or all of the components of system 2100. System 2100 may include an input 2102 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10 bit multi-component pixel values), or may be received in a compressed or encoded format. Input 2102 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, a passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0480] System 2100 may include a coding component 2104 that can implement the various encoding or coding methods described in this document. Coding component 2104 can reduce the average bit rate of the video from input 2102 to the output of coding component 2104 to generate the encoded representation of the video. Therefore, coding technology is sometimes referred to as video compression or video transcoding technology. The output of coding component 2104 can be stored or transmitted via connected communication, as shown in component 2106. Component 2108 can use the bitstream (or encoding) representation of the stored or transmitted video received at input end 2102 to generate pixel values or displayable video sent to display interface 2110. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "encoding" operations or tools, it should be understood that the encoding tools or operations are used at the encoder, and the corresponding decoding tools or operations opposite to the encoded results will be performed by the decoder.
[0481] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0482] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of blocks of video, but will not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from blocks of video to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to blocks of video will be performed using the video processing tool or mode enabled based on the decision or determination.
[0483] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when the video processing tool or mode is disabled, the encoder will not use the tool or mode in converting blocks of video to a bitstream representation of the video. In another example, when the video processing tool or mode is disabled, the decoder will process the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was disabled based on the decision or determination.
[0484] Figure 27 is a block diagram illustrating an example video encoding system 100 that may utilize the techniques of this disclosure. Figure 27 As shown, video encoding system 100 may include source device 110 and destination device 120. Source device 110, which generates encoded video data, may be referred to as a video encoding device. Destination device 120, which may decode the encoded video data generated by source device 110, may be referred to as a video decoding device. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0485] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the network 130a via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130b for access by the destination device 120.
[0486] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .
[0487] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, with destination device 120 configured to interface with an external display device.
[0488] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVM) standard, and other current and / or future standards.
[0489] Figure 28 is a block diagram illustrating an example of a video encoder 200, which may be Figure 27 The video encoder 114 in the system 100 is shown.
[0490] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 28 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0491] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0492] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode where at least one reference picture is the picture in which the current video block is located.
[0493] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated, but for the purpose of explanation, they are not shown in FIG. Figure 28 In the example, they are represented separately.
[0494] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0495] The mode selection unit 203 can, for example, select one of the intra or inter coding modes based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra and inter prediction (CIIP) mode, in which prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution (e.g., sub-pixel or integer pixel precision) for the motion vector of the block.
[0496] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.
[0497] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.
[0498] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.
[0499] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures containing the reference video blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0500] In some examples, motion estimation unit 204 may output a complete motion information set for use in a decoding process by a decoder.
[0501] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information for the current video block. For example, motion estimation unit 204 may determine that the motion information for the current video block is sufficiently similar to the motion information for an adjacent video block.
[0502] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0503] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0504] As described above, the video encoder 200 may predictively signal motion vectors. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Predication (AMVP) and Merge mode signaling.
[0505] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0506] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0507] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0508] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0509] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0510] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.
[0511] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0512] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0513] Figure 29 is a block diagram illustrating an example of a video decoder 300, which may be Figure 27 The video decoder 114 in the system 100 is shown.
[0514] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 29 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0515] exist Figure 29 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those generally performed for the video encoder 200 ( Figure 28 ) is the reverse of the encoding process described.
[0516] The entropy decoding unit 301 can retrieve a coded bitstream. The coded bitstream can include entropy-coded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine this information, for example, by implementing AMVP and Merge modes.
[0517] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0518] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used during encoding of the video block by video encoder 20. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.
[0519] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the coded video sequence.
[0520] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.
[0521] The reconstruction unit 306 can sum the residual block with the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.
[0522] From the foregoing, it will be understood that specific embodiments of the presently disclosed technology have been described herein for illustrative purposes, but that various modifications may be made without departing from the scope of the present invention. Accordingly, the presently disclosed technology is not to be limited, except as in the appended claims.
[0523] The disclosed and other aspects, examples, embodiments, modules, and functional operations described in this document may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, for execution by a data processing apparatus or to control the operation of the data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that implements a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an operating environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.
[0524] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one or more computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0525] The processes and logic flows described herein can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0526] By way of example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Non-transitory computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.
[0527] Although this patent document contains many details, these should not be construed as limitations on any subject matter or the scope of the claims, but rather as descriptions of features unique to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments, alone or in any suitable subcombination. Furthermore, although the features described above may be described as functioning in certain combinations, and even initially claimed to be so protected, in some cases one or more features in the combination may be deleted from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.
[0528] Similarly, while operations may be depicted in a particular order in the drawings, this should not be understood as requiring that these operations be performed in the particular order shown, or in sequential order, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0529] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A video processing method, comprising: obtaining refined motion information for a current video block of the video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least the weight parameters, wherein, during the generation of the final prediction block of the current video block, the weight parameters are applied to the prediction block; and performing a conversion between the current video block and a bitstream representation of a video using at least the refined motion information and the weight parameters, wherein the indication of the weight parameter is transmitted by signal using a generalized bidirectional prediction (GBI) technique or a weighted prediction technique, The weight parameters are derived using the local illumination compensation (LIC) technique. The weight parameter is associated with a shaping process of mapping a set of sample values to another set of sample values. The shaping process is disabled for a prediction block generated using the refined motion information of the current video block.
2. The method according to claim 1, wherein A field in the bitstream representation indicates the weight parameter.
3. The method according to claim 1, wherein The weight parameters are derived.
4. The method according to claim 1, wherein The DMVD scheme is implemented by applying the weight parameters to a prediction block of the current video block.
5. The method according to claim 4, wherein The converting includes calculating a prediction cost function for the current video block by first applying a weight parameter according to a bidirectional prediction BCW index having a coding unit weight to the prediction block and then calculating a prediction cost function.
6. The method according to claim 4, wherein: The converting includes calculating a prediction cost function for the current video block by first applying weight parameters according to a weighted prediction scheme to the prediction block and then calculating a prediction cost function.
7. The method according to claim 4, wherein: The converting includes calculating a prediction cost function for the current video block by first applying weight parameters according to a local illumination compensation (LIC) scheme to the prediction block and then calculating a prediction cost function.
8. The method according to any one of claims 5 to 7, wherein The prediction cost function is a gradient function.
9. The method according to any one of claims 5 to 7, wherein: The prediction cost functions are absolute difference and SAD cost functions.
10. The method according to any one of claims 5 to 7, wherein The prediction cost functions are the mean of absolute differences removal and MR-SAD cost functions.
11. The method according to claim 4, wherein The conversion includes calculating the temporal gradient and spatial gradient of the bidirectional optical flow (BIO) scheme for the current video block by first applying a weight parameter according to a bidirectional prediction BCW index with coding unit weights to the prediction block, and then calculating the temporal gradient and spatial gradient.
12. The method according to claim 4, wherein: The conversion includes calculating temporal gradients and spatial gradients of a bidirectional optical flow (BIO) scheme for the current video block by first applying weight parameters according to a weighted prediction scheme to the prediction block and then calculating temporal gradients and spatial gradients.
13. The method according to claim 4, wherein: The conversion includes calculating temporal gradient and spatial gradient of a bidirectional optical flow (BIO) scheme for the current video block by first applying weight parameters according to a local illumination compensation (LIC) scheme to the prediction block and then calculating temporal gradient and spatial gradient.
14. The method according to any one of claims 5 to 7, wherein: The prediction cost function or the computation of the temporal gradient or the spatial gradient is performed in the shaped domain.
15. The method according to claim 1, wherein The converting includes encoding the current video block into the bitstream.
16. The method according to claim 1, wherein The converting includes decoding the bitstream into the current video block.
17. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein: When executed by the processor, the instructions cause the processor to: obtaining refined motion information for a current video block of the video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least the weight parameters, wherein, during the generation of the final prediction block of the current video block, the weight parameters are applied to the prediction block; and performing a conversion between the current video block and a bitstream representation of a video using at least the refined motion information and the weight parameters, wherein the indication of the weight parameter is transmitted by signal using a generalized bidirectional prediction (GBI) technique or a weighted prediction technique, The weight parameters are derived using the local illumination compensation (LIC) technique. The weight parameter is associated with a shaping process of mapping a set of sample values to another set of sample values. The shaping process is disabled for a prediction block generated using the refined motion information of the current video block.
18. A non-transitory computer-readable storage medium storing instructions that cause a processor to: obtaining refined motion information for a current video block of the video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least the weight parameters, in, During the generation of the final prediction block of the current video block, the weight parameters are applied to the prediction block; as well as performing a conversion between the current video block and a bitstream representation of a video using at least the refined motion information and the weight parameters, wherein the indication of the weight parameter is transmitted by signal using a generalized bidirectional prediction (GBI) technique or a weighted prediction technique, The weight parameters are derived using the local illumination compensation (LIC) technique. The weight parameter is associated with a shaping process of mapping a set of sample values to another set of sample values. The shaping process is disabled for a prediction block generated using the refined motion information of the current video block.
19. A method for storing a video bitstream, comprising: obtaining refined motion information for a current video block of the video by implementing a decoder-side motion vector derivation (DMVD) scheme based on at least the weight parameters, wherein, during the generation of the final prediction block of the current video block, the weight parameters are applied to the prediction block; and generating the bitstream using at least the refined motion information and the weight parameters, storing the bitstream in a non-transitory computer-readable recording medium, wherein the indication of the weight parameter is transmitted by signal using a generalized bidirectional prediction (GBI) technique or a weighted prediction technique, The weight parameters are derived using the local illumination compensation (LIC) technique. The weight parameter is associated with a shaping process of mapping a set of sample values to another set of sample values. The shaping process is disabled for a prediction block generated using the refined motion information of the current video block.
Citation Information
Patent Citations
Decoder-side motion vector derivation
US20180278950A1