Method, device and non-transitory computer-readable storage medium for visual media processing

By adopting sub-block-based motion vector refinement technology in video encoding and decoding, and using affine motion information and reference picture index modifications, the conversion of video blocks and bitstream representations is optimized, and the problems of low encoding and decoding efficiency and high computational complexity in the prior art are solved, and more efficient video encoding and decoding is achieved.

CN115280774BActive Publication Date: 2025-08-19DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080081213.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-02
Filing Date
2020-12-02
Publication Date
2025-08-19
Estimated Expiration
2040-12-02

AI Technical Summary

Technical Problem

When handling motion vector prediction, existing video encoding and decoding technologies have problems such as low encoding and decoding efficiency and high computational complexity. Especially in efficient video encoding and decoding standards such as HEVC and VVC, it is difficult to further optimize.

Method used

Using sub-block-based motion vector refinement technology, the conversion process between video blocks and bitstream representation is optimized by using modifications of affine motion information and reference picture indexes, including determining a subset of affine merge candidates and using integer control point motion vectors to improve the motion vector prediction method.

Benefits of technology

It improves the efficiency and quality of video encoding and decoding, reduces the computational complexity, enhances the flexibility and accuracy of the encoding and decoding process, and is suitable for future video encoding and decoding standards and existing standards such as HEVC.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115280774B_ABST
    Figure CN115280774B_ABST
Patent Text Reader

Abstract

A method of visual media processing, comprising: performing a conversion between a current video block of visual media data and a bitstream representation of the visual media data according to a modification routine, wherein the current video block is encoded and decoded using affine motion information; wherein the modification routine specifies that the affine motion information encoded and decoded in the bitstream representation is modified during decoding using motion vector differences and / or reference picture indices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is an application entering the Chinese national phase based on International Patent Application No. PCT / CN2020 / 133271 filed on December 2, 2020, which claims priority to International Patent Application No. PCT / CN2019 / 122347 filed on December 2, 2019. The entire disclosure of the above application is incorporated by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to video and image encoding / decoding technologies, devices, and systems. Background Art

[0004] Despite advances in video compression, digital video remains the largest consumer of bandwidth on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth requirements for digital video usage are expected to continue to grow. Summary of the Invention

[0005] This document describes various embodiments and techniques in which video encoding or decoding is performed using sub-block based motion vector refinement.

[0006] In one example aspect, a method of visual media processing is disclosed. The method includes performing a conversion between a current video block of visual media data and a bitstream representation of the visual media data according to a modification routine, wherein the current video block is encoded and decoded using affine motion information; wherein the modification routine specifies modifying the affine motion information encoded in the bitstream representation using motion vector differences and / or reference picture indices during decoding.

[0007] In another example aspect, another method of visual media processing is disclosed. The method includes, for a conversion between a current video block of visual media data and a bitstream representation of the visual media data, determining that an affine merge candidate for the current video block is to be modified based on a modification convention; and performing the conversion based on the determination; wherein the modification convention specifies modifying one or more control point motion vectors associated with the current video block, and wherein an integer number of control point motion vectors is used.

[0008] In yet another example aspect, another method of visual media processing is disclosed. The method includes, for converting between a current video block of visual media data and a bitstream representation of the visual media data, determining that a subset of allowed affine merge candidates for the current video block is to be modified based on a modification convention; and performing the conversion based on the determination; wherein a syntax element included in the bitstream representation indicates the modification convention.

[0009] In yet another exemplary aspect, a video encoding and / or decoding apparatus is disclosed, comprising a processor configured to implement the above method.

[0010] In yet another exemplary aspect, a computer-readable medium is disclosed, wherein the computer-readable medium stores processor-executable code embodying one of the above methods.

[0011] These and other aspects are described further throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 Depicts the example derivation process for constructing the merge candidate list.

[0013] Figure 2 An example of the location of spatial merge candidates is shown.

[0014] Figure 3 An example of candidate pairs considering redundancy check for spatial merge candidates is shown.

[0015] Figures 4A-4B Example locations of the second PU for Nx2N and 2NxN partitions are shown.

[0016] Figure 5 It is an illustration of the motion vector scaling of the temporal merge candidate.

[0017] Figure 6 An example of the candidate positions of the time-domain merge candidates C0 and C1 is shown.

[0018] Figure 7 An example of a combined bi-predictive merge candidate is shown.

[0019] Figure 8 The derivation process of motion vector prediction candidates is summarized.

[0020] Figure 9 Diagram showing motion vector scaling of spatial motion vector candidates.

[0021] Figure 10 An example of ATMVP motion prediction for a CU is shown.

[0022] Figure 11 An example of a CU having four sub-blocks (AD) and its neighboring blocks (ad) is shown.

[0023] Figure 12 This is an example flow chart of encoding with different MV precisions.

[0024] Figure 13Shown are: (a) 135-degree division type (divided from the upper left corner to the lower right corner) and (b) 45-degree division mode.

[0025] Figure 14 Example locations of neighboring blocks are shown.

[0026] Figure 15 Neighboring blocks (A and L) used for context selection in TPM flag encoding and decoding are shown.

[0027] Figure 16 Shown are (a) a 4-parameter affine model and (b) a 6-parameter affine model.

[0028] Figure 17 An example of the affine MVF for each sub-block is shown.

[0029] Figure 18 Shown are: (a) a 4-parameter affine model and (b) a 6-parameter affine model.

[0030] Figure 19 The MVP of AF_INTER of the genetic affine candidate is shown.

[0031] Figure 20 The constructed affine candidate AF_INTER MVP is shown.

[0032] Figure 21 Shown are: (a) five neighboring blocks, (b) examples of control point motion vector (CPMV) prediction value derivation.

[0033] Figure 22 Example candidate positions for the affine merge mode are shown.

[0034] Figure 23 An example of spatially neighboring blocks used by ATMVP is shown.

[0035] Figure 24 An example of deriving a sub-CU motion field by applying a motion offset from a spatial neighborhood and scaling the motion information from the corresponding co-located sub-CU is shown.

[0036] Figure 25 Candidate positions for the affine merge mode are shown.

[0037] Figure 26 The modified merge list construction process is shown.

[0038] Figure 27 The sub-block MV VSB and the pixel Δv(i,j) are shown (red arrows).

[0039] Figure 28 is a block diagram of an example of a hardware platform for implementing the methods described in this document.

[0040] Figure 29 is a flow chart of an example method of video processing.

[0041] Figure 30 An example of the final motion vector expression (UMVE) process is shown.

[0042] Figure 31 is an example of a UMVE search point.

[0043] Figure 32 is a block diagram of an example video processing system in which the disclosed technology may be implemented.

[0044] Figure 33 is a block diagram illustrating an example video encoding and decoding system.

[0045] Figure 34 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0046] Figure 35 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0047] Figure 36 is a flow chart of an example method of visual media processing.

[0048] Figure 37 is a flow chart of an example method of visual media processing.

[0049] Figure 38 is a flow chart of an example method of visual media processing. DETAILED DESCRIPTION

[0050] The section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to only that section. In addition, although certain embodiments are described with reference to a general video codec or other specific video codec, the disclosed techniques are also applicable to other video codec technologies. In addition, although some embodiments describe video encoding and decoding steps in detail, it will be understood that the corresponding steps of decoding to undo the encoding and decoding will be implemented by the decoder. In addition, the term video processing includes video encoding and compression, video decoding or decompression, and video transcoding, wherein video pixels are represented from one compressed format to another compressed format or at a different compression bit rate.

[0051] 1. Introduction

[0052] This document is about video codec technology. Specifically, it is about motion vector coding in video codecs. It can be applied to existing video codec standards, such as HEVC, or to the standard to be finalized (Universal Video Codec). It may also be applicable to future video codec standards or video codecs.

[0053] 2. Preliminary Discussion

[0054] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that uses temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM) [3,4]. In April 2018, a Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC.

[0055] The latest version of the VVC draft, Versatile Video Codec (Draft 5), can be found at: phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N1001-v5.zip.

[0056] The latest VVC reference software VTM can be found at the following URL: vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0.

[0057] 2.1. Inter-frame prediction in HEVC / H.265

[0058] Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. Motion vectors can be explicitly encoded or decoded as deltas relative to the predicted value.

[0059] When a CU is coded in skip mode, a PU is associated with the CU and has no significant residual coefficients, no coded motion vector increments or reference picture indices. A merge mode is specified where the motion parameters of the current PU are derived from neighboring PUs, including spatial and temporal candidates. The merge mode can be applied to any inter-predicted PU, not only to skip mode. An alternative to the merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector prediction value), the corresponding reference picture index for each reference picture list and the reference picture list to use are explicitly signaled for each PU. In this disclosure, this mode is referred to as advanced motion vector prediction (AMVP).

[0060] When signaling indicates that one of the two reference picture lists is to be used, a PU is generated from a block of samples. This is called "unidirectional prediction." Unidirectional prediction can be used for both P slices and B slices.

[0061] When signaling indicates that two reference picture lists are to be used, a PU is generated from two sample blocks. This is called "bidirectional prediction." Bidirectional prediction is only applicable to B slices.

[0062] The following text provides details on the inter prediction modes specified in HEVC. The description begins with the merge mode.

[0063] 2.1.1. Reference Image List

[0064] In HEVC, the term inter prediction is used to refer to predictions derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the currently decoded picture. As with H.264 / AVC, a picture can be predicted from multiple reference pictures. Reference pictures used for inter prediction are organized into one or more reference picture lists. A reference index identifies which reference picture in the list can be used to create the prediction signal.

[0065] A single reference picture list List 0 is used for P slices, and two reference picture lists List 0 and List 1 are used for B slices. It should be noted that the reference pictures included in List 0 / 1 may be from past and future pictures in terms of capture / display order.

[0066] 2.1.2.Merge mode

[0067] 2.1.2.1. Merge Mode Candidate Derivation

[0068] When PU is predicted using merge mode, the index pointing to the entry in the merge candidate list is parsed from the bitstream and used to retrieve the motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:

[0069] Step 1: Initial candidate derivation

[0070] Step 1.1: Spatial Candidate Derivation

[0071] Step 1.2: Redundancy check of spatial candidates

[0072] Step 1.3: Time Domain Candidate Derivation

[0073] Step 2: Additional candidate insertions

[0074] Step 2.1: Creation of Bidirectional Prediction Candidates

[0075] Step 2.2: Insertion of zero motion candidates

[0076] These steps are also Figure 1 For spatial merge candidate derivation, at most four merge candidates are selected from the candidates located at five different positions. For temporal merge candidate derivation, at most one merge candidate is selected from two candidates. Since the number of candidates for each PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, the index of the best merge candidate is encoded using truncated unary binarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of the 2N×2N prediction unit.

[0077] The operations associated with the above steps are described in detail below.

[0078] Figure 1 An example derivation process for constructing a merge candidate list is shown.

[0079] 2.1.2.2. Spatial Candidate Derivation

[0080] In the derivation of spatial merge candidates, Figure 2Among the candidates at the positions shown, up to four merge candidates are selected. The order of derivation is A1, B1, B0, A0 and B2. Position B2 is considered only when any PU at position A1, B1, B0, A0 is unavailable (for example, because it belongs to another slice or piece) or is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only candidates with the same motion information are considered. Figure 3 The arrows in are linked pairs, and a candidate is added to the list only if the corresponding candidate for redundancy check does not have the same motion information. Another source of duplicate motion information is the "second PU" associated with a different 2N×2N partition. As an example, Figures 4A-4B The second PU is depicted for the N×2N and 2N×N cases, respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for a codec with only one PU. Similarly, position B1 is not considered when the current PU is partitioned into 2N×N.

[0081] 2.1.2.3. Time Domain Candidate Derivation

[0082] In this step, only one candidate is added to the list. In particular, in the derivation of this time domain merge candidate, the scaled motion vector is derived based on the collocated PU belonging to the picture with the smallest POC difference with the current picture within the given reference picture list. The reference picture list to be used for deriving the collocated PU is explicitly signaled in the slice header. Figure 5 As shown by the dashed line in [ ], the scaled motion vector of the temporal merge candidate is obtained, which is scaled from the motion vector of the collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal merge candidate is set equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors are obtained and combined, one for reference picture list 0 and the other for reference picture list 1, to form the bidirectional prediction merge candidate.

[0083] Figure 5 It is an illustration of the motion vector scaling of the temporal merge candidate.

[0084] In the co-located PU(Y) belonging to the reference frame, the position of the temporal candidate is selected between candidates C0 and C1, such as Figure 6 If the PU at position C0 is not available, is intra-coded, or is outside the current codec tree unit (CTU, also known as LCU, the largest codec unit) row, position C1 is used. Otherwise, position C0 is used when deriving the time domain merge candidate.

[0085] Figure 6 An example of the candidate positions of the time-domain merge candidates C0 and C1 is shown.

[0086] 2.1.2.4. Additional candidate insertions

[0087] In addition to spatial merge candidates and temporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidates and zero merge candidates. Combined bi-predictive merge candidates are generated by utilizing spatial merge candidates and temporal merge candidates. Combined bi-predictive merge candidates are used only for B slices. Combined bi-predictive candidate is generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another candidate. If the two tuples provide different motion hypotheses, they will form a new bi-predictive candidate. As an example, Figure 7 Depicted is the situation when two candidates in the original list (left) (with mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create a combined bi-predictive merge candidate that is added to the final list (right). There are many conventions regarding the combinations that are considered to generate these additional merge candidates.

[0088] Zero-motion candidates are inserted to fill the remaining entries in the merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts at zero and increases each time a new zero-motion candidate is added to the list.

[0089] More specifically, the following steps are performed in order until the merge list is full:

[0090] 1. For P slices, set the variable numRef to the number of reference pictures associated with list 0, or for B slices, set the variable numRef to the minimum number of reference pictures in the two lists;

[0091] 2. Add non-repeating zero-motion candidates:

[0092] For variable i being 0...numRef-1, add a default motion candidate to list 0 (if P slice) or both lists (if B slice) with MV set to (0,0) and reference picture index set to i.

[0093] 3. Add repeated zero motion candidates with MV set to (0,0), reference picture index of list 0 set to 0 (if P slice), and reference picture index of both lists set to 0 (if B slice).

[0094] Finally, no redundancy checks are performed on these candidates.

[0095] 2.1.3.AMVP

[0096] AMVP exploits the spatiotemporal correlation of motion vectors with neighboring PUs for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by first checking the availability of the top left temporary neighboring PU position, removing redundant candidates and adding zero vectors to keep the candidate list at a constant length. The encoder can then select the best prediction value from the candidate list and send the corresponding index indicating the selected candidate. Similarly, using merge index signaling, the index of the best motion vector candidate is encoded using a truncated unary number. In this case, the maximum value to be encoded is 2 (see Figure 8 ). In the following sections, detailed information about the derivation process of motion vector prediction candidates is provided.

[0097] 2.1.3.1. Derivation of AMVP Candidates

[0098] Figure 8 The derivation process of motion vector prediction candidates is summarized.

[0099] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For spatial motion vector candidate derivation, two motion vector candidates are finally derived based on the motion vectors of each PU located at five different positions, such as Figure 2 shown.

[0100] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different collocated positions. After making the first list of spatiotemporal candidates, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than two, motion vector candidates whose reference picture index in their associated reference picture list is greater than 1 are removed from the list. If the number of spatiotemporal motion vector candidates is less than two, an additional zero motion vector candidate is added to the list.

[0101] 2.1.3.2. Spatial Motion Vector Candidates

[0102] In the derivation of spatial motion vector candidates, a maximum of two candidates are considered among five potential candidates, which are selected from the candidate located at Figure 2The positions are derived from the PUs at the positions shown, which are the same as the positions of the motion merge. The derivation order on the left side of the current PU is defined as A0, A1 and scaled A0, scaled A1. The derivation order on the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, there are four cases on each side that can be used as motion vector candidates, two of which do not require spatial scaling and two use spatial scaling. These four different cases are summarized below.

[0103] No airspace scaling

[0104] (1) Same reference picture list and same reference picture index (same POC)

[0105] (2) Different reference picture lists, but same reference picture (same POC)

[0106] Airspace scaling

[0107] (3) Same reference picture list, but different reference pictures (different POC)

[0108] (4) Different reference picture lists and different reference pictures (different POCs)

[0109] First, check for no spatial scaling, then check for spatial scaling. Spatial scaling is considered when the POC between the reference picture of the neighboring PU and the reference picture of the current PU is different, regardless of the reference picture list. If all PUs of the left candidate are unavailable or intra-coded, scaling of the above motion vectors is allowed to facilitate the parallel derivation of the left MV candidate and the MV candidate above. Otherwise, spatial scaling of the above motion vectors is not allowed.

[0110] Figure 9 is a diagram of motion vector scaling of spatial motion vector candidates.

[0111] During spatial scaling, the motion vectors of neighboring PUs are scaled in a similar way to temporal scaling, e.g. Figure 9 The main difference is that the reference picture list and index of the current PU are given as input; the actual scaling process is the same as the time domain scaling process.

[0112] 2.1.3.3. Temporal Motion Vector Candidates

[0113] Except for the reference picture index derivation, all the processes for the derivation of temporal merge candidates are the same as those for the derivation of spatial motion vector candidates (see Figure 6 ). The reference picture index is signaled to the decoder.

[0114] 2.2. Sub-CU-based motion vector prediction method in JEM

[0115] In JEM with QTBT, each CU can have at most one set of motion parameters for each prediction direction. By dividing a large CU into sub-CUs and deriving motion information for all sub-CUs of the large CU, two sub-CU level motion vector prediction methods are considered in the encoder. The alternative temporal motion vector prediction (ATMVP) method allows each CU to obtain multiple sets of motion information from multiple blocks smaller than the current CU in the co-located reference picture. In the spatiotemporal motion vector prediction (STMVP) method, the motion vector of the sub-CU is recursively derived by using the temporal motion vector prediction value and the spatial neighboring motion vectors.

[0116] In order to preserve a more accurate motion field for sub-CU motion prediction, motion compression of reference frames is currently disabled.

[0117] Figure 10 An example of ATMVP motion prediction for a CU is shown.

[0118] 2.2.1. Alternative temporal motion vector prediction

[0119] In the alternative temporal motion vector prediction (ATMVP) method, the motion vector temporal motion vector prediction (TMVP) is modified by obtaining multiple sets of motion information (including motion vectors and reference indices) from blocks smaller than the current CU. In some embodiments, the sub-CU is a square N×N block (by default, N is set to 4).

[0120] ATMVP predicts the motion vectors of sub-CUs within a CU in two steps. The first step uses so-called temporal vectors to identify corresponding blocks in a reference picture. The reference picture is called the motion source picture. The second step is to partition the current CU into sub-CUs and obtain the motion vector and reference index for each sub-CU from its corresponding block.

[0121] In the first step, the reference picture and the corresponding block are determined by the motion information of the spatial neighboring blocks of the current CU. To avoid repeated scanning of neighboring blocks, the first merge candidate in the merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set to the temporal vector and the index of the motion source picture. In this way, in ATMVP, the corresponding block can be identified more accurately than in TMVP, where the corresponding block (sometimes called the co-located block) is always located at the bottom right or center position relative to the current CU.

[0122] In the second step, the corresponding block of the sub-CU is identified by the time domain vector in the motion source picture by adding the time domain vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the minimum motion grid covering the center sample point) is used to derive the motion information of the sub-CU. After the motion information of the corresponding N×N block is identified, it is converted into the motion vector and reference index of the current sub-CU in the same way as the TMVP of HEVC, where motion scaling and other procedures apply. For example, the decoder checks whether the low latency condition is met (i.e., the POC of all reference pictures of the current picture is less than the POC of the current picture), and may use the motion vector MVx of each sub-CU (corresponding to the motion vector of reference picture list X) to predict the motion vector MVy (where X is equal to 0 or 1, and Y is equal to 1-X).

[0123] 2.2.2. Spatiotemporal Motion Vector Prediction (STMVP)

[0124] In this method, the motion vectors of the sub-CUs are recursively derived in raster scan order. Figure 11 This concept is shown. Let us consider an 8x8 CU, which contains 4x4 sub-CUs A, B, C and D. Neighboring 4x4 blocks in the current frame are labeled a, b, c and d.

[0125] The motion derivation of sub-CU A starts with identifying its two spatial neighbors. The first neighborhood is the N×N block (block c) above sub-CU A. If this block c is not available or is internally coded, check the other N×N blocks above sub-CU A (from left to right, starting from block c). The second neighborhood is the block (block b) to the left of sub-CU A). If block b is not available or is internally coded, check the other blocks to the left of sub-CU A (from top to bottom, starting from block b). The motion information obtained from the neighboring blocks of each list is scaled to the first reference frame of the given list. Next, the temporal motion vector prediction value (TMVP) of sub-block A is derived by following the same process as the TMVP derivation specified in HEVC. The motion information of the co-located block at position D is obtained and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to 3) are averaged for each reference list. The average motion vector is designated as the motion vector of the current sub-CU.

[0126] 2.2.3. Sub-CU Motion Prediction Mode Signaling

[0127] Sub-CU modes are enabled as additional merge candidates and no additional syntax elements are needed to signal these modes. Two additional merge candidates are added to the merge candidate list of each CU to indicate ATMVP mode and STMVP mode. If the sequence parameter set indicating ATMVP and STMVP is enabled, a maximum of seven merge candidates are used. The encoding logic of the additional merge candidates is the same as that of the merge candidates in HM, which means that for each CU in a P or B slice, two additional RD checks are required for the two additional merge candidates.

[0128] In JEM, all bins of the merge index are context-coded by CABAC, whereas in HEVC, only the first bin is context-coded and the remaining bins are context-bypass coded.

[0129] 2.3. VVC inter-frame prediction method

[0130] There are several new codec tools for inter prediction improvements, such as Adaptive Motion Vector Differential Resolution (AMVR) for signaling MVD, Affine Prediction Mode, Triangular Prediction Mode (TPM), ATMVP, Generalized Bidirectional Prediction (GBI), Bidirectional Optical Flow (BIO).

[0131] 2.3.1. Adaptive motion vector differential resolution

[0132] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the PU's motion vector and the predicted motion vector) is signaled in units of quarter luma samples. In VVC, local adaptive motion vector resolution (LAMVR or AMVR for short) is introduced. In VVC, MVD can be encoded and decoded in units of quarter luma samples, integer luma samples, or four luma samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.

[0133] For CUs with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or four luma sample MV precision is used.

[0134] When the first MVD resolution flag of a CU is zero or the CU is not coded (meaning all MVDs in the CU are zero), the quarter-luminance sample MV resolution is used for the CU. When the CU uses integer-luminance sample MV precision or four-luminance sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.

[0135] In the encoder, CU-level RD checking is used to determine which MVD resolution will be used for the CU. That is, CU-level RD checking is performed three times for each MVD resolution. To speed up the encoder, the following coding scheme is applied in JEM.

[0136] · During the RD checking of a CU with normal quarter-luminance sample MVD resolution, the motion information (integer-luminance sample precision) of the current CU is stored. The stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement during the RD checking of the same CU with integer-luminance sample and 4-luminance sample MVD resolutions, so that the time-consuming motion estimation process is not repeated three times.

[0137] · Conditionally call the RD checking of a CU with 4-luminance sample MVD resolution. For a CU, when the RD cost of the integer-luminance sample MVD resolution is much larger than that of the quarter-luminance sample MVD resolution, skip the RD checking of the 4-luminance sample MVD resolution of the CU.

[0138] In Figure 12 shows the coding process. First, test the 1 / 4 pixel MV, calculate the RD cost and represent it as RDCost0, then test the integer MV and represent the RD cost as RDCost1. If RDCost1 < th * RDCost0 (where th is a positive value), then test the 4 pixel MV; otherwise, skip the 4 pixel MV. Basically, when checking the integer or 4 pixel MV, the motion information and RD cost, etc. for the 1 / 4 pixel MV are already known and can be reused to accelerate the coding process of the integer or 4 pixel MV.

[0139] In VVC, AMVR can also be applied to the affine prediction mode, where the resolution can be selected from 1 / 16 pixel, 1 / 4 pixel, and 1 pixel.

[0140] Figure 12 is the flowchart coded with different MV precisions.

[0141] 2.3.2. Triangular prediction mode

[0142] The concept of the triangular prediction mode (TPM) is to introduce a new triangular partition for motion compensation prediction. As Figure 13As shown, it divides the CU into two triangular prediction units in a diagonal or counter-diagonal direction. Each triangular prediction unit in the CU uses its own single prediction motion vector and reference frame index for inter-frame prediction, which are derived from a single single prediction candidate list. After predicting the triangular prediction unit, an adaptive weighting process is performed on the diagonal edges. Then, the transformation and quantization process are applied to the entire CU. It should be noted that this mode is only applied to merge mode (note: skip mode is regarded as a special merge mode).

[0143] 2.3.2.1.TPM Single Prediction Candidate List

[0144] The single prediction candidate list named TPM motion candidate list consists of five single prediction motion vector candidates. It is derived from seven neighboring blocks, including five spatial neighboring blocks (1 to 5) and two temporal co-located blocks (6 to 7), as shown in Figure 14 As shown in the figure. The motion vectors of the seven neighboring blocks are collected and entered into a uniprediction candidate list in the order of uniprediction motion vector, L0 motion vector of biprediction motion vector, L1 motion vector of biprediction motion vector, and average motion vector of L0 and L1 motion vectors of biprediction motion vector. If the number of candidates is less than five, a zero motion vector is added to the list. Motion candidates added to this list for TPM are called TPM candidates, and motion information derived from spatial / temporal blocks is called regular motion candidates.

[0145] More specifically, the following steps are involved:

[0146] 1) From A1, B1, B0, A0, B2, Col and Col2 (corresponding to Figure 14 Regular motion candidates are obtained from blocks 1-7 in , and a full pruning operation is performed when adding regular motion candidates from spatially neighboring blocks.

[0147] 2) Set the variable numCurrMergeCand = 0

[0148] 3) For each normal motion candidate derived from A1, B1, B0, A0, B2, Col and Col2, if not pruned and numCurrMergeCand is less than 5, if the normal motion candidate is uni-predicted (from list 0 or list 1), it will be directly added to the merge list as a TPM candidate, where numCurrMergeCand is increased by 1. This TPM candidate is named "initial uni-predicted candidate".

[0149] application Complete pruning .

[0150] 4) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if not pruned and numCurrMergeCand is less than 5, if the regular motion candidate is bi-predicted, the motion information in list 0 is added to the TPM merge list as a new TPM candidate (i.e., modified to uni-predicted in list 0), and numCurrMergeCand is increased by 1. This TPM candidate is named "truncated list 0 predicted candidate".

[0151] application Complete pruning .

[0152] 5) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if not pruned, and numCurrMergeCand is less than 5, then if the regular motion candidate is bi-predicted, the motion information from list 1 is added to the TPM merge list (i.e., modified from list 1 to uni-predicted), and numCurrMergeCand is incremented by 1. Such a TPM candidate is named "truncated list 1 predicted candidate".

[0153] application Complete pruning .

[0154] 6) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if not pruned, and numCurrMergeCand is less than 5, if the regular motion candidate is bi-predicted,

[0155] – If the slice QP of the list 0 reference picture is less than the slice QP of the list 1 reference picture, the motion information of list 1 is first scaled to the list 0 reference picture, and the average of the two MVs (one from the original list 0 and the other is the scaled MV from list 1) is added to the TPM merge list. This candidate is called the average uni-prediction of the list 0 motion candidate, and numCurrMergeCand is increased by 1.

[0156] Otherwise, first scale the motion information of list 0 to the list 1 reference picture, and add the average of the two MVs (one from the original list 1 and the other is the scaled MV from list 0) to the TPM merge list. This TPM candidate is called the average uni-prediction from list 1 motion candidate, and numCurrMergeCand is increased by 1.

[0157] application Complete pruning .

[0158] 7) If numCurrMergeCand is less than 5, add a zero motion vector candidate.

[0159] Figure 14 Example locations of neighboring blocks are shown.

[0160] When a candidate is inserted into the list, if it must be compared with all previously added candidates to determine whether it is the same as one of them, this process is called full pruning.

[0161] 2.3.2.2. Adaptive weighting process

[0162] After predicting each triangular prediction unit, an adaptive weighting process is applied to the diagonal edge between the two triangular prediction units to derive the final prediction for the entire CU. Two groups of weighting factors are defined as follows:

[0163] The first weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} for luma and chroma samples respectively;

[0164] The second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} for luma and chroma samples respectively.

[0165] The weighting factor set is selected based on the comparison of the motion vectors of the two triangular prediction units. When the reference pictures of the two triangular prediction units are different from each other or the difference in their motion vectors is greater than 16 pixels, the second weighting factor set is used. Otherwise, the first weighting factor set is used. Figure 15 An example is shown in .

[0166] 2.3.2.3. Signaling of Triangular Prediction Mode (TPM)

[0167] A one-bit flag indicating whether to use TPM may be signaled first. Afterwards, two partition mode indications (such as Figure 13 As shown) and the merge index selected for each of the two partitions are further signaled.

[0168] 2.3.2.3.1.TPM Flag Signaling

[0169] Let's denote the width and height of a luma block by W and H respectively. If W*H<64, triangular prediction mode is disabled.

[0170] When a block is encoded or decoded in affine mode, triangular prediction mode is also disabled.

[0171] When a block is encoded or decoded in merge mode, a bit flag may be signaled to indicate whether triangular prediction mode is enabled or disabled for the block.

[0172] The flag is encoded and decoded using three contexts based on the following equations:

[0173] Ctx index = ((left block L is available && L uses TPM for encoding and decoding?) 1:0) + ((upper block A is available && A uses TPM for encoding and decoding?) 1:0);

[0174] Figure 15 Neighboring blocks (A and L) used for context selection in TPM flag encoding and decoding are shown.

[0175] 2.3.2.3.2. Two partitioning modes (such as Figure 13 ), and the merge index selected for each of the two partitions

[0176] It should be noted that the partitioning mode and merge index of the two partitions are jointly coded. In some embodiments, the two partitions are restricted from using the same reference index. Therefore, there are 2 (partitioning mode) * N (maximum number of merge candidates) * (N-1) possibilities, where N is set to 5. An indication is coded and the mapping between the partitioning mode, the two merge indices and the coding indication is derived from the array defined below:

[0177] const uint8_t

[0178]

[0179] Division mode (45 degrees or 135 degrees) = g_TriangleCombination[signaled indication][0];

[0180] Merge index of candidate A = g_TriangleCombination[signaled indication][1];

[0181] Merge index of candidate B = g_TriangleCombination[signaled indication][2];

[0182] Once two motion candidates A and B are derived, the motion information of the two partitions (PU1 and PU2) can be set from A or B. Whether PU1 uses the motion information of the merge candidate A or B depends on the prediction direction of the two motion candidates. Table 1 shows the relationship between the two derived motion candidates A and B, and the two partitions.

[0183] Table 1: Motion information of partitions derived from two derived merge candidates (A, B)

[0184]

[0185] 2.3.2.3.3. Indicated entropy codec (indicated by merge_triangle_idx)

[0186] merge_triangle_idx is in the range of [0,39] (inclusive). K-order index Golomb (EG) code is used for binarization of merge_triangle_idx, where K is set to 1.

[0187] Kth Stage EG

[0188] To encode larger numbers with fewer bits (at the expense of using more bits to encode smaller numbers), it can be generalized using a non-negative integer parameter k. To encode a non-negative integer x in the order-kexp-Golomb code:

[0189] 1. Encode using the order-0exp-Golomb code described above

[0190] 2. Use binary encoding x to 2 k Modulus

[0191] Table 2: Exp-Golomb-k encoding and decoding examples

[0192]

[0193]

[0194] 2.3.3. Affine Motion Compensated Prediction

[0195] In HEVC, only the translation motion model is used for motion compensation prediction (MCP). Although in the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion and other unconventional motions. In VVC, a simplified affine transformation motion compensation prediction is applied using a 4-parameter affine model and a 6-parameter affine model. Figure 16 As shown, for the 4-parameter affine model, the affine motion field of the block is described by two control point motion vectors (CPMVs), and for the 6-parameter affine model, by 3 CPMVs.

[0196] Figure 16 An example of a simplified affine motion model is shown.

[0197] The motion vector field (MVF) of a block is described by the following equations, where the 4-parameter affine model in equation (1) (where the 4 parameters are defined as variables a, b, e, and f) and the 6-parameter affine model in equation (2) (where the 6 parameters are defined as variables a, b, c, d, e, and f) are:

[0198]

[0199] Among them (mv h 0,mv h 0) is the motion vector of the left top corner control point, and (mv h 1, mv h 1) is the motion vector of the top right corner control point, and (mv h 2, mv h 2) is the motion vector of the left bottom corner control point. All three motion vectors are called control point motion vectors (CPMVs). (x, y) represents the coordinates of the representative point relative to the left top sample point in the current block, and (mv h (x,y),mv v (x,y)) is the motion vector derived for the sample at (x,y). The CP motion vector can be signaled (as in affine AMVP mode) or dynamically derived (as in affine merge mode). w and h are the width and height of the current block. In practice, the division is implemented by right shift and rounding operations. In VTM, the representative point is defined as the center position of the sub-block, for example, when the coordinates of the left top corner of the sub-block relative to the left top sample point in the current block are (xs,ys), the coordinates of the representative point are defined as (xs+2,ys+2). For each sub-block (i.e., 4x4 in VTM), the representative point is used to derive the motion vector for the entire sub-block.

[0200] In order to further simplify the motion compensation prediction, the sub-block based affine transformation prediction is applied. In order to derive the motion vector of each M×N (M and N are both set to 4 in the current VVC) sub-block, as Figure 17 As shown, the motion vector of the center sample of each sub-block is calculated according to equations (1) and (2) and rounded to 1 / 16 fractional accuracy. A 1 / 16 pixel motion compensation interpolation filter is then applied to generate a prediction for each sub-block using the derived motion vector. The 1 / 16 pixel interpolation filter is introduced by the affine mode.

[0201] After MCP, the high-accuracy motion vector of each sub-block is rounded and saved to the same accuracy as the normal motion vector.

[0202] 2.3.3.1. Signaling of Affine Prediction

[0203] Similar to the translational motion model, due to affine prediction, there are two modes for signaling auxiliary information. They are AFFINE_INTER and AFFINE_MERGE modes.

[0204] 2.3.3.2.AF_INTER mode

[0205] AF_INTER mode can be applied to CUs with width and height greater than 8. A CU-level affine flag is signaled in the bitstream to indicate whether AF_INTER mode is used.

[0206] In this mode, for each reference picture list (list 0 or list 1), the affine AMVP candidate list consists of three affine motion prediction values in the following order, where each candidate includes the estimated CPMV of the current block. The best CPMV found at the encoder side (e.g. Figure 20 The difference between mv0, mv1, mv2 in and the estimated CPMV is signaled. In addition, the index of the affine AMVP candidate from which the estimated CPMV is derived is further signaled.

[0207] 1) Genetic affine motion prediction value

[0208] The checking order is similar to that of the spatial MVP in the HEVC AMVP list construction. First, the left inherited affine motion prediction value is derived from the first block in {A1, A0} that has been affine-coded and has the same reference picture as the current block. Second, the above inherited affine motion prediction value is derived from the first block in {B1, B0, B2} that has been affine-coded and has the same reference picture as the current block. Figure 19 Five blocks A1, A0, B1, B0, B2 are depicted in FIG.

[0209] Once a neighboring block is found to be coded or decoded using an affine mode, the CPMV of the codec covering the neighboring block is used to derive the predicted value of the CPMV of the current block. For example, if A1 is coded or decoded using a non-affine mode, and A0 is coded or decoded using a 4-parameter affine mode, the left-inherited affine MV predicted value will be derived from A0. In this case, the CPMV of the CU covering A0 (e.g. Figure 21 The top left CPMV in B and the top right CPMV denoted by ) is used to derive the estimated CPMV of the current block, given by Represents the left top (with coordinates (x0, y0)), right top (with coordinates (x1, y1)), and lower right bottom positions (with coordinates (x2, y2)) of the current block.

[0210] 2) Constructed affine motion prediction value

[0211] The constructed affine motion prediction value consists of control point motion vectors (CPMVs) derived from neighboring inter-frame codec blocks with the same reference picture, such as Figure 20 If the current affine motion model is 4-parameter affine, the number of CPMVs is 2. Otherwise, if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. CPMV on the top left Derived from the MV at the first block in the group {A, B, C} that is inter-coded and has the same reference picture as the current block. CPMV at the top right Derived from the MV of the first block in the group {D, E} that is inter-coded and has the same reference picture as the current block. CPMV at the bottom left Derived from the MV of the first block in the group {F, G} that is inter-coded and has the same reference picture as the current block.

[0212] - If the current affine motion model is 4-parameter affine, only if and have been established, i.e. and The constructed affine motion prediction value is inserted into the candidate list only when it is used as the estimated CPMV of the left top (with coordinates (x0, y0)) and right top (with coordinates (x1, y1)) positions of the current block.

[0213] - If the current affine motion model is 6-parameter affine, only if and When all are established, the constructed affine motion prediction value will be inserted into the candidate list, that is, and The CPMVs used as estimates of the left top (with coordinates (x0, y0)), right top (with coordinates (x1, y1)), and right top corner (with coordinates (x2, y2)) positions of the current block.

[0214] When inserting the constructed affine motion predictors into the candidate list, no pruning process is applied.

[0215] 3) Normal AMVP motion prediction value

[0216] Until the number of affine motion predictors reaches the maximum value, the following applies.

[0217] 1) By setting all CPMVs equal to (If available), derive affine motion prediction values.

[0218] 2) By setting all CPMVs equal to (If available), derive affine motion prediction values.

[0219] 3) By setting all CPMVs equal to (If available), derive affine motion prediction values.

[0220] 4) Derive affine motion prediction values by setting all CPMVs equal to HEVC TMVP (if available).

[0221] 5) Derive the affine motion prediction value by setting all CPMVs to zero MV.

[0222] Note that the affine motion predictor has been derived in the construction

[0223] In AF_INTER mode, when using the 4 / 6 parameter affine mode, 2 / 3 control points are required, and therefore 2 / 3 MVDs need to be encoded and decoded for these control points, as shown in Figure 18 In JVET-K0337, it is suggested to derive MV as follows, that is, to predict mvd1 and mvd2 based on mvd0.

[0224]

[0225] in mvd i and mv1 are the predicted motion vector, motion vector difference, and motion vector of the left top pixel (i=0), right top pixel (i=1), or left bottom pixel (i=2), respectively. Figure 18 (b) Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two components, i.e., newMV = mvA + mvB, and the two components of newMV are set to (xA + xB) and (yA + yB), respectively.

[0226] 2.3.3.3.AF_MERGE Mode

[0227] When a CU is applied in AF_MERGE mode, it obtains the first block coded in affine mode from the valid neighboring reconstructed blocks. Figure 21 As shown in (a), the order of selecting candidate blocks is from left, top, top right, bottom left to top left (in order represented by A, B, C, D, E). Figure 21 As represented by A0 in (b), the adjacent left bottom block is encoded and decoded in affine mode, and the control point (CP) motion vector mv0 of the left top corner, upper right corner and left bottom corner of the adjacent CU / PU containing block A is obtained N 、mv1 Nand mv2 N . And based on mv0 N 、mv1 N and mv2 N Calculate the motion vector mv0 of the top left corner / top right corner / bottom left corner of the current CU / PU C 、mv1 C and mv2 C (Only for 6-parameter affine model). It should be noted that in VTM-2.0, the sub-block at the top left corner (e.g., a 4×4 block in VTM) stores mv0, and if the current block is affine coded, the sub-block at the top right corner stores mv1. If the current block is coded with a 6-parameter affine model, the sub-block at the bottom left corner stores mv2; otherwise (using a 4-parameter affine model), LB stores mv2′. The other sub-blocks store the MV for MC.

[0228] Export the current CU mv0 C 、mv1 C and mv2 C After the CPMV of the current CU is calculated, the MVF of the current CU is generated according to the simplified affine motion model equations (1) and (2). In order to identify whether the current CU is coded or decoded using the AF_MERGE mode, the affine flag is signaled in the bitstream when at least one neighboring block is coded or decoded in the affine mode.

[0229] In JVET-L0142 and JVET-L0632, the affine merge candidate list is constructed by the following steps:

[0230] 1) Insert inherited affine candidates

[0231] An inherited affine candidate is one that is derived from the affine motion model of its valid neighboring affine codec block. The two largest inherited affine candidates are derived from the affine motion model of the neighboring block and inserted into the candidate list. For the left predictor, the scan order is {A0, A1}; for the top predictor, the scan order is {B0, B1, B2}.

[0232] 2) Insert the constructed affine candidate

[0233] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), the constructed affine candidate is inserted into the candidate list. The constructed affine candidate refers to a candidate constructed by combining the neighboring motion information of each control point.

[0234] a) First, Figure 22The motion information of the control points is derived from the specified spatial and temporal neighborhoods shown. CPk (k = 1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, and B3 are the spatial locations used to predict CPk (k = 1, 2, 3); T is the temporal location used to predict CP4.

[0235] The coordinates of CP1, CP2, CP3, and CP4 are (0,0), (W,0), (H,0), and (W,H), respectively, where W and H are the width and height of the current block.

[0236] The motion information for each control point is obtained according to the following priority order:

[0237] - For CP1, the priority is B2->B3->A2. If available, use B2. Otherwise, if B2 is available, use B3. If neither B2 nor B3 is available, use A2. If none of the three candidates are available, motion information for CP1 cannot be obtained.

[0238] - For CP2, the check priority is B1->B0.

[0239] - For CP3, the check priority is A1->A0.

[0240] -For CP4, use T.

[0241] b) Secondly, combinations of control points are used to construct affine merge candidates.

[0242] I. Motion information of three control points is required to construct a 6-parameter affine candidate. The three control points can be selected from the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combination {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4} will be converted into a 6-parameter motion model represented by the top left, top right, and bottom left control points.

[0243] II. The motion information of two control points is required to construct a 4-parameter affine candidate. These two control points can be selected from one of two combinations ({CP1, CP2}, {CP1, CP3}). These two combinations will be converted into a 4-parameter motion model represented by the left top and right top control points.

[0244] III. The combinations of constructed affine candidates are inserted into the candidate list in the following order: {CP1, CP2, CP3}, {CP1, CP2, CP4}, {CP1, CP3, CP4}, {CP2, CP3, CP4}, {CP1, CP2}, {CP1, CP3}

[0245] i. For each combination, check the reference index of list X for each CP. If they are the same, then this combination has a valid CPMV for list X. If the combination does not have a valid CPMV for both list 0 and list 1, then this combination is marked as invalid. Otherwise, it is valid, and the CPMV is placed in the sub-block merge list.

[0246] 3) Fill in with zero affine motion vector candidates

[0247] If the number of candidates in the affine merge candidate list is less than 5, for the sub-block merge candidate list, the MV is set to (0, 0) and the prediction direction is set to the 4-parameter merge candidate of uni-prediction (for P slices) and bi-prediction (for B slices) from list 0.

[0248] 2.3.4. Merge List Design in VVC

[0249] There are three different merge list construction processes supported in VVC:

[0250] 1) Sub-block merge candidate list: This includes ATMVP and affine merge candidates. Both affine and ATMVP modes share a common merge list construction process. ATMVP and affine merge candidates can be added sequentially. The sub-block merge list size is signaled in the slice header and has a maximum value of 5.

[0251] 2) Single prediction TPM merge list: For triangular prediction mode, a merge list construction process is shared between the two partitions, even though the two partitions can choose their own merge candidate indices. When constructing this merge list, the spatial neighbors and two temporal blocks of the block are checked. In our IDF, the motion information from the spatial neighbors and the temporal blocks are called regular motion candidates. These regular motion candidates are further used to derive multiple TPM candidates. Note that the transformation is performed at the whole block level, even though the two partitions may use different motion vectors to generate their own prediction blocks.

[0252] The single-prediction TPM merge list size is fixed to 5.

[0253] 3) Regular merge list: A common merge list construction process is used for the remaining codec blocks. Spatial / temporal / HMVP, paired bi-prediction merge candidates, and zero-motion candidates can be inserted sequentially. The regular merge list size is signaled in the slice header and has a maximum value of 6.

[0254] 2.3.4.1. Sub-block merge candidate list

[0255] It is recommended that all sub-block related motion candidates be placed in a separate merge list in addition to the regular merge list for non-sub-block merge candidates.

[0256] The sub-block related motion candidates are placed in a separate merge list, which is named "sub-block merge candidate list".

[0257] In one example, the sub-block merge candidate list includes affine merge candidates, ATMVP candidates and / or sub-block based STMVP candidates.

[0258] 2.3.4.1.1.JVET-L0278

[0259] In this paper, the ATMVP merge candidate in the normal merge list is moved to the first position of the affine merge list. In this way, all merge candidates in the new list (i.e., the sub-block based merge candidate list) are based on the sub-block codec tool.

[0260] 2.3.4.1.2.ATMVP in JVET-N1001

[0261] ATMVP is also known as sub-block based temporal motion vector prediction (SbTMVP).

[0262] In JVET-N1001, in addition to the regular merge candidate list, a special merge candidate list is added, called the sub-block merge candidate list (also known as the affine merge candidate list). The sub-block merge candidate list is filled with candidates in the following order:

[0263] a. ATMVP candidate (may be available or not);

[0264] b. Inherited affine candidates;

[0265] c. Constructed affine candidates, including constructed affine candidates based on TMVP, which use MV in the co-located reference picture;

[0266] d. Filled with zero MV 4-parameter affine model

[0267] VTM supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve the motion vector prediction and merge mode of the CU in the current picture. The same collocated picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in two main aspects:

[0268] 1. TMVP predicts CU-level motion, while SbTMVP predicts sub-CU-level motion;

[0269] 2. Whereas TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is the right bottom block or the center block relative to the current CU), SbTMVP applies motion shifting before obtaining the temporal motion information from the co-located picture, where the motion shifting is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.

[0270] exist Figure 23 and Figure 24 The SbTMVP process is shown in Figure 2. SbTMVP predicts the motion vector of the sub-CU in the current CU in two steps. In the first step, check Figure 23 The spatial neighborhood A1 in . If A1 has a motion vector that uses the co-located picture as its reference picture, this motion vector is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0,0).

[0271] In the second step, if Figure 24 As shown, the motion shift identified in step 1 is applied (ie, added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the collocated picture. Figure 24 The example in assumes that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block in the co-located picture (the minimum motion grid covering the center sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the co-located sub-CU, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.

[0272] In VTM, a combined sub-block based merge list containing both SbTMVP candidates and affine merge candidates is used for signaling of sub-block based merge mode. SbTMVP mode is enabled / disabled by the sequence parameter set (SPS) flag. If SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list is signaled in the SPS, and the maximum allowed size of the sub-block based merge list is 5 in VTM4.

[0273] The sub-CU size used in SbTMVP is fixed to 8×8, and like the affine merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.

[0274] The encoding logic of the additional SbTMVP merge candidate is the same as that of other merge candidates, that is, for each CU in a P slice or a B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.

[0275] Figure 23-24 An example of the SbTMVP process in VVC is shown.

[0276] The maximum number of candidates in the sub-block merge candidate list is denoted as MaxNumSubblockMergeCand.

[0277] 2.3.4.1.3. Syntax / Semantics Related to Sub-Block Merge Lists

[0278] 7.3.2.3 Sequence Parameter Setting RBSP Syntax

[0279]

[0280]

[0281]

[0282]

[0283]

[0284] 7.3.5.1 General Strip Header Syntax

[0285]

[0286]

[0287]

[0288]

[0289]

[0290]

[0291] 7.3.5.1Merge Data Syntax

[0292]

[0293]

[0294]

[0295] five_minus_max_num_subblock_merge_cand specifies the maximum number of sub-block based merging motion vector prediction (MVP) candidates supported in a slice, subtracted from 5. When five_minus_max_num_subblock_merge_cand is not present, it is inferred to be equal to 5-sps_sbtmvp_enabled_flag. The maximum number of sub-block based merging MVP candidates, MaxNumSubblockMergeCand, is derived as follows:

[0296] MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand

[0297] The value of MaxNumSubblockMergeCand should be in the range of 0 to 5, inclusive.

[0298] 8.5.5.2 Derivation of Motion Vectors and Reference Indices in Sub-Block Merge Mode

[0299] The inputs to this process are:

[0300] - the luma position (xCb, yCb) of the top left sample of the current luma codec block relative to the top left luma sample of the current picture,

[0301] - Two variables cbWidth and cbHeight that specify the width and height of the luma codec block.

[0302] The output of this process is:

[0303] - the number of luma codec sub-blocks in the horizontal direction numSbX and the vertical direction numSbY,

[0304] - reference indices refIdxL0 and refIdxL1,

[0305] - prediction lists utilize the flag arrays predFlagL0[xSbIdx][ySbIdx] and predFlagL1[xSbIdx][ySbIdx],

[0306] - Luma sub-block motion vector arrays mvL0[xSbIdx][ySbIdx] and mvL1[xSbIdx][ySbIdx] with 1 / 16 fractional sample accuracy, where xSbIdx=0…numSbX-1, ySbIdx=0…numSbY-1,

[0307] - Chroma sub-block motion vector arrays mvCL0[xSbIdx][ySbIdx] and mvCL1[xSbIdx][ySbIdx] with 1 / 32 fractional sample accuracy, where xSbIdx=0…numSbX-1, ySbIdx=0…numSbY-1,

[0308] - Double prediction weight index bcwIdx.

[0309] The variables numSbX, numSbY and the subblock merging candidate list subblockMergeCandList are derived through the following sequential steps:

[0310] 1. When sps_sbtmvp_enabled_flag is equal to 1, the following applies:

[0311] - For the derivation of availableFlagA1, refIdxLXA1, predFlagLXA1, and mvLXA1, the following applies:

[0312] - The luma position (xNbA1, yNbA1) within the adjacent luma codec block is set equal to (xCb-1, yCb+cbHeight-1).

[0313] - Call the block availability derivation procedure specified in clause 6.4.X [Ed.(BB): Neighboring block availability check procedure tbd] with the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the neighboring luma position (xNbA1, yNbA1) as input, and the output is assigned to the block availability flag availableA1.

[0314] - The available variables FlagA1, refIdxLXA1, predFlagLXA1 and mvLXA1 are derived as follows:

[0315] If availableA1 is equal to FALSE, availableFlagA1 is set equal to 0, both components of mvLXA1 are set equal to 0, refIdxLXA1 is set equal to -1 and predFlagLXA1 is set equal to 0, where X is 0 or 1, and bcwIdxA1 is set equal to 0.

[0316] Otherwise, availableFlagA1 is set equal to 1 and the following assignments are made:

[0317] mvLXA1=MvLX[xNbA1][yNbA1] (8-485)

[0318] refIdxLXA1=RefIdxLX[xNbA1][yNbA1] (8-486)

[0319] predFlagLXA1=PredFlagLX[xNbA1][yNbA1] (8-487)

[0320] - Call the derivation process of the sub-block based temporal merging candidate specified in clause 8.5.5.3, with the luma position (xCb, yCb), luma codec block width cbWidth, luma codec block height cbHeight, availability flag availableFlagA1, reference index refIdxLXA1, prediction list usage flag predFlagLXA1 and motion vector mvLXA1 as input, and output is the availability flag availableFlagSbC ol, the number of luma codec sub-blocks numSbX in the horizontal direction and the number of luma codec sub-blocks numSbY in the vertical direction, the reference index refIdxLXSbCol, the luma motion vector mvLXSbCol[xSbIdx][ySbIdx], and the prediction list usage flag predFlagLXSbCol[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1 and X is 0 or 1.

[0321] 2. When sps_affine_enabled_flag is equal to 1, the sample positions (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2), (xNbB3, yNbB3) and the variables numSbX and numSbY are derived as follows:

[0322] (xA0,yA0)=(xCb-1,yCb+cbHeight) (8-488)

[0323] (xA1,yA1)=(xCb-1,yCb+cbHeight-1) (8-489)

[0324] (xA2,yA2)=(xCb-1,yCb) (8-490)

[0325] (xB0,yB0)=(xCb+cbWidth,yCb-1) (8-491)

[0326] (xB1,yB1)=(xCb+cbWidth-1,yCb-1) (8-492)

[0327] (xB2,yB2)=(xCb-1,yCb-1) (8-493)

[0328] (xB3,yB3)=(xCb,yCb-1) (8-494)

[0329] numSbX=cbWidth>>2 (8-495)

[0330] numSbY=cbHeight>>2 (8-496)

[0331] 3. When sps_affine_enabled_flag is equal to 1, the variable availableFlagA is set equal to FALSE, and the following applies to (xNbA0, yNbA0) to (xNbA1, yNbA1) (xNbA0, yNbA0) k ,yNbA k ):

[0332] - Invoke the block availability derivation procedure specified in clause 6.4.X [Ed.(BB): Neighboring Block Availability Check Procedure tbd], where the current luma position (xCurr, yCurr) and the neighboring luma position (xNbA) are set equal to (xCb, yCb). k ,yNbA k ) as input, and the output is assigned to the block availability flag availableA k .

[0333] -When availableA k Equal to TRUE, MotionModelIdc[xNbA k ][yNbA k] is greater than 0, and availableFlagA is equal to FALSE, the following applies:

[0334] -The variable availableFlagA is set equal to TRUE, and motionModelIdcA is set equal to MotionModelIdc[xNbA k ][yNbA k ], (xNb, yNb) is set equal to (CbPosX[xNbA k ][yNbA k ],CbPosY[xNbA k ][yNbA k ]), nbW is set equal to CbWidth[xNbA k ][yNbA k ], nbH is set equal to CbHeight[xNbA k ][yNbA k ], numCpMv is set equal to MotionModelIdc[xNbA k ][yNbA k ]+1, and bcwIdxA is set equal to BcwIdx[xNbA k ][yNbA k ].

[0335] -For X replaced by 0 or 1, the following applies:

[0336] -When PredFlagLX[xNbA k ][yNbA k ] is equal to 1, the derivation process of the brightness affine control point motion vector of the neighboring block specified in clause 8.5.5 is called, where the brightness codec block position (xCb, yCb), the brightness codec block width and height (cbWidth, cbHeight), the neighboring brightness codec block position (xNb, yNb), the neighboring brightness codec block width and height (nbW, nbH), and the number of control point motion vectors numCpMv are taken as input, and the control point motion vector prediction value candidate cpMvLXA[cpIdx] (where cpIdx = 0..numCpMv-1) is taken as output.

[0337] - Make the following allocations:

[0338] predFlagLXA=PredFlagLX[xNbA k ][yNbA k ] (8-497)

[0339] refIdxLXA=RefIdxLX[xNbAk][yNbAk] (8-498)

[0340] 4. When sps_affine_enabled_flag is equal to 1, the variable availableFlagB is set equal to FALSE, and the following applies to (xNbB0, yNbB0) to (xNbB2, yNbB2) k ,yNbB k ):

[0341] - Invoke the block availability derivation procedure specified in clause 6.4.X [Ed.(BB): Neighboring Block Availability Check Procedure tbd], where the current luma position (xCurr, yCurr) and the neighboring luma position (xNbB) are set equal to (xCb, yCb). k ,yNbB k ) as input, and the output is assigned to the block availability flag availableB k .

[0342] -When availableB k Equal to TRUE, MotionModelIdc[xNbB k ][yNbB k ] is greater than 0, and availableFlagB is equal to FALSE, the following applies:

[0343] -The variable availableFlagB is set equal to TRUE, and motionModelIdcB is set equal to MotionModelIdc[xNbB k ][yNbB k ], (xNb, yNb) is set equal to (CbPosX[xNbAB][yNbB k ],CbPosY[xNbB k ][yNbB k ]), nbW is set equal to CbWidth[xNbB k ][yNbB k ], nbH is set equal to CbHeight[xNbB k ][yNbB k ], numCpMv is set equal to MotionModelIdc[xNbB k ][yNbB k ]+1, and bcwIdxB is set equal to BcwIdx[xNbB k ][yNbBk ].

[0344] -For X replaced by 0 or 1, the following applies:

[0345] -When PredFlagLX[xNbB k ][yNbB k ] is equal to TRUE, the derivation process of the luminance affine control point motion vector from the neighboring blocks specified in clause 8.5.5.5 is called, where the luminance codec block position (xCb, yCb), the luminance codec block width and height (cbWidth, cbHeight), the neighboring luminance codec block position (xNb, yNb), the neighboring luminance codec block width and height (nbW, nbH), and the number of control point motion vectors numCpMv are taken as input, and the control point motion vector prediction value candidates cpMvLXB[cpIdx] (where cpIdx = 0..numCpMv-1) are taken as output.

[0346] – Make the following assignments:

[0347] predFlagLXB=PredFlagLX[xNbB k ][yNbB k ] (8-499)

[0348] refIdxLXB=RefIdxLX[xNbB k ][yNbB k ] (8-500)

[0349] 5. When sps_affine_enabled_flag is equal to 1, the derivation process of the constructed affine control point motion vector merging candidates specified in clause 8.5.5.6 is called, with the luma codec block position (xCb, yCb), luma codec block width and height (cbWidth, cbHeight), availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3 as input, and the availability flag availableFlagConstK, reference index refIdxLXConstK, prediction list utilization flag predFlagLXConstK, motion model index motionModelIdcConstK, dual prediction weight index bcwIdxConstK and cpMvpLXConstK[cpIdx] (where X is 0 or 1, K = 1..6, cpIdx = 0..2) as output.

[0350] 6. The initial sub-block merging candidate list subblockMergeCandList is constructed as follows: i = 0

[0351] if(availableFlagSbCol)

[0352] subblockMergeCandList[i++] = SbCol

[0353] if(availableFlagA && i < MaxNumSubblockMergeCand)

[0354] subblockMergeCandList[i++] = A

[0355] if(availableFlagB && i < MaxNumSubblockMergeCand)

[0356] subblockMergeCandList[i++] = B

[0357] if(availableFlagConst1 && i < MaxNumSubblockMergeCand)

[0358] subblockMergeCandList[i++] = Const1 (8 - 501)

[0359] if(availableFlagConst2 && i < MaxNumSubblockMergeCand)

[0360] subblockMergeCandList[i++] = Const2

[0361] if(availableFlagConst3 && i < MaxNumSubblockMergeCand)

[0362] subblockMergeCandList[i++] = Const3

[0363] if(availableFlagConst4 && i < MaxNumSubblockMergeCand)

[0364] subblockMergeCandList[i++] = Const4

[0365] if(availableFlagConst5&&i <MaxNumSubblockMergeCand)

[0366] subblockMergeCandList[i++]=Const5

[0367] if(availableFlagConst6&&i <MaxNumSubblockMergeCand)

[0368] subblockMergeCandList[i++]=Const6

[0369] 7. The variables numCurrMergeCand and numOrigMergeCand are set to the number of merging candidates in subblockMergeCandList.

[0370] 8. When numCurrMergeCand is less than MaxNumSubblockMergeCand, repeat the following operations until numCurrMergeCand is equal to MaxNumSubblockMergeCand, where mvZero[0] and mvZero[1] are both equal to 0:

[0371] -Export zeroCand as follows m Reference index, prediction list utilization flag and motion vector of , where m is equal to (numCurrMergeCand)-numOrigMergeCand):

[0372] refIdxL0ZeroCand m =0 (8-502)

[0373] predFlagL0ZeroCand m =1 (8-503)

[0374] cpMvL0ZeroCand m [0] = mvZero (8-504)

[0375] cpMvL0ZeroCand m [1]=mvZero (8-505)

[0376] cpMvL0ZeroCand m [2]=mvZero (8-506)

[0377] refIdxL1ZeroCand m =(slice_type==B)? 0:-1(8-507)

[0378] predFlagL1ZeroCand m =(slice_type==B)? 1:0(8-508)

[0379] cpMvL1ZeroCand m [0] = mvZero (8-509)

[0380] cpMvL1ZeroCand m [1]=mvZero (8-510)

[0381] cpMvL1ZeroCand m [2]=mvZero (8-511)

[0382] motionModelIdcZeroCand m =1 (8-512)

[0383] bcwIdxZeroCand m =0 (8-513)

[0384] -Candidate zeroCand m , where m, which is equal to (numCurrMergeCand - numOrigMergeCand), is added to the end of subblockMergeCandList and numCurrMergeCand is incremented by 1, as follows:

[0385] subblockMergeCandList[numCurrMergeCand++]=zeroCand m (8-514)

[0386] The variables refIdxL0, refIdxL1, predFlagL0[xSbIdx][ySbIdx], predFlagL1[xSbIdx][ySbIdx], mvL0[xSbIdx][ySbIdx], mvL1[xSbIdx][ySbIdx], mvCL0[xSbIdx][ySbIdx], and mvCL1[xSbIdx][ySbIdx] are derived as follows, where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1:

[0387] - If subblockMergeCandList[merge_subblock_idx[xCb][yCb]] is equal to SbCol, then the bi-prediction weight index bcwIdx is set equal to 0 and the following applies (X is 0 or 1):

[0388] refIdxLX=refIdxLXSbCol (8-515)

[0389] - For xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, the following applies:

[0390] predFlagLX[xSbIdx][ySbIdx]=predFlagLXSbCol[xSbIdx][ySbIdx] (8-516)

[0391] mvLX[xSbIdx][ySbIdx][0]=mvLXSbCol[xSbIdx][ySbIdx][0](8-517)

[0392] mvLX[xSbIdx][ySbIdx][1]=mvLXSbCol[xSbIdx][ySbIdx][1](8-518)

[0393] - When predFlagLX[xSbIdx][ySbIdx] is equal to 1, the derivation process of chroma motion vector in clause 8.5.2.13 is called, with mvLX[xSbIdx][ySbIdx] and refIdxLX as input, and the output is mvCLX[xSbIdx][ySbIdx].

[0394] - For x = xCb..xCb + cbWidth-1 and y = yCb..yCb + cbHeight-1, the following assignments are made:

[0395] MotionModelIdc[x][y]=0 (8-519)

[0396] Otherwise (subblockMergeCandList[merge_subblock_idx[xCb][yCb]] is not equal to SbCol), the following applies (where X is 0 or 1):

[0397] - When N is the candidate at position merge_subblock_idx[xCb][yCb] in the subblock merging candidate list subblockMergeCandList (N=subblockMergeCandList[merge_subblock_idx[xCb][yCb]]), the following assignment is performed:

[0398] refIdxLX=refIdxLXN (8-520)

[0399] predFlagLX[0][0]=predFlagLXN (8-521)

[0400] cpMvLX[0]=cpMvLXN[0] (8-522)

[0401] cpMvLX[1]=cpMvLXN[1] (8-523)

[0402] cpMvLX[2]=cpMvLXN[2] (8-524)

[0403] numCpMv=motionModelIdxN+1 (8-525)

[0404] bcwIdx=bcwIdxN (8-526)

[0405] - For xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, the following applies:

[0406] predFlagLX[xSbIdx][ySbIdx]=predFlagLX[0][0](8-527)

[0407] - When predFlagLX[0][0] is equal to 1, the derivation of motion vector arrays from affine control point motion vectors specified in subclause 8.5.5.9 is called with the luma codec block position (xCb, yCb), the luma codec block width cbWidth, the luma prediction block height cbHeight, the number of control point motion vectors numCpMv, the control point motion vectors cpMvLX[cpIdx] (where cpIdx is 0..2), and the number of luma codec subblocks in the horizontal direction numSbX and the number of luma codec subblocks in the vertical direction numSbY as input, and the luma subblock motion vector array mvLX[xSbIdx][ySbIdx] and the chroma subblock motion vector array mvCLX[xSbIdx][ySbIdx] (where xSbIdx = 0..numSbX-1, ySbIdx = 0..numSbY-1) as output.

[0408] - For x = xCb..xCb+cbWidth-1 and y = yCb..yCb+cbHeight-1, make the following assignments:

[0409] MotionModelIdc[x][y]=numCpMv-1 (8-528)

[0410] 8.5.5.6 Derivation of Constructed Affine Control Point Motion Vector Merging Candidates

[0411] The inputs to this process are:

[0412] -Specifies the luma position (xCb, yCb) of the top left sample of the current luma codec block relative to the top left luma sample of the current picture,

[0413] -Specify the two variables cbWidth and cbHeight of the width and height of the current brightness codec block,

[0414] -Availability flags availableA0, availableA1, availableA2, availableB0, availableB1, availableB2, availableB3,

[0415] - Sample point positions (xNbA0, yNbA0), (xNbA1, yNbA1), (xNbA2, yNbA2), (xNbB0, yNbB0), (xNbB1, yNbB1), (xNbB2, yNbB2) and (xNbB3, yNbB3).

[0416] The output of this process is:

[0417] - Availability flag availableFlagConstK of the constructed affine control point motion vector merging candidate, where K = 1..6,

[0418] – Reference index refIdxLXConstK, where K = 1..6, X is 0 or 1,

[0419] – The prediction list uses the flag predFlagLXConstK, where K = 1..6, X is 0 or 1,

[0420] – Affine motion model index motionModelIdcConstK, where K = 1..6,

[0421] – Dual prediction weight index bcwIdxConstK, where K = 1..6,

[0422] – Constructed affine control point motion vector cpMvLXConstK[cpIdx], where cpIdx=0..2, K=1..6 and X is 0 or 1.

[0423] The first (top left) control point motion vector cpMvLXCorner[0], reference index refIdxLXCorner[0], prediction list utilization flag predFlagLXCorner[0], bi-prediction weight index bcwIdxCorner[0] and availability flag availableFlagCorner[0] are derived as follows, where X is 0 and 1:

[0424] - The availability flag availableFlagCorner[0] is set equal to FALSE.

[0425] - The following applies to (xNbTL,yNbTL), where TL is replaced by B2, B3, and A2:

[0426] - When availableTL equals TRUE and availableFlagCorner[0] equals FALSE, the following applies (where X is 0 and 1):

[0427] refIdxLXCorner[0]=RefIdxLX[xNbTL][yNbTL](8-572)

[0428] predFlagLXCorner[0]=PredFlagLX[xNbTL][yNbTL] (8-573)

[0429] cpMvLXCorner[0]=MvLX[xNbTL][yNbTL] (8-574)

[0430] bcwIdxCorner[0]=BcwIdx[xNbTL][yNbTL] (8-575)

[0431] availableFlagCorner[0]=TRUE (8-576)

[0432] The second (top right) control point motion vector cpMvLXCorner[1], reference index refIdxLXCorner[1], prediction list utilization flag predFlagLXCorner[1], bi-prediction weight index bcwIdxCorner[1] and availability flag availableFlagCorner[1] are derived as follows, where X is 0 and 1:

[0433] - The availability flag availableFlagCorner[1] is set equal to FALSE.

[0434] - The following applies to (xNbTR,yNbTR), where TR is replaced by B1 and B0:

[0435] - When availableTR equals TRUE and availableFlagCorner[1] equals FALSE, the following applies (where X is 0 and 1):

[0436] refIdxLXCorner[1]=RefIdxLX[xNbTR][yNbTR] (8-577)

[0437] predFlagLXCorner[1]=PredFlagLX[xNbTR][yNbTR] (8-578)

[0438] cpMvLXCorner[1]=MvLX[xNbTR][yNbTR] (8-579)

[0439] bcwIdxCorner[1]=BcwIdx[xNbTR][yNbTR] (8-580)

[0440] availableFlagCorner[1]=TRUE (8-581)

[0441] The third (bottom left) control point motion vector cpMvLXCorner[2], reference index refIdxLXCorner[2], prediction list utilization flag predFlagLXCorner[2], bi-prediction weight index bcwIdxCorner[2] and availability flag availableFlagCorner[2] are derived as follows, where X is 0 and 1:

[0442] - The availability flag availableFlagCorner[2] is set equal to FALSE.

[0443] - The following applies to (xNbBL,yNbBL), where BL is replaced by A1 and A0:

[0444] - When availableBL equals TRUE and availableFlagCorner[2] equals FALSE, the following applies (where X is 0 and 1):

[0445] refIdxLXCorner[2]=RefIdxLX[xNbBL][yNbBL] (8-582)

[0446] predFlagLXCorner[2]=PredFlagLX[xNbBL][yNbBL] (8-583)

[0447] cpMvLXCorner[2]=MvLX[xNbBL][yNbBL] (8-584)

[0448] bcwIdxCorner[2]=BcwIdx[xNbBL][yNbBL] (8-585)

[0449] availableFlagCorner[2]=TRUE (8-586)

[0450] The fourth (collocated right bottom) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list utilization flag predFlagLXCorner[3], bi-prediction weight index bcwIdxCorner[3] and availability flag availableFlagCorner[3] are derived as follows, where X is 0 and 1:

[0451] - The reference index refIdxLXCorner[3] (where X is 0 or 1) of the time-domain merging candidate is set equal to 0.

[0452] - Export the variables mvLXCol and availableFlagLXCol (where X is 0 or 1) as follows:

[0453] If slice_temporal_mvp_enabled_flag is equal to 0, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.

[0454] Otherwise (slice_temporal_mvp_enabled_flag is equal to 1), the following applies:

[0455] xColBr=xCb+cbWidth

[0456] (8-587)

[0457] yColBr=yCb+cbHeight

[0458] (8-588)

[0459] - If yCb >> CtbLog2SizeY equals yColBr >> CtbLog2SizeY, yColBr is less than pic_height_in_luma_samples and xColBr is less than pic_width_in_luma_samples, then the following applies:

[0460] - The variable colCb specifies the luma codec block that covers the modification position given by ((xColBr>>3)<<3, (yColBr>>3)<<3) within the collocated picture specified by ColPic.

[0461] - The luma position (xColCb, yColCb) is set equal to the top left sample of the collocated luma codec block specified by colCb relative to the top left luma sample of the collocated picture specified by ColPic.

[0462] - Invoke the derivation process of the co-located motion vector specified in clause 8.5.2.12 with as input currCb, colCb, (xColCb, yColCb), refIdxLXCorner[3] and sbFlag set equal to 0, and the output assigned to mvLXCol and availableFlagLXCol.

[0463] Otherwise, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0.

[0464] - Export the variables availableFlagCorner[3], predFlagL0Corner[3], cpMvL0Corner[3], and predFlagL1Corner[3] as follows:

[0465] availableFlagCorner[3]=availableFlagL0Col (8-589)

[0466] predFlagL0Corner[3]=availableFlagL0Col (8-590)

[0467] cpMvL0Corner[3]=mvL0Col (8-591)

[0468] predFlagL1Corner[3]=0

[0469] (8-592)

[0470] - When slice_type is equal to B, the variables availableFlagCorner[3], predFlagL1Corner[3], and cpMvL1Corner[3] are derived as follows:

[0471] availableFlagCorner[3]=availableFlagL0Col||availableFlagL1Col (8-593)

[0472] predFlagL1Corner[3]=availableFlagL1Col (8-594)

[0473] cpMvL1Corner[3]=mvL1Col (8-595)

[0474] bcwIdxCorner[3]=0 (8-596)

[0475] When sps_affine_type_flag is equal to 1, the first four constructed affine control point motion vector merging candidates ConstK (where K=1..4) including the availability flag availableFlagConstK, the reference index refIdxLXConstK, the prediction list utilization flag predFlagLXConstK, the affine motion model index motionModelIdcConstK, and the constructed affine control point motion vector cpMvLXConstK[cpIdx] (where cpIdx=0..2 and X is 0 or 1) are derived as follows:

[0476] 1. When availableFlagCorner[0] is equal to TRUE, availableFlagCorner[1] is equal to TRUE, and availableFlagCorner[2] is equal to TRUE, the following applies:

[0477] -For X replaced by 0 or 1, the following applies:

[0478] -Export the variable availableFlagLX as follows:

[0479] - availableFlagLX is set equal to TRUE if all of the following conditions are TRUE:

[0480] -predFlagLXCorner[0] is equal to 1

[0481] -predFlagLXCorner[1] is equal to 1

[0482] -predFlagLXCorner[2] is equal to 1

[0483] -refIdxLXCorner[0] is equal to refIdxLXCorner[1]

[0484] -refIdxLXCorner[0] is equal to refIdxLXCorner[2]

[0485] - Otherwise, availableFlagLX is set equal to FALSE.

[0486] -When availableFlagLX is equal to TRUE, the following assignments are made:

[0487] predFlagLXConst1=1 (8-597)

[0488] refIdxLXConst1=refIdxLXCorner[0] (8-598)

[0489] cpMvLXConst1[0]=cpMvLXCorner[0] (8-599)

[0490] cpMvLXConst1[1]=cpMvLXCorner[1] (8-600)

[0491] cpMvLXConst1[2]=cpMvLXCorner[2] (8-601)

[0492] -Derive the dual prediction weight index bcwIdxConst1 as follows:

[0493] - If availableFlagL0 is equal to 1 and availableFlagL1 is equal to 1, the derivation process of the dual prediction weight index for the constructed affine control point motion vector merging candidate specified in clause 8.5.5.10 is called, with the dual prediction weight indices bcwIdxCorner[0], bcwIdxCorner[1] and bcwIdxCorner[2] as input, and the output is assigned to the dual prediction weight index bcwIdxConst1.

[0494] Otherwise, the bi-prediction weight index bcwIdxConst1 is set equal to 0.

[0495] -Export the variables availableFlagConst1 and motionModelIdcConst1 as follows:

[0496] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst1 is set equal to TRUE and motionModelIdcConst1 is set equal to 2.

[0497] Otherwise, availableFlagConst1 is set equal to FALSE and motionModelIdcConst1 is set equal to 0.

[0498] 2. When availableFlagCorner[0] is equal to TRUE, availableFlagCorner[1] is equal to TRUE, and availableFlagCorner[3] is equal to TRUE, the following applies:

[0499] -For X replaced by 0 or 1, the following applies:

[0500] - Export the variable availableFlagLX as follows:

[0501] - availableFlagLX is set to TRUE if all of the following conditions are TRUE:

[0502] -predFlagLXCorner[0] is equal to 1

[0503] -predFlagLXCorner[1] is equal to 1

[0504] -predFlagLXCorner[3] is equal to 1

[0505] -refIdxLXCorner[0] is equal to refIdxLXCorner[1]

[0506] -refIdxLXCorner[0] is equal to refIdxLXCorner[3]

[0507] - Otherwise, availableFlagLX is set equal to FALSE.

[0508] -When availableFlagLX is equal to TRUE, the following assignments are made:

[0509] predFlagLXConst2=1 (8-602)

[0510] refIdxLXConst2=refIdxLXCorner[0] (8-603)

[0511] cpMvLXConst2[0]=cpMvLXCorner[0] (8-604)

[0512] cpMvLXConst2[1]=cpMvLXCorner[1] (8-605)

[0513] cpMvLXConst2[2]=cpMvLXCorner[3]+cpMvLXCorner[0]-cpMvLXCorner[1]

[0514] (8-606)

[0515] cpMvLXConst2[2][0]=Clip3(-2 17 ,2 17-1,cpMvLXConst2[2][0]) (8-607)

[0516] cpMvLXConst2[2][1]=Clip3(-2 17 ,2 17 -1,cpMvLXConst2[2][1]) (8-608)

[0517] -Derive the dual prediction weight index bcwIdxConst2 as follows:

[0518] - If availableFlagL0 is equal to 1 and availableFlagL1 is equal to 1, the derivation process of the dual prediction weight index for the constructed affine control point motion vector merging candidate specified in clause 8.5.5.10 is called, with the dual prediction weight indices bcwIdxCorner[0], bcwIdxCorner[1] and bcwIdxCorner[3] as input, and the output is assigned to the dual prediction weight index bcwIdxConst2.

[0519] Otherwise, the bi-prediction weight index bcwIdxConst2 is set equal to 0.

[0520] -Export the variables availableFlagConst2 and motionModelIdcConst2 as follows:

[0521] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst2 is set equal to TRUE and motionModelIdcConst2 is set equal to 2.

[0522] Otherwise, availableFlagConst2 is set equal to FALSE and motionModelIdcConst2 is set equal to 0.

[0523] 3. When availableFlagCorner[0] is equal to TRUE, availableFlagCorner[2] is equal to TRUE, and availableFlagCorner[3] is equal to TRUE, the following applies:

[0524] -For X replaced by 0 or 1, the following applies:

[0525] - Export the variable availableFlagLX as follows:

[0526] - availableFlagLX is set equal to TRUE if all of the following conditions are TRUE:

[0527] -predFlagLXCorner[0] is equal to 1

[0528] -predFlagLXCorner[2] is equal to 1

[0529] -predFlagLXCorner[3] is equal to 1

[0530] -refIdxLXCorner[0] is equal to refIdxLXCorner[2]

[0531] -refIdxLXCorner[0] is equal to refIdxLXCorner[3]

[0532] - Otherwise, availableFlagLX is set equal to FALSE.

[0533] -When availableFlagLX is equal to TRUE, the following assignments are made:

[0534] predFlagLXConst3=1 (8-609)

[0535] refIdxLXConst3=refIdxLXCorner[0] (8-610)

[0536] cpMvLXConst3[0]=cpMvLXCorner[0] (8-611)

[0537] cpMvLXConst3[1]=cpMvLXCorner[3]+cpMvLXCorner[0]-cpMvLXCorner[2] (8-612)

[0538] cpMvLXConst3[1][0]=Clip3(-2 17 ,2 17 -1,cpMvLXConst3[1][0]) (8-613)

[0539] cpMvLXConst3[1][1]=Clip3(-2 17 ,2 17 -1,cpMvLXConst3[1][1]) (8-614)

[0540] cpMvLXConst3[2]=cpMvLXCorner[2] (8-615)

[0541] -Derive the dual prediction weight index bcwIdxConst3 as follows:

[0542] - If availableFlagL0 is equal to 1 and availableFlagL1 is equal to 1, the derivation process of the dual prediction weight index for the constructed affine control point motion vector merging candidate specified in clause 8.5.5.10 is called, with the dual prediction weight indices bcwIdxCorner[0], bcwIdxCorner[2] and bcwIdxCorner[3] as input, and the output is assigned to the dual prediction weight index bcwIdxConst3.

[0543] Otherwise, the bi-prediction weight index bcwIdxConst3 is set equal to 0.

[0544] -Export the variables availableFlagConst3 and motionModelIdcConst3 as follows:

[0545] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst3 is set equal to TRUE and motionModelIdcConst3 is set equal to 2.

[0546] Otherwise, availableFlagConst3 is set equal to FALSE and motionModelIdcConst3 is set equal to 0.

[0547] 4. When availableFlagCorner[1] equals TRUE, availableFlagCorner[2] equals TRUE, and availableFlagCorner[3] equals TRUE, the following applies:

[0548] -For X replaced by 0 or 1, the following applies:

[0549] - Export the variable availableFlagLX as follows:

[0550] - availableFlagLX is set equal to TRUE if all of the following conditions are TRUE:

[0551] -predFlagLXCorner[1] is equal to 1

[0552] -predFlagLXCorner[2] is equal to 1

[0553] -predFlagLXCorner[3] is equal to 1

[0554] -refIdxLXCorner[1] is equal to refIdxLXCorner[2]

[0555] -refIdxLXCorner[1] is equal to refIdxLXCorner[3]

[0556] - Otherwise, availableFlagLX is set equal to FALSE.

[0557] -When availableFlagLX is equal to TRUE, the following assignments are made:

[0558] predFlagLXConst4=1 (8-616)

[0559] refIdxLXConst4=refIdxLXCorner[1] (8-617)

[0560] cpMvLXConst4[0]=cpMvLXCorner[1]+cpMvLXCorner[2]-cpMvLXCorner[3] (8-618)

[0561] cpMvLXConst4[0][0]=Clip3(-2 17 ,2 17 -1,cpMvLXConst4[0][0]) (8-619)

[0562] cpMvLXConst4[0][1]=Clip3(-2 17 ,2 17 -1,cpMvLXConst4[0][1]) (8-620)

[0563] cpMvLXConst4[1]=cpMvLXCorner[1] (8-621)

[0564] cpMvLXConst4[2]=cpMvLXCorner[2] (8-622)

[0565] -Derive the dual prediction weight index bcwIdxConst4 as follows:

[0566] - If availableFlagL0 is equal to 1 and availableFlagL1 is equal to 1, the derivation process of the dual prediction weight index for the constructed affine control point motion vector merging candidate specified in clause 8.5.5.10 is called, with the dual prediction weight indices bcwIdxCorner[1], bcwIdxCorner[2] and bcwIdxCorner[3] as input, and the output is assigned to the dual prediction weight index bcwIdxConst4.

[0567] Otherwise, the bi-prediction weight index bcwIdxConst4 is set equal to 0.

[0568] -Export the variables availableFlagConst4 and motionModelIdcConst4 as follows:

[0569] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst4 is set equal to TRUE and motionModelIdcConst4 is set equal to 2.

[0570] Otherwise, availableFlagConst4 is set equal to FALSE and motionModelIdcConst4 is set equal to 0.

[0571] The last two constructed affine control point motion vector merging candidates ConstK (where K = 5..6) including the availability flag availableFlagConstK, the reference index refIdxLXConstK, the prediction list utilization flag predFlagLXConstK, the affine motion model index motionModelIdcConstK, and the constructed affine control point motion vector cpMvLXConstK[cpIdx] (where cpIdx = 0..2 and X is 0 or 1) are derived as follows:

[0572] 5. When availableFlagCorner[0] is equal to TRUE and availableFlagCorner[1] is equal to TRUE, the following applies:

[0573] -For X replaced by 0 or 1, the following applies:

[0574] - Export the variable availableFlagLX as follows:

[0575] - availableFlagLX is set equal to TRUE if all of the following conditions are TRUE:

[0576] -predFlagLXCorner[0] is equal to 1

[0577] -predFlagLXCorner[1] is equal to 1

[0578] -refIdxLXCorner[0] is equal to refIdxLXCorner[1]

[0579] - Otherwise, availableFlagLX is set equal to FALSE.

[0580] -When availableFlagLX is equal to TRUE, the following assignments are made:

[0581] predFlagLXConst5=1 (8-623)

[0582] refIdxLXConst5=refIdxLXCorner[0] (8-624)

[0583] cpMvLXConst5[0]=cpMvLXCorner[0] (8-625)

[0584] cpMvLXConst5[1]=cpMvLXCorner[1] (8-626)

[0585] -Derive the dual prediction weight index bcwIdxConst5 as follows:

[0586] - If availableFlagL0 is equal to 1, availableFlagL1 is equal to 1, and bcwIdxCorner[0] is equal to bcwIdxCorner[1], then bcwIdxConst5 is set equal to bcwIdxCorner[0].

[0587] Otherwise, the bi-prediction weight index bcwIdxConst5 is set equal to 0.

[0588] - Export the variables availableFlagConst5 and motionModelIdcConst5 as follows:

[0589] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst5 is set equal to TRUE and motionModelIdcConst5 is set equal to 1.

[0590] Otherwise, availableFlagConst5 is set equal to FALSE and motionModelIdcConst5 is set equal to 0.

[0591] 6. When availableFlagCorner[0] is equal to TRUE and availableFlagCorner[2] is equal to TRUE, the following applies:

[0592] -For X replaced by 0 or 1, the following applies:

[0593] - Export the variable availableFlagLX as follows:

[0594] - availableFlagLX is set equal to TRUE if all of the following conditions are TRUE:

[0595] -predFlagLXCorner[0] is equal to 1

[0596] -predFlagLXCorner[2] is equal to 1

[0597] -refIdxLXCorner[0] is equal to refIdxLXCorner[2]

[0598] - Otherwise, availableFlagLX is set equal to FALSE.

[0599] - When availableFlagLX equals TRUE, the following applies:

[0600] - The second control point motion vector cpMvLXCorner[1] is derived as follows:

[0601]

[0602] - Invoke the motion vector rounding process specified in clause 8.5.2.14 with mvX set equal to cpMvLXCorner[1], rightShift set equal to 7, leftShift set equal to 0 as inputs, and the rounded cpMvLXCorner[1] as output.

[0603] - Make the following assignments:

[0604] predFlagLXConst6=1 (8-629)

[0605] refIdxLXConst6=refIdxLXCorner[0] (8-630)

[0606] cpMvLXConst6[0]=cpMvLXCorner[0] (8-631)

[0607] cpMvLXConst6[1]=cpMvLXCorner[1] (8-632)

[0608] cpMvLXConst6[0][0]=Clip3(-2 17 ,2 17 -1,cpMvLXConst6[0][0])(8-633)

[0609] cpMvLXConst6[0][1]=Clip3(-2 17 ,2 17 -1,cpMvLXConst6[0][1])(8-634)

[0610] cpMvLXConst6[1][0]=Clip3(-2 17 ,2 17 -1,cpMvLXConst6[1][0])(8-635)

[0611] cpMvLXConst6[1][1]=Clip3(-2 17 ,2 17 -1,cpMvLXConst6[1][1])(8-636)

[0612] -Derive the dual prediction weight index bcwIdxConst6 as follows:

[0613] - If availableFlagL0 is equal to 1, availableFlagL1 is equal to 1, and bcwIdxCorner[0] is equal to bcwIdxCorner[2], then bcwIdxConst6 is set equal to bcwIdxCorner[0].

[0614] Otherwise, the bi-prediction weight index bcwIdxConst6 is set equal to 0.

[0615] -Export the variables availableFlagConst6 and motionModelIdcConst6 as follows:

[0616] If availableFlagL0 or availableFlagL1 is equal to 1, availableFlagConst6 is set equal to TRUE and motionModelIdcConst6 is set equal to 1.

[0617] Otherwise, availableFlagConst6 is set equal to FALSE and motionModelIdcConst6 is set equal to 0.

[0618] 2.3.4.2. Conventional merge list

[0619] Different from the merge list design, in VVC, a history-based motion vector prediction (HMVP) method is adopted.

[0620] In HMVP, the motion information of the previously coded blocks is stored. The motion information of the previously coded blocks is defined as HMVP candidates. Multiple HMVP candidates are stored in a table called HMVP table, and this table is dynamically maintained during the encoding / decoding process. When encoding / decoding a new slice starts, the HMVP table is cleared. Whenever there is an inter-coded block, the associated motion information is added to the last entry of the table as a new HMVP candidate. Figure 25 The overall encoding and decoding flow is depicted in .

[0621] HMVP candidates can be used in the AMVP and merge candidate list construction process. Figure 26 The modified merge candidate list construction process is depicted (shown using a dotted box). When the merge candidate list is not full after the TMVP candidate is inserted, the HMVP candidates stored in the HMVP table can be used to fill the merge candidate list. Taking into account that a block usually has a higher correlation with the nearest neighboring block in terms of motion information, the HMVP candidates in the table are inserted in descending order of index. The last entry in the table is added to the list first, and the first entry is added to the end. Similarly, redundancy elimination also applies to HMVP candidates. Once the total number of available merge candidates reaches the maximum number of merge candidates allowed to be signaled, the merge candidate list construction process is terminated.

[0622] Figure 25 Candidate positions for the affine merge mode are shown.

[0623] Figure 26The modified merge list construction process is shown.

[0624] 2.3.5.JVET-N0236

[0625] This paper proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the prediction samples are refined by adding differences derived from the optical flow equation. This is called optical flow prediction refinement (PROF). The proposed method achieves pixel-level granularity for inter-frame prediction without increasing memory access bandwidth.

[0626] To achieve finer motion compensation granularity, this paper proposes a method for refining sub-block-based affine motion compensation predictions using optical flow. After performing sub-block-based affine motion compensation, the luma prediction samples are refined by adding the difference values derived from the optical flow equation. The proposed PROF (Optical Flow Prediction Refinement) is described in the following four steps.

[0627] Step 1) Perform sub-block based affine motion compensation to generate sub-block predictions I(I,j).

[0628] Step 2) Use a 3-tap filter [-1, 0, 1] to calculate the spatial gradient g of the sub-block prediction at each sample point x (i,j) and g y (i,j).

[0629] g x (i,j)=I(i+1,j)-I(i-1,j)

[0630] g y (i,j)=I(i,j+1)-I(i,j-1)

[0631] For gradient calculation, the sub-block prediction is extended by one pixel on each side. To reduce memory bandwidth and complexity, pixels on the extended boundary are copied from the nearest integer pixel location in the reference image. Thus, additional interpolation in the padded area is avoided.

[0632] Step 3) Calculate the brightness prediction refinement through the optical flow equation.

[0633] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j)

[0634] where Δv(i,j) is the difference between the pixel MV calculated for the sample position (i,j) (denoted by v(i,j)) and the sub-block MV of the sub-block to which the pixel (i,j) belongs, as Figure 27shown.

[0635] Since the affine model parameters and the pixel position relative to the sub-block center are not changed between sub-blocks, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let x and y be the horizontal and vertical offsets from the pixel position to the sub-block center, Δv(i,j) can be derived by the following equation,

[0636]

[0637] For a 4-parameter affine model,

[0638]

[0639] For the 6-parameter affine model,

[0640]

[0641] Where (v 0x ,v 0y )、(v 1x ,v 1y )、(v 2x ,v 2y ) are the left top, right top, and left bottom control point motion vectors, w and h are the width and height of the CU.

[0642] Step 4) Finally, the luma prediction refinement is added to the sub-block prediction I(i,j). The final prediction I' is generated as the following equation.

[0643] I′(i,j)=I(i,j)+ΔI(i,j)

[0644] 2.3.6. PCT / CN2018 / 125420 and PCT / CN2018 / 116889 Regarding Improvements to ATMVP

[0645] In these documents, several methods for making the design of ATMVP more rational and efficient have been disclosed, and their entire contents are incorporated herein by reference.

[0646] 2.3.7.MMVD of VVC

[0647] The Ultimate Motion Vector Expression (UMVE) is adopted in VVC and the Audio Video Standard (AVS). UMVE is also known as MVDMerge (MMVD). UMVE uses the proposed motion vector expression method for skip or merge mode.

[0648] UMVE reuses the same merge candidates as those included in the regular merge candidate list in VVC. Among the merge candidates, basic candidates can be selected and further extended by the proposed motion vector expression method. See for the UMVE process Figure 30 and Figure 31 .

[0649] UMVE provides a new method for representing motion vector difference (MVD), in which the starting point, motion magnitude and motion direction are used to represent MVD.

[0650] The proposed technique uses the merge candidate list as is, but only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE expansion.

[0651] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.

[0652] Table 1. Basic candidate IDX\

[0653] Basic Candidate IDX 0 1 2 3 Nth MVP First MVP Second MVP The third MVP Fourth MVP

[0654] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled. In VVC, there are two basic candidates.

[0655] The distance index is the motion magnitude information. The distance index indicates the predefined distance from the start point information. The predefined distances are as follows:

[0656] Table 2a. Distance IDX

[0657] Distance from IDX 0 1 2 3 4 5 6 7 Pixel distance 1 / 4 pixel 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels 16 pixels 32 pixels

[0658] During the entropy encoding and decoding process, the distance IDX is binarized using a truncated unary code in binary bits as follows:

[0659] Table 2b. Distance IDX binarization

[0660] Distance from IDX 0 1 2 3 4 5 6 7 binary bit 0 10 110 1110 11110 111110 1111110 1111111

[0661] In arithmetic coding, the first bin is coded with a probability context, and the following bins are coded with an equal probability model (also known as bypass coding).

[0662] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent one of the four directions shown below.

[0663] Table 3. Direction IDX

[0664] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0665] The UMVE flag is signaled immediately after the skip or merge flag is sent. If the skip or merge flag is true, the UMVE flag is parsed. If the UMVE flag is 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is 1, it indicates AFFINE mode, but if it is not 1, the skip / merge index is parsed for the VTM's skip / merge mode.

[0666] 2.3.8. Affine MMVD proposed in PCT / CN2018 / 115633

[0667] In PCT / CN2018 / 115633 (incorporated herein by reference), it is suggested

[0668] 1-way affine merge codec block signaling flag to indicate whether the merged affine model should be further modified. The affine model can be defined by CPMV or affine parameters.

[0669] 2 If the merged affine model is indicated to be modified for the affine merge codec block, one or more modification indices are signaled.

[0670] 3 If the merged affine model is indicated as modified for the affine merge codec block, the CPMV (i.e., MV0, MV1 for the 4-parameter affine model, MV0, MV1, MV2 for the 6-parameter affine model) can add an offset of MV0 derived from the signaled modification index (for MV0, expressed as Off0 = (Off0x, Off0y), for MV1, expressed as Off1 = (Off1x, Off1y), and for MV2, expressed as Off2 = (Off2x, Off2y)), or one or more direction indices and one or more distance indices.

[0671] If the merged affine model is indicated to be modified for the affine merge codec block, the parameters (i.e., a, b, e, f for the 4-parameter affine model and a, b, c, d, e, f for the 6-parameter affine model) are added with an offset (Offa for a, Offb for b, Offc for c, Offd for d, Offe for e and Offf for f), derived from the signaled modification index, or one or more symbol flags and one or more distance indices.

[0672] 3. Examples of Problems Solved by the Embodiments

[0673] In the current design of VVC, the sub-block based prediction mode has the following problems:

[0674] 1) The Affine AMVR flag in SPS may be turned on when the regular AMVR is off;

[0675] 2) When affine mode is off, the affine AMVR flag in SPS may be turned on;

[0676] 3) When ATMVP is not applied, MaxNumSubblockMergeCand is not set appropriately.

[0677] 4) When TMVP is disabled for a slice and ATMVP is enabled for a sequence, the co-located pictures of B slices are not identified, however, co-located pictures are required in the ATMVP process.

[0678] 5) Both TMVP and ATMVP need to obtain motion information from reference pictures. In the current design, they are assumed to be the same, which may be suboptimal.

[0679] 6) PROF should have flag to control it on / off.

[0680] 4. Example Embodiments

[0681] The following detailed inventions should be considered as examples to explain the general concept. These inventions should not be interpreted in a narrow way. In addition, these inventions can be combined in any way.

[0682] The method described below may also be applicable to other types of motion candidate lists (such as AMVP candidate lists).

[0683] 1. Whether to signal the control information of the affine AMVR may depend on whether the affine prediction is applied.

[0684] a) In one example, if affine prediction is not applied, the control information of affine AMVR is not signaled.

[0685] b) In one example, if affine prediction is not applied in a conforming bitstream, affine AMVR should be disabled (eg, the use of affine AMVR should be signaled as an error).

[0686] c) In one example, if affine prediction is not applied, the signaled control information of the affine AMVR may be ignored and inferred to be not applied.

[0687] 2. Whether to signal the control information of the affine AMVR may depend on whether the conventional AMVR is applied.

[0688] a) In one example, if conventional AMVR is not applied, no affine

[0689] AMVR control information.

[0690] b) In one example, if regular AMVR is not applied in the conforming bitstream, then affine AMVR should be disabled (eg, use of the affine AMVR flag is signaled as false).

[0691] c) In one example, if conventional AMVR is not applied, the control information signaled by affine AMVR may be ignored and inferred to be not applied.

[0692] d) In one example, an indication of adaptive motion vector resolution (e.g., a flag) can be signaled in a sequence / picture / slice / slice group / slice / brick / other video unit to control the use of AMVR (e.g., conventional AMVR (i.e., AMVR applied to translational motion) and affine AMVR (i.e., AMVR applied to affine motion)) for multiple codec methods.

[0693] i. In one example, such indication may be signaled in the SPS / DPS / VPS / PPS / picture header / slice header / slice group header.

[0694] ii. Alternatively, in addition, whether to signal an indication of using conventional AMVR and / or affine AMVR may depend on the indication.

[0695] 1) In one example, when such indication indicates that adaptive motion vector resolution is disabled, signaling indicating the use of conventional AMVR may be skipped.

[0696] 2) In one example, when such indication indicates that adaptive motion vector resolution is disabled, the signaling indicating the use of affine AMVR may be skipped.

[0697] iii. Alternatively, furthermore, whether to signal an indication of use of affine AMVR may depend on the indication and use of affine prediction mode.

[0698] 1) For example, if the affine prediction mode is disabled, such indication may be skipped.

[0699] iv. In one example, if the current slice / slice group / picture can only be predicted from the previous pictures and may be derived as false, such indication may not be signaled.

[0700] v. In one example, if the current slice / slice group / picture can only be predicted from the following pictures and may be derived as false, such indication may not be signaled.

[0701] vi. In one example, when the current slice / slice group / picture can be predicted from the previous and following pictures, such indication can be signaled.

[0702] 3. Whether to signal the control information of affine AMVR may depend on whether conventional AMVR is applied, and whether affine prediction is applied.

[0703] a) In one example, if affine prediction is not applied or conventional AMVR is not applied, the control information of affine AMVR is not signaled.

[0704] i. In one example, if affine prediction is not applied in a conforming bitstream or regular AMVR is not applied in a conforming bitstream, then affine AMVR should be disabled (eg, use of affine AMVR should be signaled as an error).

[0705] ii. In one example, if affine prediction is not applied or conventional AMVR is not applied, the signaled control information of affine AMVR may be ignored and inferred to be not applied.

[0706] b) In one example, if affine prediction is not applied and conventional AMVR is not applied, the control information of affine AMVR is not signaled.

[0707] i. In one example, if affine prediction is not applied in a conforming bitstream and regular AMVR is not applied in a conforming bitstream, then affine AMVR should be disabled (eg, use of affine AMVR should be signaled as an error).

[0708] ii. In one example, if affine prediction is not applied and conventional AMVR is not applied, the signaled control information of affine AMVR may be ignored and inferred to be not applied.

[0709] 4. The maximum number of candidates in the sub-block merge candidate list (denoted as MaxNumSubblockMergeCand) may depend on whether ATMVP is enabled. Whether ATMVP is enabled may only be indicated by sps_sbtmvp_enabled_flag in SPS.

[0710] a) For example, whether ATMVP is enabled may depend not only on a flag signaled in the sequence level (e.g., sps_sbtmvp_enabled_flag in SPS), but may also depend on one or more syntax elements signaled in any other video unit at the sequence / picture / slice / slice group / slice level, such as VPS, DPS, APS, PPS, slice header, slice group header, picture header, etc.

[0711] i. Alternatively, whether ATMVP is enabled can be derived implicitly without signaling.

[0712] ii. For example, if TMVP is not enabled for a picture, slice, or slice group, ATMVP will not be enabled for the picture, slice, or slice group.

[0713] b) For example, whether and / or how syntax elements related to MaxNumSubblockMergeCand (eg, five_minus_max_num_subblock_merge_cand) are signaled may depend on whether ATMVP is enabled.

[0714] i. For example, if ATMVP is not enabled, five_minus_max_num_subblock_merge_cand may be restricted in the conforming bitstream.

[0715] 1) For example, if ATMVP is not enabled, five_minus_max_num_subblock_merge_cand is not allowed to be equal to a fixed number. In both examples, the fixed number can be 0 or 5.

[0716] 2) For example, if ATMVP is not enabled, five_minus_max_num_subblock_merge_cand is not allowed to be greater than a fixed number. In one example, the fixed number may be 4.

[0717] 3) For example, if ATMVP is not enabled, five_minus_max_num_subblock_merge_cand is not allowed to be less than a fixed number. In one example, the fixed number may be 1.

[0718] ii. For example, when five_minus_max_num_subblock_merge_cand does not exist, five_minus_max_num_subblock_merge_cand may be set to five_minus_max_num_subblock_merge_cand-(ATMVP enabled? 0:1), where whether ATMVP is enabled depends not only on a flag in the SPS (eg, sps_sbtmvp_enabled_flag).

[0719] c) For example, MaxNumSubblockMergeCand may be derived based on one or more syntax elements (eg, five_minus_max_num_subblock_merge_cand) and whether ATMVP is enabled.

[0720] i. For example, MaxNumSubblockMergeCand can be derived as MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand-(ATMVP enabled? 0:1).

[0721] d) When ATMVP is enabled and affine motion prediction is disabled, MaxNumSubblockMergeCand can be set to 1.

[0722] 5. A default candidate (with translation and / or affine motion) can be added to the sub-block merge candidate list. The default candidate can use a prediction type, such as sub-block prediction or full-block prediction.

[0723] a) In one example, the full-block prediction of the default candidate may follow a translational motion model (eg, the full-block prediction of a regular merge candidate).

[0724] b) In one example, the sub-block prediction of the default candidate may follow a translational motion model (eg, the sub-block prediction of the ATMVP candidate).

[0725] c) In one example, the sub-block prediction of the default candidate may follow an affine motion model (eg, the sub-block prediction of the affine merge candidate).

[0726] d) In one example, the default candidate may have an affine flag equal to 0.

[0727] i. Alternatively, the default candidate may have an affine flag equal to 1.

[0728] e) In one example, subsequent processing on a block may depend on whether the block is coded with a default candidate.

[0729] i. In one example, the block is considered to be coded using whole-block prediction (e.g., the selected default candidate is predicted using whole-block prediction), and

[0730] 1) For example, PROF may not be applied to a block.

[0731] 2) For example, DMVR (decoding side motion vector refinement) can be applied to blocks.

[0732] 3) For example, BDOF (Bidirectional Optical Flow) can be applied to the blocks.

[0733] 4) For example, deblocking filtering may not be applied to boundaries between sub-blocks in a block.

[0734] ii. In one example, the block is considered to be coded using sub-block prediction (e.g., the selected default candidate is sub-block predicted), and

[0735] 1) For example, PROF can be applied to a block.

[0736] 2) For example, DMVR (decoding side motion vector refinement) may not be applied to the block.

[0737] 3) For example, BDOF (Bidirectional Optical Flow) may not be applied to a block.

[0738] 4) For example, deblocking filtering can be applied to the boundaries between sub-blocks in a block.

[0739] iii. In one example, the block is considered to be coded using translational prediction, and

[0740] 1) For example, PROF may not be applied to a block.

[0741] 2) For example, DMVR (decoding side motion vector refinement) can be applied to blocks.

[0742] 3) For example, BDOF (Bidirectional Optical Flow) can be applied to the blocks.

[0743] 4) For example, deblocking filtering may not be applied to boundaries between sub-blocks in a block.

[0744] iv. In one example, the block is considered to be encoded and decoded using affine prediction, and

[0745] 1) For example, PROF can be applied to a block.

[0746] 2) For example, DMVR (decoding side motion vector refinement) may not be applied to the block.

[0747] 3) For example, BDOF (Bidirectional Optical Flow) may not be applied to a block.

[0748] 4) For example, deblocking filtering can be applied to the boundaries between sub-blocks in a block.

[0749] f) In one example, one or more default candidates may be put into the sub-block merge candidate list.

[0750] i. For example, the first type of default candidates using whole-block prediction and the second type of default candidates using sub-block prediction can both be put into the sub-block merge candidate list.

[0751] ii. For example, the first type of default candidates using translation prediction and the second type of default candidates using affine prediction can both be put into the sub-block merge candidate list.

[0752] iii. The maximum number of default candidates of each type may depend on whether ATMVP is enabled and / or whether affine prediction is enabled.

[0753] g) In one example, for B slices, the default candidate may have zero motion vectors for all sub-blocks, bi-prediction is applied, and both reference picture indices are set to 0.

[0754] h) In one example, for P slices, the default candidate may have zero motion vectors for all subblocks, uni-prediction is applied, and the reference picture index is set to 0.

[0755] i) Which default candidate is put into the sub-block merge candidate list may depend on the use of ATMVP and / or affine prediction mode.

[0756] i. In one example, when affine prediction mode is enabled, a default candidate with an affine motion model with an affine flag equal to 1 (eg, all CPMVs equal to 0) may be added.

[0757] ii. In one example, when affine prediction mode and ATMVP are enabled, default candidates with translational motion models having an affine flag equal to 0 (e.g., zero MV) and / or default candidates with affine motion models having an affine flag equal to 1 (e.g., all CPMVs equal to 0) may be added.

[0758] 1) In one example, a default candidate with a translational motion model may be added before a default candidate with an affine motion model.

[0759] iii. In one example, when affine prediction and ATMVP are enabled, default candidates with translational motion models having an affine flag equal to 0 may be added, while default candidates with affine motion models are not added.

[0760] j) When the sub-block merge candidate is not satisfied after checking the ATMVP candidate and / or the spatial / temporal / constructed affine merge candidate, the above method can be applied.

[0761] 6. Information about ATMVP (eg, whether ATMVP is enabled for a slice or slice group or picture) may be signaled in a slice header or slice group header or slice header.

[0762] a) In one example, the collocated pictures for ATMVP may be different from the collocated pictures used for TMVP.

[0763] b) In one example, information about ATMVP may not be signaled for I slice or I slice group or I picture.

[0764] c) In one example, information about ATMVP may be signaled only if ATMVP is signaled to be enabled at the sequence level (e.g., sps_sbtmvp_enabled_flag is equal to 1).

[0765] d) In one example, if TMVP is disabled for a slice or slice group or picture, information about ATMVP may not be signaled for the slice or slice group or picture.

[0766] i. For example, in this case, ATMVP may be inferred to be prohibited.

[0767] e) If TMVP is disabled for a slice (or slice group or picture), ATMVP may be inferred to be disabled for the slice (or slice group or picture) regardless of information signaled using ATMVP.

[0768] 7. Whether to add sub-block based temporal merging candidates (eg, temporal affine motion candidates) may depend on the use of TMVP.

[0769] a) Alternatively, it may depend on the value of sps_temporal_mvp_enabled_flag.

[0770] b) Alternatively, it may depend on the value of slice_temporal_mvp_enabled_flag.

[0771] c) When sps_temporal_mvp_enabled_flag or slice_temporal_mvp_enabled_flag is true, sub-block based temporal merging candidates may be added to sub-block merging candidates.

[0772] i. Alternatively, when sps_temporal_mvp_enabled_flag and slice_temporal_mvp_enabled_flag are both true, sub-block based temporal merging candidates may be added to the sub-block merging candidates.

[0773] ii. Alternatively, when sps_temporal_mvp_enabled_flag or slice_temporal_mvp_enabled_flag is false, sub-block based temporal merging candidates should not be added to the sub-block merge candidates.

[0774] d) Alternatively, an indication of adding sub-block based time domain merging candidates may be signaled in a sequence / picture / slice / slice group / slice / brick / other video unit.

[0775] i. Alternatively, furthermore, temporal motion vector prediction (eg, sps_temporal_mvp_enabled_flag and / or slice_temporal_mvp_enabled_flag) may be signaled conditionally depending on its use.

[0776] 8. Indication of co-located reference pictures, for example, depending on the use of multiple codecs that require access to temporal motion information, it is possible to conditionally signal from which reference picture list the co-located reference picture is derived (e.g., collocated_from_l0_flag) and / or the reference index of the co-located reference picture.

[0777] a) In one example, the condition is that one of ATMVP or TMVP is enabled.

[0778] b) In one example, the condition is that one of ATMVP or TMVP or affine motion information prediction is enabled.

[0779] 9. In one example, sub-block based temporal merging candidates can be put into the sub-block merge candidate list only when both ATMVP and TMVP are enabled for the current picture / slice / slice group.

[0780] a) Alternatively, sub-block based temporal merging candidates can be put into the sub-block merge candidate list only when ATMVP is enabled for the current picture / slice / slice group.

[0781] 10.MaxNumSubblockMergeCand may depend on whether sub-block based temporal merging candidates can be used.

[0782] a) Alternatively, MaxNumSubblockMergeCand may depend on whether TMVP can be used.

[0783] b) For example, if the sub-block based time domain merging candidate (or TMVP) cannot be used, then MaxNumSubblockMergeCand shall not be greater than 4.

[0784] i. For example, if the sub-block based time domain merging candidate (or TMVP) cannot be used and ATMVP can be used, then MaxNumSubblockMergeCand shall not be greater than 4.

[0785] ii. For example, if the sub-block based time domain merging candidate (or TMVP) cannot be used and the ATMVP cannot be used, then MaxNumSubblockMergeCand shall not be greater than 4.

[0786] iii. For example, if the sub-block based time domain merging candidate (or TMVP) cannot be used and the ATMVP cannot be used, then MaxNumSubblockMergeCand shall not be greater than 3.

[0787] 11. One or more syntax elements indicating whether and / or how to perform PROF can be signaled in any video unit at the sequence / picture / slice / slice group / slice / CTU row / CTU / CU / PU / TU level, such as VPS, DPS, APS, PPS, slice header, slice group header, picture header, CTU, CU, PU, etc.

[0788] a) In one example, one or more syntax elements (eg, a flag indicating whether PROF is enabled) may be conditionally signaled based on other syntax elements (eg, a syntax element indicating whether affine prediction is enabled).

[0789] i. For example, when affine prediction is disabled, the syntax element indicating whether PROF is enabled may not be signaled, and PROF is inferred to be disabled.

[0790] b) In one example, when affine prediction is disabled in a conforming bitstream, the syntax element indicating whether PROF is enabled must be set to PROF disabled.

[0791] c) In one example, when affine prediction is disabled in a conforming bitstream, the syntax element signaling whether PROF is enabled is ignored and PROF is inferred to be disabled.

[0792] d) In one example, a syntax element may be signaled to indicate whether PROF applies to uni-prediction only.

[0793] Affine algorithm based on MMVD

[0794] Method for Affine with MMVD

[0795] 12. The decoded affine motion information (eg, due to AMVP affine / Merge affine mode) may be further modified with motion vector differences and / or reference picture indices further signaled in the bitstream.

[0796] a) Alternatively, the motion vector differences and / or reference picture indices may be derived dynamically.

[0797] b) Alternatively, furthermore, the above method may be called only when the current block is encoded or decoded using the affine merge mode (excluding the AMVP affine mode).

[0798] c) Alternatively, furthermore, the above method may be called only when the decoded merge index corresponds to a sub-block merge candidate (eg, an affine merge candidate).

[0799] 13. If the affine merge candidate is indicated as modified for the affine merge codec block, then N CPMVs are modified, where N is an integer value.

[0800] a) In one example, N is equal to 2.

[0801] b) In one example, N is equal to 3.

[0802] c) In one example, N depends on the affine model. For example, if a 4-parameter model is used, N is equal to 2. If a 6-parameter model is used, N is equal to 3.

[0803] d) Alternatively, in addition, a syntax may be signaled to indicate whether all CPMVs associated with the block need to be modified.

[0804] i. Alternatively, furthermore, the syntax may be signaled conditionally, for example depending on whether the current block is coded in affine mode, affine AMVP mode or affine Merge mode.

[0805] e) Alternatively, furthermore, the syntax may be context coded or bypass coded.

[0806] 14. In one example, whether to modify the first CPMV of an affine merge candidate can be controlled separately from the second CPMV.

[0807] a) In one example, a first flag may be signaled to indicate whether a first CPMV of an affine merge candidate is modified, and a second flag may be signaled to indicate whether a second CPMV of an affine merge candidate is modified.

[0808] i. In one example, these two flags can be signaled independently.

[0809] b) In one example, a syntax element may be signaled to indicate whether each CPMV of an affine merge candidate is modified.

[0810] i. For example, the first value of the syntax element may indicate that no CPMV is modified. The second value of the syntax element may indicate that the first CPMV is modified but the second CPMV is not modified. The third value of the syntax element may indicate that the first CPMV is not modified but the second CPMV is modified. The fourth value of the syntax element may indicate that the first CPMV is modified and the second CPMV is modified.

[0811] c) In one example, a syntax element may be signaled to indicate whether at least one of a plurality of CPMVs associated with the current block is modified.

[0812] i. If the syntax element indicates that at least one is to be modified, then for each of the CPMVs, another syntax element may be further signaled.

[0813] d) Alternatively, furthermore, the syntax / flags may be context coded or bypass coded.

[0814] 15. In one example, whether the affine merge candidate is modified may be signaled first, and then the candidate index may be signaled.

[0815] 16. In one example, the affine merge candidate index may be signaled first, and then whether the affine merge candidate is modified may be signaled.

[0816] 17. Only the first K affine merge candidates can be indicated to be modified, where K is an integer greater than 0 and not greater than the number of allowed affine merge candidates.

[0817] a) In one example, it may be first signaled whether the affine merge candidate is modified, and then the candidate index may be signaled with a maximum value K-1.

[0818] b) In one example, the affine merge candidate index may be signaled first, and then only when the affine merge candidate index is not greater than K-1, it may be signaled whether the affine merge candidate is modified.

[0819] c) Alternatively, only the last K affine merge candidates may be indicated to be modified.

[0820] i. In one example, the affine merge candidate index may be signaled first, and then only if the selected affine merge candidate is from one of the last K affine merge candidates, whether the affine merge candidate is modified may be signaled.

[0821] 18. If the merged affine model is indicated to be modified for the affine merge codec block, the CPMV (i.e., MV0, MV1 for the 4-parameter affine model, MV0, MV1, MV2 for the 6-parameter affine model) can be added with an offset (for MV0, expressed as Off0 = (Off0x, Off0y), for MV1, expressed as Off1 = (Off1x, Off1y), and for MV2, expressed as Off2 = (Off2x, Off2y)), derived from the signaled modification index, or one or more direction indices and one or more distance indices.

[0822] a) In one example, Off0′=(Off0x′, Off0y′) for MV0 and / or Off1′=(Off1x′, Off1y′) for MV1 and / or Off2′=(Off2x′, Off2y′) for MV2 are first derived using one or more direction indices and one or more distance indices, wherein OffQ′ is derived depending on the direction index and / or distance index associated with MVQ, independently of the direction index and / or distance index associated with CPMV other than MVQ, where Q is equal to 0, 1, or 2, and then Off0 and / or Off1 and / or Off2 are derived based on Off0′ and / or Off1′ and / or Off2′. In other words, OffQ can be derived based on OffQ′ and OffP′, where P and Q belong to {0, 1, 2} and Q may not be equal to P.

[0823] i. In one example, OffQ' is derived from the direction index and / or distance index associated with the MVQ, in the same way that the MVD is derived in the MMVD method.

[0824] ii. In one example, OffQ=OffQ′, where Q is an integer such as 0, 1, or 2.

[0825] iii. In one example, OffQ=Off0′+OffQ′, where Q is an integer such as 1 or 2.

[0826] 19. The modified CPMV can be stored and used to process subsequent blocks.

[0827] a) In one example, the subsequent block is in the same picture as the current block.

[0828] b) In one example, the subsequent block and the current block are located in different pictures.

[0829] c) Alternatively, the modified CPMV is not stored, but the decoded CPMV is stored and used to process subsequent blocks.

[0830] 20. An indication of whether the CPMV needs to be modified for the current block (eg, indicated by MMVD_AFFINE_flag) may be used to process subsequent blocks.

[0831] a) In one example, the MMVD_AFFINE_flag may be inherited by subsequent blocks.

[0832] b) In one example, the MMVD_AFFINE_flag can be used to model the context for encoding and decoding the indication of subsequent blocks.

[0833] 21. The modified affine merge candidate can be inserted into the affine merge list (or sub-block merge list), and the affine merge index (or sub-block merge index) can be used to indicate whether to use or / and which modified affine merge candidate to use.

[0834] 22. New affine merge candidates can be generated from N (N>1) existing affine merge candidates and can be further inserted into the affine merge list or sub-block merge list.

[0835] a) In one example, the CPMVs of N (N>1) affine merge candidates may be averaged or weighted averaged to generate a new affine merge candidate. For example, N=2 or 3 or 4.

[0836] 23. In one example, the modified distance of CPMV may be signaled as zero.

[0837] a) In one example, if the modified distance is signaled as zero, the modified direction is not signaled.

[0838] 24. In one example, a flag is signaled to indicate whether the CPMV of no affine merge candidate is modified.

[0839] a) In one example, if the CPMV indicating no affine merge candidate is modified, no modification information is signaled for the affine merge candidate.

[0840] b) In one example, if at least one CPMV indicating an affine merge candidate is modified, and all CPMVs before the last CPMV are indicated as not modified, it is inferred that the last CPMV is modified with a distance greater than zero without signaling.

[0841] 5. Examples

[0842] For all the following embodiments, syntax elements may be signaled at different levels (eg, in SPS / PPS / slice header / picture header / slice group header / slice or other video units).

[0843] 5.1. Example #1: Example of Syntax Design for sps_affine_amvr_enabled_flag in SPS / PPS / Slice Header / Slice Group Header

[0844]

[0845] Alternatively,

[0846]

[0847] Alternatively,

[0848]

[0849]

[0850] Alternatively,

[0851]

[0852] 5.2. Example #2: Example of semantics on five_minus_max_num_subblock_merge_cand

[0853] five_minus_max_num_subblock_merge_cand specifies the maximum number of sub-block based merging motion vector prediction (MVP) candidates supported in a slice, subtracted from 5. When five_minus_max_num_subblock_merge_cand is not present, it is inferred to be equal to 5 – (sps_sbtmvp_enabled_flag && slice_temporal_mvp_enabled_flag). The maximum number of sub-block based merging MVP candidates, MaxNumSubblockMergeCand, is derived as follows: MaxNumSubblockMergeCand = 5 – five_minus_max_num_subblock_merge_cand. The value of MaxNumSubblockMergeCand should be in the range of 0 to 5, inclusive.

[0854] Alternatively, the following may apply:

[0855] When ATMVP is enabled and affine is disabled, the value of MaxNumSubblockMergeCand should be in the range of 0 to 1 (inclusive). When affine is enabled, the value of MaxNumSubblockMergeCand should be in the range of 0 to 5 (inclusive).

[0856] 5.3. Example #3: Example of syntax elements of ATMVP in a slice header (or slice group header)

[0857]

[0858]

[0859] Alternatively,

[0860]

[0861] 5.4. Example #4: Example of syntax elements for sub-block-based time-domain merging candidates in a slice header (or slice group header)

[0862] if(slice_temporal_mvp_enabled_flag&&MaxNumSubblockMergeCand>0) sub_block_tmvp_merge_candidate_enalbed_flag u(1)

[0863] sub_block_tmvp_merge_candidate_enalbed_flag specifies whether sub-block based temporary merging candidates can be used.

[0864] If not present, sub_block_tmvp_merge_candidate_enalbed_flag is inferred to be 0.

[0865] 8.5.5.6 Derivation of Constructed Affine Control Point Motion Vector Merging Candidates

[0866]

[0867] The fourth (collocated bottom right) control point motion vector cpMvLXCorner[3], reference index refIdxLXCorner[3], prediction list utilization flag predFlagLXCorner[3], bi-prediction weight index bcwIdxCorner[3], and availability flag availableFlagCorner[3] are derived as follows, where X is 0 and 1:

[0868] – The reference index refIdxLXCorner[3] (X is 0 or 1) of the time-domain merging candidate is set equal to 0.

[0869] – Export the variables mvLXCol and availableFlagLXCol (where X is 0 or 1) as follows:

[0870] – If sub_block_tmvp_merge_candidate_enalbed_flag is equal to 0, or both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0.

[0871] – Otherwise (sub_block_tmvp_merge_candidate_enalbed_flag is equal to 1), the following applies:

[0872]

[0873] 5.5. Example #5: Example of syntax / semantics for controlling PROF in SPS

[0874] sps_affine_enabled_flag u(1) if(sps_affine_enabled_flag) sps_prof_flag u(1)

[0875] sps_prof_flag specifies whether PROF can be used for inter prediction. If sps_prof_flag is equal to 0, PROF is not applied. Otherwise (sps_prof_flag is equal to 1), PROF is applied. When not present, the value of sps_prof_flag is inferred to be equal to 0.

[0876] Figure 281 is a block diagram of a video processing device 1400. Device 1400 can be used to implement one or more methods described herein. Device 1400 can be embodied in a smartphone, tablet computer, computer, Internet of Things (IoT) receiver, etc. Device 1400 may include one or more processors 1402, one or more memories 1404, and video processing hardware 1406. Processor 1402 can be configured to implement one or more methods described in this document. One or more memories 1404 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 1406 can be used to implement some of the techniques described in this document in hardware circuits. In some embodiments, video processing hardware 1406 can be partially or entirely located within processor 1402 (e.g., a graphics processor).

[0877] Figure 29 is a flow chart of an example method 2900 for video processing. The method 2900 includes performing (2902) a conversion between a current video block of a video and a bitstream representation of the video using an affine adaptive motion vector resolution technique, such that the bitstream representation selectively includes control information related to the affine adaptive motion vector resolution technique based on a rule.

[0878] The following list of examples provides embodiments that can solve the technical problems described in this document as well as other problems.

[0879] 1. A method for video processing, comprising: using an affine adaptive motion vector resolution technique to perform conversion between a current video block of a video and a bitstream representation of the video, so that the bitstream representation selectively includes control information related to the affine adaptive motion vector resolution technique based on a rule.

[0880] 2. The method of example 1, wherein the rule provides for including the control information if affine prediction is used during the conversion, and omitting the control information if affine prediction is not used during the conversion.

[0881] 3. The method of example 1, wherein the rule further provides that, in the case where affine prediction is not applied to the transformation, an adaptive motion vector resolution step is used for exclusion during the transformation.

[0882] Additional examples and embodiments related to the above examples are provided in Section 4, Item 1.

[0883] 4. The method of example 1, wherein the rule provides for including or omitting the control information based on whether a conventional adaptive motion vector resolution step is used during conversion.

[0884] 5. The method of example 4, wherein the rule provides for omitting the control information if a conventional adaptive motion vector resolution step is not applied during conversion.

[0885] 6. The method of example 1, wherein the control information includes a same field indicating use of multiple adaptive motion vector resolution techniques during conversion.

[0886] Additional examples and embodiments related to the above examples are provided in Section 4, Item 2.

[0887] 7. The method of example 1, wherein the rule provides for including or omitting the control information based on whether conventional adaptive motion vector resolution and affine prediction are used during conversion.

[0888] 8. The method of example 7, wherein the rule provides for omitting the control information if conventional adaptive motion vector resolution and affine prediction are not applied during conversion.

[0889] Additional examples and embodiments related to the above examples are provided in Section 4, Item 3.

[0890] 9. A method of video processing, comprising: during conversion between a current video block and a bitstream representation, determining a sub-block merge candidate list for conversion, wherein a maximum number of candidates in the sub-block merge candidate list depends on whether alternative temporal motion vector prediction (ATMVP) is applied to the conversion; and performing the conversion using the sub-block merge candidate list.

[0891] 10. The method of example 9, wherein a field in the bitstream representation indicates whether the alternative temporal motion vector prediction should be used for the conversion.

[0892] 11. The method of example 10, wherein the field is at a sequence level, a video parameter set level, a picture parameter set level, a slice level, a slice group level, or a picture header level.

[0893] 12. The method of example 9, wherein, if ATMVP is applied to the transformation and affine prediction is disabled for the transformation, the maximum number of candidates is set to 1.

[0894] Additional examples and embodiments related to the above examples are provided in Section 4, Item 4.

[0895] 13. A method of video processing, comprising: during conversion between a current video block and a bitstream representation, appending one or more default merge candidates to a sub-block merge candidate list for conversion; and performing the conversion using the sub-block merge candidate list to which the one or more default merge candidates are appended.

[0896] 14. The method of example 13, wherein the default candidate is associated with the sub-block prediction type.

[0897] 15. The method of example 14, wherein the sub-block prediction type comprises prediction based on a translational motion model or an affine motion model.

[0898] 16. The method of example 13, wherein the default candidate is associated with the whole block prediction type.

[0899] 17. The method of example 14, wherein the whole block prediction type comprises prediction based on a translational motion model or an affine motion model.

[0900] Additional examples and embodiments related to the above examples are provided in Section 4, Item 5.

[0901] 18. A method of video processing, comprising: during a conversion between a current video block of a video and a bitstream representation, determining the suitability of an alternative temporal motion vector prediction (ATMVP) for the conversion, wherein one or more bits in the bitstream representation correspond to the determination; and performing the conversion based on the determination.

[0902] 19. The method of example 18, wherein the one or more bits are included at a picture header, a slice header, or a slice group header.

[0903] 20. The method of examples 18-19, wherein converting a collocated picture using ATMVP is different from another collocated picture used for video conversion using temporal motion vector prediction (TMVP).

[0904] Additional examples and embodiments related to the above examples are provided in Section 4, Item 6.

[0905] 21. A method of video processing, comprising: selectively constructing a sub-block merge candidate list based on a condition associated with a temporal motion vector prediction (TMVP) step or an alternative temporal motion vector prediction (ATMVP); and performing conversion between a current video block and a bitstream representation of the current video block based on the sub-block merge candidate list.

[0906] 22. The method of example 21, wherein the condition corresponds to the presence of a flag in the bitstream representation at a sequence parameter set level or a slice level or a slice level or a brick level.

[0907] 23. The method of example 21, wherein the sub-block merge candidate list is constructed using the sub-block based temporal merging candidates only when alternative motion vector prediction and TMVP steps are enabled for the picture, slice or slice group to which the current video block belongs.

[0908] 24. The method of Example 21, wherein the sub-block merge candidate list is constructed using the sub-block based temporal merging candidates only when the ATMVP and TMVP steps are enabled for the picture, slice or slice group to which the current video block belongs.

[0909] 25. The method of Example 21, wherein the sub-block merge candidate list is constructed using the sub-block based temporal merging candidates only when ATMVP is enabled for the picture, slice, or slice group to which the current video block belongs and the TMVP step is disabled.

[0910] Additional examples and embodiments related to the above examples are provided in Section 4, Items 7 and 9.

[0911] 26. The method of examples 21-25, wherein the flag in the bitstream representation is included or omitted based on whether a sub-block based time-domain merging candidate is used during conversion.

[0912] Additional examples and embodiments related to the above examples are provided in Section 4, Item 10.

[0913] 27. A method of video processing, comprising: selectively performing a conversion between a current video block of a video and a bitstream representation of the video using rule-based prediction refinement using optical flow (PROF), wherein the rules include (1) including or omitting a field in the bitstream representation, or (2) whether affine prediction is applied to the conversion.

[0914] 28. The method of example 27, wherein the rule provides that PROF is prohibited due to prohibition of affine prediction for the transformation.

[0915] 29. The method of example 27, wherein, if affine prediction is disabled, PROF is inferred to be disabled for the transition.

[0916] 30. The method of example 27, wherein the rule further provides that PROF is only used for uni-prediction based on a corresponding flag in the bitstream representation.

[0917] For example, Section 4, Item 11 provides additional examples and embodiments related to the above examples.

[0918] 31. The method of example 30, wherein the flag is included at a video parameter set, a picture parameter set, an adaptation parameter set, a slice header, a slice group header, a picture header, a codec tree unit, a codec unit, or a prediction unit level.

[0919] 32. A method of video processing, comprising: during a conversion between a current video block of a video and a bitstream representation of the video, modifying decoded affine motion information of the current video block according to a modification rule; and performing the conversion by modifying the decoded affine motion information according to the modification rule.

[0920] 33. The method of example 32, wherein the modification rule specifies modifying the decoded affine motion information using motion vector differences and / or reference picture indices, and wherein the syntax element in the bitstream representation corresponds to the modification rule.

[0921] 34. The method of any of Examples 32-33, wherein modifying comprises modifying N control point motion vectors, where N is a positive integer.

[0922] 35. The method of any of Examples 32-34, wherein the modification rule provides for modifying the first control point motion vector differently than the second control point motion vector.

[0923] For example, Section 4, Items 12-24 provide additional examples and embodiments related to the above examples.

[0924] 36. The method of any of Examples 1-35, wherein converting comprises encoding the video into a bitstream representation.

[0925] 37. The method of any of Examples 1-35, wherein converting comprises parsing and decoding the bitstream representation to generate the video.

[0926] 38. A video processing device comprising a processor configured to implement one or more of Examples 1 to 37.

[0927] 39. A computer-readable medium having code stored thereon, which, when executed by a processor, causes the processor to implement any one or more of the methods of Examples 1 to 37.

[0928] In the listing of examples in this document, the term converting may refer to generating a bitstream representation for or from a current video block. The bitstream representation need not represent a contiguous group of bits and may be divided into bits included in a header field or bits included in a codeword representing coded pixel value information.

[0929] In the above examples, the rules may be predefined and the encoder and decoder are known.

[0930] Figure 32is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values) or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (e.g., Ethernet, passive optical network (PON), etc.), and wireless interfaces (e.g., Wi-Fi or cellular interfaces).

[0931] System 1900 may include a codec component 1904 that can implement various encoding or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a coded representation of the video. Therefore, codec technology is sometimes referred to as video compression or video transcoding technology. The output of codec component 1904 can be stored or transmitted via connected communication, as represented by component 1906. The stored or communicated bitstream (or coded) representation of the video received at input 1902 can be used by component 1908 to generate pixel values or displayable video that is sent to display interface 1910. The process of generating a user-viewable video from a bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it will be understood that the codec tools or operations are used at the encoder, and the corresponding decoding tools or operations that reverse the encoding results will be performed by the decoder.

[0932] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document can be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0933] Figure 33 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0934] like Figure 33 As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data, which may be referred to as a video encoding device. The destination device 120 may decode the encoded video data generated by the source device 110, which may be referred to as a video decoding device.

[0935] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0936] The video source 112 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. A codec picture is a codec representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the destination device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the destination device 120.

[0937] Destination device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0938] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, configured to interface with an external display device.

[0939] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or future standards.

[0940] Figure 34 is a block diagram illustrating an example of a video encoder 200, which may be Figure 33 The video encoder 114 in the system 100 is shown.

[0941] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 34In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0942] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 which may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214.

[0943] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is a picture in which the current video block is located.

[0944] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, may be highly integrated, but for the purpose of explanation, are not shown in FIG. Figure 5 are represented separately in the example.

[0945] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0946] The mode selection unit 203 can, for example, select one of the coding modes, i.e., intra or inter, based on the error result, and provide the resulting intra or inter coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra prediction and inter prediction (CIIP) modes, where the prediction is based on an inter prediction signal and an intra prediction signal. The mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel precision or integer pixel precision) in the case of inter prediction.

[0947] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.

[0948] For example, motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0949] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 or list 1. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0950] In other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1 containing the reference video block, and a motion vector indicating the spatial displacement between the reference video block and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0951] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder.

[0952] In some examples, motion estimation unit 204 may not output a full set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0953] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0954] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0955] As discussed above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0956] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0957] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0958] In other examples, such as in skip mode, the current video block may not have residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0959] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0960] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0961] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0962] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0963] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0964] Figure 35 is a block diagram illustrating an example of a video decoder 300, which may be Figure 33 The video decoder 114 in the system 100 is shown.

[0965] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 35 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0966] exist Figure 35 In the example of FIG, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally the reverse of the encoding pass described with respect to the video encoder 200 (e.g., Figure 34 ).

[0967] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., coded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information based on the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge modes.

[0968] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0969] Motion compensation unit 302 may calculate interpolated values of sub-integer pixels of a reference block using interpolation filters used by video encoder 20 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.

[0970] The motion compensation unit 302 may use some syntax information to determine the size of blocks used to encode frames and / or slices of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, a mode indicating how each partition is to be encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the coded video sequence.

[0971] The intra prediction unit 303 can form a prediction block based on spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0972] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation.

[0973] Some embodiments of this document are now presented in a clause-based format.

[0974] 1. A method of visual media processing (e.g., Figure 26 3600) as described in conjunction with Example 12 in Section 4 of this document, comprising:

[0975] A conversion between a current video block of visual media data and a bitstream representation of the visual media data is performed (step 3602) according to a modification rule, wherein the current video block is encoded and decoded using affine motion information; wherein the modification rule provides for modifying the affine motion information encoded and decoded in the bitstream representation using motion vector differences and / or reference picture indices during decoding.

[0976] 2. The method of clause 1, wherein the motion vector differences and / or reference picture indices used in modifying the decoded affine motion information comprise a plurality of motion vector differences and / or a plurality of reference picture indices explicitly included in the bitstream representation.

[0977] 3. The method of clause 1, wherein the motion vector difference and / or reference picture index used in modifying the decoded affine motion information is associated with an affine merge mode, wherein the motion vector difference and / or reference picture index is derived based on the motion information of the merge candidate.

[0978] 4. The method of clause 1, wherein the modification rule is applied only when the current block is processed using an affine merge mode to which an affine advanced motion vector prediction (AMVP) mode is not applied.

[0979] 5. The method of clause 1, wherein the modification rule is applied only when processing the current block using a merge index corresponding to a sub-block merge candidate.

[0980] 6. The method of any one or more of clauses 1-5, wherein the syntax elements in the bitstream representation correspond to modification rules.

[0981] 7. A method for visual media processing ( Figure 37 3700 as discussed in conjunction with Example Embodiment 13 in Section 4 of this document, comprising:

[0982] For a conversion between a current video block of visual media data and a bitstream representation of the visual media data, determining (step 3702) that an affine merge candidate for the current video block is subject to modification based on a modification rule; and

[0983] A conversion is performed (step 3704) based on the determination; wherein the modification rule provides for modifying one or more control point motion vectors associated with the current video block, and wherein an integer number of control point motion vectors is used.

[0984] 8. The method of clause 7, wherein the affine model applied to the current video block is a 4-parameter model, and the modification rule provides for modifying two (2) CPMVs associated with the current video block.

[0985] 9. The method of clause 7, wherein the affine model applied to the current video block is a 6-parameter model, and the modification rule provides for modifying three (3) CPMVs associated with the current video block.

[0986] 10. The method of clause 7, wherein the modification rule provides for modification of all one or more control point motion vectors (CPMVs) associated with the current video block, and wherein the syntax element is included in the bitstream representation corresponding to the modification rule.

[0987] 11. The method of any one or more of clauses 7-10, wherein the modification rule provides for modifying two (2) CPMVs associated with the current video block and / or three (3) CPMVs associated with the current video block.

[0988] 12. The method of any one or more of clauses 7-11, wherein the syntax element is selectively included in the bitstream representation based on whether the current video block is processed using affine mode and / or affine AMVP mode and / or affine merge mode.

[0989] 13. The method of clause 12, wherein the syntax elements are represented in the bitstream representation using context coding or bypass coding.

[0990] 14. The method of any one or more of clauses 7-13, wherein the one or more control point motion vectors include a first control point motion vector and a second control point motion vector, and wherein the modification rule provides for modifying the first control point motion vector differently from the second control point motion vector.

[0991] 15. The method of clause 14, wherein a first syntax element is included in the bitstream representation to indicate modification of the first control point motion vector, and a second syntax element is included in the bitstream representation to indicate modification of the second control point motion vector.

[0992] 16. The method of clause 15, wherein the first syntax element and the second syntax element are independent of each other.

[0993] 17. A method according to any one or more of clauses 7-16, wherein the modification rule provides for selectively modifying or excluding modification of the one or more control point motion vectors of the affine merge candidate included in the affine merge candidate, and wherein a syntax element corresponding to the modification rule is included in the bitstream representation.

[0994] 18. The method of clause 17, wherein the first value of the syntax element indicates that none of the one or more control point motion vectors of the affine merge candidate are to be modified.

[0995] 19. The method of clause 17, wherein the second value of the syntax element indicates that the first control point motion vector is to be modified and the second control point motion vector is to remain unchanged.

[0996] 20. The method of clause 17, wherein the third value of the syntax element indicates that the first control point motion vector is to remain unchanged and the second control point motion vector is to be modified.

[0997] 21. The method of clause 17, wherein the fourth value of the syntax element indicates that both the first control point motion vector and the second control point motion vector are to be modified.

[0998] 22. The method of clause 17, wherein the syntax element indicates whether at least one of the one or more control point motion vectors is to be modified.

[0999] 23. The method of clause 22, wherein the syntax element is a first syntax element indicating whether at least one of the one or more control point motion vectors is to be modified, and wherein, in the case where the first syntax element indicates that at least one control point motion vector is to be modified, the second syntax element indicates whether a particular control point motion vector is to be modified, and wherein the second syntax element is included in the bitstream representation.

[1000] 24. A method according to any one or more of clauses 7-22, wherein, in order, a first syntax element corresponding to whether at least one of the one or more control point motion vectors is modified and a second syntax element corresponding to a merge index of an affine merge candidate are included in the bitstream representation.

[1001] 25. The method of clause 24, wherein the first syntax element is included in the bitstream representation before the second syntax element is included.

[1002] 26. The method of clause 24, wherein the second syntax element is included in the bitstream representation before the first syntax element is included.

[1003] 27. The method of clause 23, wherein the first syntax element and / or the second syntax element are represented in the bitstream representation using context coding.

[1004] 28. The method of clause 23, wherein the first syntax element and / or the second syntax element are represented in the bitstream representation using a bypass codec.

[1005] 29. The method of clause 7, wherein modifying one or more control point motion vectors (CPMVs) of an affine merge candidate of the current video block comprises adding an offset value to the one or more control point motion vectors, and wherein the number of the one or more control point motion vectors used depends on the affine model.

[1006] 30. The method of clause 29, wherein the offset value is derived using a modification index and / or a direction index and / or a distance index associated with one or more control point motion vectors.

[1007] 31. The method of clause 29, wherein the modification index and / or the direction index and / or the distance index are included in the bitstream representation.

[1008] 32. The method of clause 29, wherein the affine model is a 4-parameter model defined using control point vectors MV0 and MV1, and wherein the offset values of MV0 and MV1 are represented as Off0 = (Off0x, Off0y) and Off1 = (Off1x, Off1y), respectively.

[1009] 33. The method of claim 29, wherein the affine model is a 6-parameter model defined using control point vectors MV0, MV1, and MV2, and wherein the offset values of MV0, MV1, and MV2 are represented as Off0 = (Off0x, Off0y), Off1 = (Off1x, Off1y), and Off2 = (Off2x, Off2y), respectively.

[1010] 34. The method of clause 30, wherein the offset value is a final offset value, further comprising:

[1011] calculating an initial offset value of the first control point motion vector, wherein the initial offset value of the first control point motion vector depends only on a first modification index and / or a first direction index and / or a first distance index associated with the first control point motion vector, and is independent of modification indices and / or direction indices and / or distance indices of other control point vectors;

[1012] Calculating an initial offset value of the second control point motion vector, wherein the initial offset value of the second control point motion vector depends only on the second modification index and / or the second direction index and / or the second distance index associated with the second control point motion vector, and is independent of the modification index and / or the direction index and / or the distance index of other control point vectors; and

[1013] The final offset value of the first control point motion vector is calculated using the initial offset value of the first control point motion vector and the initial offset value of the second control point motion vector.

[1014] 35. The method of clause 34, wherein the affine model is a 4-parameter model defined using control point vectors MV0 and MV1, wherein the initial offset value and the final offset value of MV0 are expressed as Off0' = (Off0x', Off0y') and Off0 = (Off0x, Off0y), and wherein the initial offset value and the final offset value of MV1 are expressed as Off1' = (Off1x', Off1y') and Off1 = (Off1x, Off1y), and wherein Off0 is derived using Off0' and Off1', and wherein Off1 is derived using Off0' and Off1'.

[1015] 36. The method of clause 34, wherein the affine model is a 6-parameter model defined using control point vectors MV0, MV1, and MV2, wherein the initial offset value and the final offset value of MV0 are expressed as Off0' = (Off0x', Off0y') and Off0 = (Off0x, Off0y), wherein the initial offset value and the final offset value of MV1 are expressed as Off1' = (Off1x', Off1y') and Off1 = (Off1x, Off1y), wherein the initial offset value and the final offset value of MV2 are expressed as Off2' = (Off2x', Off2y') and Off2 = (Off2x', Off2y'), and wherein Off0 is derived using Off0', Off1', Off2', wherein Off1 is derived using Off0', Off1', Off2', and wherein Off2 is derived using Off0', Off1', Off2'.

[1016] 37. The method of clause 34, wherein the final offset value of the control point motion vector is the same as the initial offset value of the control point motion vector.

[1017] 38. The method of clause 34, wherein the final offset value of the second control point motion vector is the sum of the initial offset value of the first control point motion vector and the initial offset value of the second control point motion vector.

[1018] 39. The method of clause 34, wherein the final offset value of the third control point motion vector is the sum of the initial offset value of the first control point motion vector and the initial offset value of the third control point motion vector.

[1019] 40. The method of any one or more of clauses 7-39, wherein, when modifying one or more control point motion vectors (CPMVs), the resulting control point motion vectors are stored for use in processing other video blocks of the visual media data.

[1020] 41. The method of clause 40, wherein the other blocks of visual media data and the current video block are located in the same picture.

[1021] 42. The method of clause 40, wherein the other blocks of visual media data and the current video block are located in different pictures.

[1022] 43. The method of any one or more of clauses 7-39, wherein, when modifying one or more control point motion vectors (CPMVs), the resulting control point motion vectors are not stored, and wherein the unmodified decoded control point motion vectors are stored for use in processing other video blocks of the visual media data.

[1023] 44. The method of clause 22, wherein a syntax element indicating whether at least one of the one or more control point motion vectors is to be modified for a current video block is used in processing subsequent video blocks.

[1024] 45. The method of clause 44, wherein the syntax element corresponds to MMVD_AFFINE_flag.

[1025] 46. The method of any one or more of clauses 7-39, further comprising:

[1026] After modification, the affine merge candidate is inserted into a list, where the index is used to indicate whether a member included in the list has been modified and / or to identify the member included in the list.

[1027] 47. The method of clause 46, wherein the list is an affine merge list and the index is an affine merge index.

[1028] 48. The method of clause 46, wherein the list is a sub-block merge list and the index is a sub-block merge index.

[1029] 49. The method of clause 46, further comprising:

[1030] Generates a new affine merge candidate based on the affine merge candidates included in the list.

[1031] 50. The method of clause 49, wherein one or more new affine merge candidates are added to the list.

[1032] 51. The method of any one or more of clauses 49-50, wherein one or more control point motion vectors of the affine merge candidates in the list are used to generate a new affine merge candidate.

[1033] 52. The method of clause 51, wherein an average and / or weighted average of one or more control point motion vectors of the affine merge candidates in the list is used to generate a new affine merge candidate.

[1034] 53. The method of clause 7, wherein the one or more control point motion vectors have an associated direction index and / or distance index, wherein the direction index and / or distance index are included in the bitstream representation.

[1035] 54. The method of clause 53, wherein the value of the distance index included in the bitstream representation is zero.

[1036] 55. The method of clause 54, wherein if the value of the distance index is zero, the direction index is excluded from the bitstream representation.

[1037] 56. The method of clause 17, wherein the syntax element indicates that none of the one or more control point vectors of the affine merge candidate are to be modified.

[1038] 57. The method of clause 56, wherein, in the event that one or more control point vectors are not to be modified, modification information for the affine merge candidate is excluded from the bitstream representation.

[1039] 58. The method of clause 17, wherein, when a syntax element indicates that at least one of the one or more control point vectors of an affine merge candidate is to be modified, and a further field indicates that none of the one or more control point vectors before the last one of the one or more control point motion vectors is to be modified, it is determined that the last one of the one or more control point vectors is to be modified.

[1040] 59. The method of clause 58, wherein the distance index of a last one of the one or more control point vectors is assigned a value greater than zero.

[1041] 60. The method of any one or more of clauses 44-45, wherein the MMVD_AFFINE_flag of the current block is encoded or decoded using a context derived from at least one MMVD_AFFINE_flag of a previously decoded block.

[1042] 61. The method of clause 58, wherein the syntax element indicating whether the distance index of a last one of the one or more control point vectors is equal to 0 is excluded from the bitstream representation.

[1043] 62. A method of visual media processing (e.g., Figure 38 3800 as discussed in conjunction with the example embodiments in Section 4 of this document), comprising:

[1044] For a conversion between a current video block of visual media data and a bitstream representation of the visual media data, determining (step 3802) a subset of allowed affine merge candidates for the current video block that are subject to modification based on a modification rule; and

[1045] Transforming is performed (step 3804) based on the determination; wherein a syntax element included in the bitstream representation indicates a modification rule.

[1046] 63. The method of clause 62, wherein the subset of allowed affine merge candidates corresponds to the first K affine merge candidates included in the list.

[1047] 64. The method of clause 63, wherein K is greater than zero and less than the total number of allowed affine merge candidates.

[1048] 65. The method of clause 64, wherein the syntax element is a first syntax element indicating whether the affine merge candidate is to be modified, and wherein a second syntax element corresponding to the merge index of the affine merge candidate is included in the bitstream representation.

[1049] 66. The method of clause 65, wherein the first syntax element is included in the bitstream representation before including the second syntax element, and wherein the maximum value of the second syntax element is K-1.

[1050] 67. The method of clause 65, wherein the first syntax element is selectively included in the bitstream representation after including the second syntax element only if the value of the second syntax element is not greater than K-1.

[1051] 68. The method of clause 62, wherein the subset of allowed affine merge candidates corresponds to the last K affine merge candidates included in the list.

[1052] 69. The method of clause 65, wherein the first syntax element is selectively included in the bitstream representation after including the second syntax element only if the affine merge candidate belongs to one of the last K affine merge candidates included in the list.

[1053] 70. The method of any of clauses 1-69, wherein converting comprises encoding the video into a bitstream representation.

[1054] 71. The method of any of clauses 1-69, wherein converting comprises parsing and decoding the bitstream representation to generate the video.

[1055] 72. A video processing device comprising a processor configured to implement one or more of clauses 1 to 69.

[1056] 73. A computer-readable medium having stored thereon code which, when executed by a processor, causes the processor to perform the method of any one or more of clauses 1 to 69.

[1057] 74. A computer-readable medium having stored thereon a bitstream representation of a video, the bitstream representation being generated by the method of any one or more of clauses 1 to 69.

[1058] In this document, the terms "video processing" or "visual media processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits that are co-located or dispersed across different locations within the bitstream as defined by the syntax. For example, a macroblock may be encoded based on error residual values from a transform and a codec, and bits from headers and other fields in the bitstream may also be used. Furthermore, as described in the above solutions, during conversion, a decoder may parse the bitstream knowing that some fields may or may not be present based on this determination. Similarly, an encoder may determine whether to include certain syntax fields and generate a codec representation accordingly by including the syntax fields or excluding the syntax fields from the codec representation. It will be understood that the disclosed techniques may be embodied in a video encoder or decoder to improve compression efficiency using techniques including the use of sub-block based motion vector refinement.

[1059] It will be appreciated that the disclosed techniques may be embodied in a video encoder or decoder to improve compression efficiency using techniques including the use of sub-block based motion vector refinement.

[1060] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuitry or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or any combination thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing device or to control its operation. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or any combination thereof. The term "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or any combination thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[1061] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a file portion that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the relevant program, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program may be deployed for execution on a single computer or on multiple computers, with the computers being located at a single site or distributed across multiple sites and interconnected by a communications network.

[1062] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[1063] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from read-only memory or random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data (e.g., magnetic, magneto-optical, or optical disks). However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[1064] Although this patent document contains many details, these details should not be interpreted as limitations on the scope of any subject matter or content that may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular technologies. Certain features described in this patent document in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, although the above-mentioned features may be described as working in a particular combination, or even initially claimed as such, in some cases, one or more features from the claimed combination may be deleted from the claimed combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[1065] Similarly, while operations may be described in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all illustrated operations be performed, in order to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[1066] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A method for processing visual media, comprising: For a conversion between a current video block of visual media data and a bitstream of the visual media data, determining an affine candidate for the current video block to be modified based on a modification rule; as well as performing said converting based on said determining; wherein the modification rule provides for modifying one or more control point motion vectors associated with the current video block; wherein modifying the one or more control point motion vectors of the affine candidate for the current video block comprises: adding one or more offset values to the one or more control point motion vectors, each of the one or more offset values being derived by using a direction index and a distance index associated with the control point motion vector; When the affine model of the current video block is a 6-parameter model defined using three control point motion vectors MV0, MV1 and MV2, the one or more control point motion vectors to be modified are MV0 and MV1.

2. The method according to claim 1, wherein The direction index and the distance index are included in the bitstream.

3. The method according to claim 1, wherein The N control point motion vectors are modified by using the one or more offset values.

4. The method according to claim 3, wherein: N=2。 5. The method according to claim 3, wherein: When the affine model of the current video block is a 4-parameter model defined using two control point motion vectors MV0 and MV1, the one or more control point motion vectors to be modified are MV0 and MV1.

6. The method according to claim 1, wherein A syntax element in the bitstream is used to indicate whether to modify the one or more control point motion vectors.

7. The method according to claim 3, wherein: N is predefined.

8. The method according to claim 6, wherein: The syntax element is context coding.

9. The method according to claim 1, wherein: The modified control point motion vector is stored for processing other video blocks of the visual media data.

10. The method according to any one of claims 1 to 9, wherein The converting includes encoding the current video block into the bitstream.

11. The method according to any one of claims 1 to 9, wherein The converting includes decoding the bitstream to generate the current video block.

12. A video data processing apparatus comprising a processor and a non-transitory memory having instructions thereon, wherein: The instructions, when executed by the processor, cause the processor to: For a conversion between a current video block of visual media data and a bitstream of the visual media data, determining an affine candidate for the current video block to be modified based on a modification rule; as well as performing said converting based on said determining; wherein the modification rule provides for modifying one or more control point motion vectors associated with the current video block; wherein modifying the one or more control point motion vectors of the affine candidate for the current video block comprises: adding one or more offset values to the one or more control point motion vectors, each of the one or more offset values being derived by using a direction index and a distance index associated with the control point motion vector; When the affine model of the current video block is a 6-parameter model defined using three control point motion vectors MV0, MV1 and MV2, the one or more control point motion vectors to be modified are MV0 and MV1.

13. The device according to claim 12, wherein The direction index and the distance index are included in the bitstream.

14. The device according to claim 12, wherein The N control point motion vectors are modified by using the one or more offset values.

15. The device according to claim 14, wherein N=2。 16. The device according to claim 14, wherein When the affine model of the current video block is a 4-parameter model defined using two control point motion vectors MV0 and MV1, the one or more control point motion vectors to be modified are MV0 and MV1.

17. A non-transitory computer-readable storage medium having stored therein instructions that cause a processor to: For a conversion between a current video block of visual media data and a bitstream of the visual media data, determining that an affine candidate for the current video block is modified based on a modification rule; and performing said converting based on said determining; in, The modification rules provide for modifying one or more control point motion vectors associated with the current video block; wherein modifying the one or more control point motion vectors of the affine candidate for the current video block comprises: adding one or more offset values to the one or more control point motion vectors, each of the one or more offset values being derived by using a direction index and a distance index associated with the control point motion vector; When the affine model of the current video block is a 6-parameter model defined using three control point motion vectors MV0, MV1 and MV2, the one or more control point motion vectors to be modified are MV0 and MV1.

18. The storage medium according to claim 17, wherein: The N control point motion vectors are modified by using the one or more offset values, and N=2.

19. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method executed by a video processing apparatus, wherein: The method comprises: For a current video block of visual media data, determining that an affine candidate of the current video block is modified based on a modification rule; and generating the bitstream based on the determination; wherein the modification rule provides for modifying one or more control point motion vectors associated with the current video block; wherein modifying the one or more control point motion vectors of the affine candidate for the current video block comprises: adding one or more offset values to the one or more control point motion vectors, each of the one or more offset values being derived by using a direction index and a distance index associated with the control point motion vector; When the affine model of the current video block is a 6-parameter model defined using three control point motion vectors MV0, MV1 and MV2, the one or more control point motion vectors to be modified are MV0 and MV1.

20. A method for storing a bitstream of a video, comprising: For a current video block of visual media data, determining that an affine candidate of the current video block is modified based on a modification rule; generating the bitstream based on the determination; as well as storing the bitstream in a non-transitory computer-readable recording medium; wherein the modification rule provides for modifying one or more control point motion vectors associated with the current video block; wherein modifying the one or more control point motion vectors of the affine candidate for the current video block comprises: adding one or more offset values to the one or more control point motion vectors, each of the one or more offset values being derived by using a direction index and a distance index associated with the control point motion vector; When the affine model of the current video block is a 6-parameter model defined using three control point motion vectors MV0, MV1 and MV2, the one or more control point motion vectors to be modified are MV0 and MV1.

Citation Information

Patent Citations

  • Affine motion prediction for video coding

    CN109155855A

  • Motion vector generation for affine motion model for video coding

    IN201947016713A