USING CO-LOCATED BLOCKS IN SUB-BLOCK TEMPORAL MOTION VECTOR PREDICTION MODE - Patent application

Sub-block-based temporal motion vector prediction tools optimize motion vector prediction in video coding, addressing inefficiencies and complexity, thereby enhancing video coding efficiency and quality.

JP7741139B2Active Publication Date: 2025-09-17DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023117792
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-22
Filing Date
2023-07-19
Publication Date
2025-09-17
Estimated Expiration
2039-11-22

Smart Images

  • Figure 0007741139000027
    Figure 0007741139000027
  • Figure 0007741139000028
    Figure 0007741139000028
  • Figure 0007741139000029
    Figure 0007741139000029
Patent Text Reader

Abstract

To provide a device, a system, and a method for digital video encoding and decoding including an inter-prediction method based on sub-blocks.SOLUTION: A method for video processing includes: determining whether to add the maximum number of candidates in a merge candidate list based on subblocks and / or a temporal motion vector prediction candidate (SbTMVP) based on the subblocks to the merge candidate list based on the subblocks on the basis of whether temporal motion vector prediction (TMVP) is valid for use during conversion or whether to use the current picture referring (CPR) encoding mode for the conversion between the current block of a video and the bitstream representation of the video; and performing the conversion on the basis of the determination.SELECTED DRAWING: Figure 30
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) Book Wish This is a divisional application based on Japanese Patent Application No. 2021-526695, which is based on International Patent Application No. PCT / CN2019 / 120301 filed on November 22, 2019, and this International Patent Application is Priority and benefit of International Patent Application No. PCT / CN2018 / 116889 filed on November 22, 2018, International Patent Application No. PCT / CN2018 / 125420 filed on December 29, 2018, International Patent Application No. PCT / CN2019 / 100396 filed on August 13, 2019, and International Patent Application No. PCT / CN2019 / 107159 filed on September 22, 2019 Mainly stretch All the above patents The entire disclosure of the application is incorporated by reference as part of the disclosure of this specification.

[0002] This specification relates to image and video encoding and decoding. [Background technology]

[0003] Despite advances in video compression, digital video still accounts for the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demands for digital video use are expected to continue to grow. Summary of the Invention

[0004] Devices, systems, and methods related to digital video coding are described, including sub-block-based inter prediction methods. The described methods may be applied to both existing video coding standards (e.g., High Efficiency Video Coding (HEVC) and / or Universal Video Coding (VVC)) and future video coding standards or video codecs.

[0005] In one exemplary aspect, the disclosed technology may be used to provide a method of video processing that includes determining whether to add a maximum number of candidates (ML) in a sub-block-based merge candidate list and / or sub-block-based temporal motion vector prediction (SbTMVP) candidates to the sub-block-based merge candidate list for conversion between a current block of video and a bitstream representation of the video based on whether temporal motion vector prediction (TMVP) is enabled during conversion or whether a current picture reference (CPR) coding mode is used for conversion, and performing the conversion based on the determination.

[0006] In another representative aspect, the disclosed technology may be used to provide a method of video processing that includes determining a maximum number (ML) of candidates in a sub-block-based merge candidate list based on whether temporal motion vector prediction (TMVP), sub-block-based temporal motion vector prediction (SbTMVP) tools, and affine coding modes are enabled for use during the conversion for converting between a current block of video and a bitstream representation of the video, and performing the conversion based on the determination.

[0007] In another representative aspect, the disclosed techniques may be used to provide a method of video processing that includes determining that sub-block-based motion vector prediction (SbTMVP) mode is disabled for conversion between a current block of a first video segment of a video and a bitstream representation of the video because temporal motion vector prediction (TMVP) mode is disabled for conversion at the first video segment level, and performing the conversion based on the determination, where the bitstream representation conforms to a format that specifies whether to include an indication of SbTMVP mode and / or a position of the indication of SbTMVP mode relative to the indication of TMVP mode in a merge candidate list.

[0008] In another exemplary aspect, the disclosed technology may be used to provide a method of video processing that includes converting between a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool and a bitstream representation of the video, selectively masking coordinates of the current block or a corresponding position of the current block using a mask based on compression of motion vectors associated with the SbTMVP tool or the TMVP tool, and applying the mask includes bitwise ANDing values ​​of the coordinates and values ​​of the mask.

[0009] In another exemplary aspect, the disclosed techniques may be used to provide a method of video processing that includes determining, based on one or more characteristics of a current block of a video segment of a video, a valid corresponding region of the current block for applying a sub-block-based motion vector prediction (SbTMVP) tool based on the current block, and converting between the current block and a bitstream representation of the video based on the determination.

[0010] In another exemplary aspect, the disclosed technology may be used to provide a method of video processing that includes determining a default motion vector for a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool and converting between the current block and a bitstream representation of the video based on the determination, where the default motion vector is determined when no motion vector is available from a block that includes a corresponding position in the co-located picture associated with a central position of the current block.

[0011] In another representative aspect, the disclosed techniques may be used to provide a method of video processing that includes inferring, for a current block of a video segment of a video, that a current picture of the current block is a reference picture with index M in a reference picture list X, where M and X are integers, and that a sub-block-based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool is disabled for the video segment if X=0 or X=1, and converting between the current block and a bitstream representation of the video based on the inference.

[0012] In another representative aspect, the disclosed techniques may be used to provide a method of video processing that includes determining, for a current block of video, that application of a sub-block-based temporal motion vector prediction (SbTMVP) tool is enabled if a current picture of the current block is a reference picture with index M in a reference picture list X, where M and X are integers, and converting between the current block and a bitstream representation of the video based on the determination.

[0013] In another representative aspect, the disclosed techniques may be used to provide a method of video processing, the method including converting between a current block of video and a bitstream representation of the video, the current block being coded using a sub-block-based coding tool, the converting including coding a sub-block merge index in a uniform manner using multiple bins (N) when a sub-block-based temporal motion vector prediction (SbTMVP) tool is enabled or disabled.

[0014] In another representative aspect, the disclosed techniques may be used to provide a method of video processing that includes determining, for a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool, a motion vector that the SbTMVP tool uses to locate a corresponding block in a picture different from the current picture that contains the current block, and converting between the current block and a bitstream representation of the video based on the determination.

[0015] In another representative aspect, the disclosed techniques may be used to provide a method of video processing that includes determining whether to insert a zero motion affine merge candidate into a sub-block merge candidate list for transforming between a current block of video and a bitstream representation of the video based on whether affine prediction is enabled for the transform of the current block, and performing the transform based on the determination.

[0016] In another representative aspect, the disclosed techniques may be used to provide a method of video processing, the method including inserting a zero motion non-affine padding candidate into a sub-block merging candidate list if the sub-block merging candidate list is not filled, for converting between a current block of video and a bitstream representation of the video using a sub-block merging candidate list, and performing the conversion after the insertion.

[0017] In another exemplary aspect, the disclosed techniques may be used to provide a method of video processing that includes determining a motion vector for conversion between a current block of video and a bitstream representation of the video using a rule that determines that the motion vector is derived from one or more motion vectors of blocks containing corresponding positions in a co-located picture, and performing the conversion based on the motion vector.

[0018] In yet another exemplary aspect, a video encoder apparatus is disclosed, the video encoder apparatus including a processing unit configured to implement the methods described herein.

[0019] In yet another exemplary aspect, a video decoder apparatus is disclosed, the video decoder apparatus including a processing unit configured to implement the methods described herein.

[0020] In yet another aspect, a computer readable medium is disclosed having code stored thereon that, when executed by a processing device, causes the processing device to implement a method described herein.

[0021] These and other aspects are described herein. [Brief explanation of the drawings]

[0022] [Figure 1] FIG. 10 is a diagram illustrating an example of a process for deriving a merge candidate list. [Figure 2] A diagram showing example locations of spatial merge candidates. [Figure 3] A diagram showing example candidate pairs considered in a spatial merge candidate redundancy check. [Figure 4A] FIG. 1 illustrates an exemplary location for a second prediction unit (PU) of an N×2N partition. [Figure 4B] FIG. 1 illustrates an exemplary location for the second prediction unit (PU) of a 2N×N partition. [Figure 5] FIG. 1 illustrates an example of motion vector scaling for temporal merge candidates. [Figure 6] Figure showing example candidate locations C0 and C1 for temporal merge candidates. [Figure 7] FIG. 10 shows an example of combined bi-predictive merge candidates. [Figure 8] FIG. 10 is a diagram showing an example of a process for deriving motion vector prediction candidates; [Figure 9] FIG. 1 illustrates an example of motion vector scaling for spatial motion vector candidates. [Figure 10]FIG. 1 illustrates an example of alternative temporal motion vector prediction (ATMVP) motion estimation for a CU. [Figure 11] FIG. 1 shows an example of one CU with four sub-blocks (AD) and their neighboring blocks. [Figure 12] Flowchart of an example of encoding with different MV precision [Figure 13A] Diagram showing 135-degree split type (split from upper left corner to lower right corner) [Figure 13B] Diagram showing 45-degree division pattern [Figure 14] A diagram showing examples of neighboring block locations [Figure 15] Diagram showing examples of upper and left and right blocks [Figure 16A] A diagram showing two control point motion vectors (CPMVs) [Figure 16B] Diagram showing three examples of CPMV [Figure 17] A diagram showing an example of an affine motion vector field (MVF) for each subblock. [Figure 18A] An example of a four-parameter affine model [Figure 18B] An example of a six-parameter affine model [Figure 19] A diagram showing an example of the MVP of AF_INTER for inherited affine candidates. [Figure 20] A diagram showing an example of constructing an affine motion predictor with AF_INTER. [Figure 21A] FIG. 10 is a diagram showing an example of control point motion vectors in affine coding in AF_MERGE. [Figure 21B] FIG. 10 is a diagram showing an example of control point motion vectors in affine coding in AF_MERGE. [Figure 22] A diagram showing examples of candidate positions for affine merge modes. [Figure 23] FIG. 1 shows an example of an intra-picture block copy operation. [Figure 24] FIG. 10 shows examples of valid corresponding regions in co-located pictures. [Figure 25]FIG. 1 shows an exemplary flowchart for history-based motion vector prediction. [Figure 26] A diagram showing the modified merge list construction process. [Figure 27] FIG. 1 illustrates an exemplary embodiment of a proposed valid region when the current block is in a basic region. [Figure 28] FIG. 1 illustrates an exemplary embodiment of a valid region when the current block is not within a fundamental region. [Figure 29A] FIG. 10 shows an example of a location for identifying existing default motion information. [Figure 29B] FIG. 1 shows an example of a proposed location for identification of default motion information. [Figure 30] 1 is a flowchart illustrating an example of a video processing method. [Figure 31] 1 is a flowchart illustrating an example of a video processing method. [Figure 32] 1 is a flowchart illustrating an example of a video processing method. [Figure 33] 1 is a flowchart illustrating an example of a video processing method. [Figure 34] 1 is a flowchart illustrating an example of a video processing method. [Figure 35] 1 is a flowchart illustrating an example of a video processing method. [Figure 36] 1 is a flowchart illustrating an example of a video processing method. [Figure 37] 1 is a flowchart illustrating an example of a video processing method. [Figure 38] 1 is a flowchart illustrating an example of a video processing method. [Figure 39] 1 is a flowchart illustrating an example of a video processing method. [Figure 40] 1 is a flowchart illustrating an example of a video processing method. [Figure 41] 1 is a flowchart illustrating an example of a video processing method. [Figure 42] 1 is a flowchart illustrating an example of a video processing method. [Figure 43] FIG. 1 is a block diagram illustrating an example hardware platform for implementing the visual media decoding or visual media encoding techniques described herein. [Figure 44] 1 is a block diagram of an exemplary video processing system in which the disclosed techniques may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0023] This specification provides various techniques that can be used by a decoder of a video bitstream to improve the quality of the decompressed or decoded digital video or images. Furthermore, a video encoder may implement these techniques during the encoding process to reconstruct the decoded frames for use in further encoding.

[0024] Section headings are used herein for ease of understanding, but are not intended to limit the embodiments and techniques to the corresponding section, and thus embodiments from one section may be combined with examples from other sections.

[0025] 1. Overview

[0026] This patent specification relates to video coding technology. Specifically, the present invention relates to motion vector coding in video coding. The present invention can be applied to existing video coding standards such as HEVC or standards to be finalized (e.g., general-purpose video coding). The present invention can also be applied to future video coding standards or video codecs.

[0027] 2. Preface

[0028] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and H.265 / HEVC. The video coding standard, H.262, is based on a hybrid video coding architecture that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by the JVET and incorporated into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) was formed between VCEG (Q6 / 16) and ISO / IEC JTC_1 SC29 / WG11 (MPEG) to work on a VVC standard with the goal of reducing the bitrate by 50% compared to HEVC.

[0029] The latest version of the VVC draft, namely Generalized Video Coding (Draft 3), can be found at:

[0030] http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 12_Macao / wg11 / JVET-L1001-v2.zip The latest reference software for VVC, called VTM, can be found here:

[0031] https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-3.0rC1

[0032] 2.1 Inter Prediction in HEVC / H.265

[0033] Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists may be signaled using inter_pred_idc. The motion vector may be explicitly coded as a delta to the predictor.

[0034] When a CU is coded in skip mode, one PU is associated with this CU, has no significant residual coefficients, and has no coded motion vector delta or reference picture index. A merge mode is specified, whereby motion parameters for the current PU are obtained from neighboring PUs, including spatial and temporal candidates. Merge mode can be applied to any inter-predicted PU, not just skip mode. An alternative to merge mode is explicit transmission of motion parameters, explicitly signaling, for each PU, the motion vector (more precisely, the motion vector difference compared to the motion vector predictor (MVD)), which is the reference picture index corresponding to each reference picture list and the use of the reference picture list. This mode is referred to as advanced motion vector prediction (AMVP) in this disclosure.

[0035] If the signaling indicates using one of two reference picture lists, the PU is generated from a block of one sample. This is called "uni-prediction." Uni-prediction is available for both P slices and B slices.

[0036] If the signaling indicates that both reference picture lists are to be used, the PU is generated from a block of two samples. This is called "bi-prediction." Bi-prediction is only available for B slices.

[0037] The inter prediction modes defined in HEVC will be described in detail below, with merge mode being first described.

[0038] 2.1.1 Reference Picture List

[0039] In HEVC, the term inter-prediction is used to indicate a prediction derived from data elements (e.g., sample values ​​or motion vectors) of reference pictures other than the currently decoded picture. Similar to H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter-prediction are organized into one or more reference picture lists. A reference index identifies which reference picture in the list to use to generate the prediction signal.

[0040] One reference picture list, List0, is used for P slices, and two reference picture lists, List0 and List1, are used for B slices. Note that the reference pictures included in List0 / 1 may be in capture / display order or may be from past and future pictures.

[0041] 2.1.2 Merge Mode

[0042] 2.1.2.1 Deriving Merge Mode Candidates

[0043] When predicting a PU using merge mode, an index pointing to an entry in the merge candidate list is parsed from the bitstream and used to look up the motion information. The construction of this list is specified in the HEVC standard and can be summarized based on the following sequence of steps:

[0044] Step 1: Derive initial candidates Step 1.1: Spatial candidate derivation Step 1.2: Spatial candidate redundancy check Step 1.3: Temporal candidate derivation Step 2: Insert additional candidates Step 2.1: Creating bi-prediction candidates Step 2.2: Insertion of zero motion candidates

[0045] These steps are also shown schematically in Figure 1. For spatial merge candidate derivation, a maximum of four merge candidates are selected from candidates at five different locations. For temporal merge candidate derivation, a maximum of one merge candidate is selected from two candidates. Since the decoder assumes a fixed number of candidates per PU, if the number of candidates obtained in step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header, additional candidates are generated. Since the number of candidates is fixed, a truncated unary binarization (TU) is used to encode the index of the best merge candidate. If the CU size is equal to 8, all PUs of the current CU share one merge candidate list, which is the same as the merge candidate list for the 2N × 2N prediction units.

[0046] The operations associated with the above steps are now described in detail.

[0047] 2.1.2.2 Spatial candidate derivation

[0048] In deriving spatial merge candidates, up to four merge candidates are selected from the candidates located at the positions shown in FIG. 2. The derivation order is A1, B1, B0, A0, B2. Position B2 is considered only if any of the PUs located at positions A1, B1, B0, and A0 are unavailable (e.g., because they belong to another slice or tile) or are intra-coded. After adding the candidate located at position A1, the remaining candidates are subjected to a redundancy check, which ensures that candidates with identical motion information are removed from the list, improving coding efficiency. To reduce computational complexity, the aforementioned redundancy check does not consider all possible candidate pairs. Instead, only the pairs connected by the arrows in FIG. 3 are considered, and a candidate is added to the list only if the corresponding candidate used in the redundancy check does not have the same motion information. Another source of overlapping motion information is a "second PU" associated with a partition different from 2N×2N. As an example, FIGS. 4A and 4B show second PUs for N×2N and 2N×N cases, respectively. When dividing the current PU into Nx2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate makes the bi-prediction units have the same motion information, and it is redundant to have only one PU in one coding unit. Similarly, when dividing the current PU into 2NxN, the candidate at position B1 is not considered.

[0049] 2.1.2.3 Temporal candidate derivation

[0050] In this step, only one candidate is added to the list. Specifically, in deriving this temporal merge candidate, a scaled motion vector is derived based on the co-located PU belonging to the picture with the smallest POC difference with the current picture in a given reference picture list. The reference picture list used to derive the co-located PU is explicitly signaled in the slice header. As shown by the dotted line in Figure 5, the scaled motion vector of the temporal merge candidate is obtained. It is scaled from the motion vector of the co-located PU using the POC distances tb and td. tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located PU and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. The practical implementation of this scaling process is described in the HEVC specification. For B slices, two motion vectors are taken, one for reference picture list 0 and one for reference picture list 1, and combined to form a bi-predictive merge candidate.

[0051] FIG. 5 is a diagram illustrating the scaling of motion vectors of temporal merge candidates.

[0052] At the same position PU(Y) belonging to the reference frame, a temporal candidate position is selected between candidate C0 and candidate C1, as shown in Figure 6. If the PU at position C0 is unavailable, intra-coded, or outside the current coding tree unit (CTU a.k.a. LCU, largest coding unit) row, then position C1 is used. Otherwise, position C0 is used to derive the temporal merge candidate.

[0053] FIG. 6 shows examples of candidate positions C0 and C1 for temporal merge candidates.

[0054] 2.1.2.4 Inserting Additional Candidates

[0055] In addition to spatiotemporal merge candidates, there are two additional types of merge candidates: combined bi-predictive merge candidates and zero merge candidates. Spatiotemporal merge candidates are utilized to generate combined bi-predictive merge candidates. Combined bi-predictive merge candidates are used only for B slices. A combined bi-predictive candidate is generated by combining the first reference picture list motion parameters of a first candidate with the second reference picture list motion parameters of another candidate. If these two tuples provide different motion hypotheses, these tuples form a new bi-predictive candidate. As an example, FIG. 7 illustrates the use of two candidates with mvL0, refIdxL0 or mvL1, refIdxL1 in the original list (left) to generate a combined bi-predictive merge candidate that is added to the final list (right). Various rules exist for the combinations considered to generate these additional merge candidates.

[0056] The MaxNumMergeCand capacity is hit by inserting zero motion candidates to fill the remaining entries in the merge candidate list. These candidates have a spatial displacement of zero and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list.

[0057] Specifically, the following steps are performed in order until the merge list is full. 1. Set the variable numRef to either the number of reference pictures associated with list 0 for P slices or the minimum number of reference pictures in the two lists for B slices. 2. Add non-repetitive motion-zero candidates. If variable i is between 0 and numRef-1, add the default motion candidate with MV set to (0,0) and reference picture index set to i to list 0 (for P slices) and to both lists (for B slices). 3. Set MV to (0,0), set the reference picture index of list 0 to 0 (for P slices), and add a recursive zero motion candidate with the reference picture index of both lists set to 0 (for B slices).

[0058] Finally, no redundancy checks are performed on these candidates.

[0059] 2.1.3 Advanced Motion Vector Prediction (AMVP)

[0060] AMVP exploits the spatiotemporal correlation of motion vectors with neighboring PUs, which is used for explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by checking the availability of neighboring PU positions on the left and above, removing redundant candidates, and adding zero vectors to keep the length of the candidate list constant. The encoder can then select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to merge index signaling, the index of the best motion vector candidate is coded using a shortened unary term. In this case, the maximum value to be coded is 2 (see Figure 8). The following sections provide details on the process of deriving motion vector prediction candidates.

[0061] 2.1.3.1 Derivation of AMVP Candidates

[0062] FIG. 8 summarizes the process of deriving motion vector prediction candidates.

[0063] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. To derive spatial motion vector candidates, as shown in Figure 2, two motion vector candidates are ultimately derived based on the motion vectors of each PU at five different positions.

[0064] To derive a temporal motion vector candidate, select one motion vector candidate from two candidates derived based on two different co-located positions. After creating a first spatio-temporal candidate list, remove duplicate motion vector candidates in the list. If the number of candidates is greater than two, remove motion vector candidates whose reference picture index in the associated reference picture list is greater than 1 from the list. If the number of spatio-temporal motion vector candidates is less than two, add an additional zero motion vector candidate to the list.

[0065] 2.1.3.2 Spatial Motion Vector Candidates

[0066] In deriving spatial motion vector candidates, up to two candidates are considered among the five candidates derived from PUs located as shown in Figure 2, and their positions are the same as the positions of the motion merge. The derivation order for the left side of the current PU is specified as A0, A1, scaled A0, scaled A1. The derivation order for the top side of the current PU is specified as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates: two cases that do not require spatial scaling and two cases that use spatial scaling. The four different cases can be summarized as follows:

[0067] No spatial scaling - (1) The same reference picture list and the same reference picture index (same POC) - (2) Different reference picture lists but the same reference picture (same POC) Spatial scaling - (3) Same reference picture list but different reference pictures (different POC) - (4) Different reference picture lists and different reference pictures (different POCs)

[0068] First check the non-spatial scaling case, then spatial scaling. Regardless of the reference picture list, if the POC differs between the reference picture of the neighboring PU and the reference picture of the current PU, spatial scaling is considered. If all PUs of the left candidate are not available or are intra-coded, scaling the upper motion vector helps in parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not allowed for the upper motion vector.

[0069] In the spatial scaling process, the motion vectors of neighboring PUs are scaled in the same manner as in temporal scaling, as shown in Figure 9. The main difference is that the reference picture list and index of the current PU are given as input, and the actual scaling process is the same as in temporal scaling.

[0070] 2.1.3.3 Temporal Motion Vector Candidates

[0071] Except for deriving the reference picture index, all the processes for deriving temporal merge candidates are the same as the processes for deriving spatial motion vector candidates (see FIG. 6). The reference picture index is signaled to the decoder.

[0072] 2.2 Sub-CU-based motion vector prediction method in JEM

[0073] In JEM with QTBT, each CU can have at most one motion parameter set for each prediction direction. In the encoder, two sub-CU level motion vector prediction methods are considered by dividing a large CU into sub-CUs and deriving motion information for all sub-CUs of the large CU. The alternative temporal motion vector prediction (ATMVP) method allows each CU to derive multiple sets of motion information from multiple blocks smaller than the current CU in the aligned reference picture. In the spatio-temporal motion vector prediction (STMVP) method, the temporal motion vector predictor and spatial neighbor motion vectors are used to recursively derive the motion vectors of sub-CUs.

[0074] To maintain a more accurate motion field for sub-CU motion estimation, reference frame motion compression is currently disabled.

[0075] FIG. 10 shows an example of ATMVP motion prediction for a CU.

[0076] 2.2.1 Alternative Temporal Motion Vector Prediction

[0077] In alternative temporal motion vector prediction (ATMVP), the motion vector temporal motion vector prediction (TMVP) method is modified by retrieving a set of multiple motion information (including motion vectors and reference indices) from a block smaller than the current CU. In some implementations, a sub-CU is an NxN square block (N is set to 4 by default).

[0078] ATMVP predicts motion vectors for sub-CUs within a CU in two steps. In the first step, a corresponding block in a reference picture is identified by a temporal vector. This reference picture is called a motion source picture. In the second step, the current CU is divided into sub-CUs, and the motion vector and reference index of each sub-CU are obtained from the block corresponding to each sub-CU.

[0079] In the first step, the reference picture and corresponding block are determined based on the motion information of the spatially neighboring blocks of the current CU. To avoid repeated scanning of neighboring blocks, the first merge candidate in the merge candidate list of the current CU is used. The first available motion vector and its associated reference index are set to the temporal vector and the index of the motion source picture. Thus, ATMVP can identify corresponding blocks more accurately than TMVP, and the corresponding block (sometimes called the aligned block) is always located in the bottom right or center position relative to the current CU.

[0080] In the second step, the corresponding block of the sub-CU is identified by the temporal vector in the motion source picture by adding the temporal vector to the coordinates of the current CU. For each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) is used to derive the motion information of the sub-CU. After identifying the motion information of the corresponding NxN block, it is converted into the motion vector and reference index of the current sub-CU, similar to TMVP in HEVC, and motion scaling and other procedures are applied. For example, the decoder checks whether the low-delay condition (i.e., the POC of all reference pictures of the current picture is smaller than the POC of the current picture) is met, and possibly calculates the motion vector MV. x (motion vectors corresponding to reference picture list X) are used to calculate the motion vector MV of each sub-CU. y (X equals 0 or 1, and Y equals 1-X).

[0081] 2.2.2 Spatiotemporal Motion Vector Prediction (STMVP)

[0082] In this method, the motion vectors of sub-CUs are derived recursively along the raster scan order. Figure 11 illustrates this concept. Consider an 8x8 CU that contains four 4x4 sub-CUs, A, B, C, and D. The neighboring 4x4 blocks in the current frame are labeled a, b, c, and d.

[0083] The motion derivation for sub-CU A begins by identifying its two spatial neighbors. The first neighbor is the NxN block above sub-CU A (block c). If block c is unavailable or intra-coded, check the other NxN blocks above sub-CU A (starting from block c, left to right). The second neighbor is the block to the left of sub-CU A (block b). If block b is unavailable or intra-coded, check the other blocks to the left of sub-CU A (starting from block b, top to bottom). The motion information obtained from the neighboring blocks in each list is scaled to the first reference frame of the given list. Next, the temporal motion vector predictor (TMVP) for sub-block A is derived according to the same procedure as the TMVP derivation specified in HEVC. The motion information of the co-located block at position D is taken and scaled accordingly. Finally, after retrieving and scaling the motion information, all available motion vectors (up to three) are averaged separately for each reference list. This averaged motion vector is the motion vector for the current sub-CU.

[0084] 2.2.3 Sub-CU Motion Estimation Mode Signaling

[0085] Sub-CU modes are enabled as additional merge candidates, and no additional syntax elements are required to signal the mode. Two additional merge candidates are added to the merge candidate list for each CU to represent ATMVP and STMVP modes. If the sequence parameter set indicates that ATMVP and STMVP are enabled, up to seven merge candidates are used. The encoding logic for the additional merge candidates is the same as for merge candidates in HM, i.e., for each CU in a P or B slice, two or more RD checks are required for the two additional merge candidates.

[0086] In JEM, all binary values ​​of the merge index are context coded by CABAC, whereas in HEVC, only the first binary value is context coded and the remaining binary values ​​are context bypass coded.

[0087] 2.3 Inter-Prediction Method in VVC

[0088] There are several new coding tools to improve inter prediction such as adaptive motion vector difference resolution (AMVR) to signal MVD, affine prediction mode, triangular prediction mode (TPM), ATMVP, generalized bi-prediction (GBI), and bidirectional optical flow (BIO).

[0089] 2.3.1 Adaptive Motion Vector Difference Resolution

[0090] In HEVC, when use_integer_mv_flag is 0 in the slice header, the motion vector difference (MVD) (the difference between a motion vector and a PU's predicted motion vector) is signaled in units of quarter luma samples. In VVC, locally adaptive motion vector resolution (LAMVR) is introduced. In VVC, MVD can be coded in units of quarter luma samples, integer luma samples, or four luma samples (i.e., quarter pixel, one pixel, four pixels). MVD resolution is controlled at the coding unit (CU) level, and an MVD resolution flag is conditionally signaled for each CU that has at least one non-zero MVD module.

[0091] For a CU with at least one non-zero MVD module, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. If the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or 4 luma sample MV precision is used.

[0092] If the first MVD resolution flag of a CU is zero or not coded for the CU (i.e., all MVDs in the CU are zero), then 1 / 4 luma sample MV resolution is used for the CU. If the CU uses integer luma sample MV precision or 4 luma sample MV precision, round the MVPs in the CU's AMVP candidate list to the corresponding precision.

[0093] In the encoder, the CU-level RD check is used to determine which MVD resolution to use for the CU. That is, the CU-level RD check is performed three times for each MVD resolution. To speed up the encoder, the following coding method is applied in JEM:

[0094] During the RD check of a CU with normal 1 / 4 luma sample MVD resolution, the motion information (integer luma sample precision) of the current CU is stored. During the RD check of the same CU with integer luma sample and 4 luma sample MVD resolution, the stored motion information (after rounding) is used as the starting point for further small-range motion vector refinement, so that the time-consuming motion estimation process is not duplicated three times.

[0095] Conditionally invoke RD check for CUs with 4 luma sample MVD resolution. For a CU, if the RD cost of integer luma sample MVD resolution is much larger than that of 1 / 4 luma sample MVD resolution, then the RD check of 4 luma sample MVD resolution for the CU is skipped.

[0096] The symbolic processing is shown in FIG. 12. First, the 1 / 4-pixel MV is tested, the RD cost is calculated, and represented as RDCost0. Next, the integer MV is tested, and the RD cost is represented as RDCost1. If RDCost1 < th * RDCost0 (where th is a positive value), the 4-pixel MV is tested; otherwise, the 4-pixel MV is skipped. Basically, when checking the integer or 4-pixel MV, the motion information and RD cost, etc. for the 1 / 4-pixel MV are already known, and this can be reused to speed up the encoding process of the integer or 4-pixel MV.

[0097] 2.3.2 Triangular Prediction Mode

[0098] The concept of the Triangular Prediction Mode (TPM) is to introduce a new triangular partition for motion compensation prediction. As shown in FIGS. 13A and 13B, the CU is divided into two triangular prediction units in the diagonal or anti-diagonal direction. Each triangular prediction unit in the CU is inter-predicted using its own single prediction motion vector and reference frame index derived from one single prediction candidate list. After predicting the triangular prediction unit, an adaptive weighting process is performed on the diagonal edge. Then, the transform and quantization processes are performed on the entire CU. Note that this mode is only applicable to the merge mode (note that the skip mode is treated as a special merge mode).

[0099] FIGS. 13A and 13B are explanatory diagrams showing the division of the CU into two triangular prediction units (two division patterns). FIG. 13A: 135-degree division type (division from the upper left corner to the lower right corner), FIG. 13B: 45-degree division pattern.

[0100] 2.3.2.1 Single Prediction Candidate List of TPM

[0101] The single-prediction candidate list, called the TPM motion candidate list, consists of five single-prediction motion vector candidates. It is derived from seven neighboring blocks, including five spatially neighboring blocks (1-5) and two temporally co-located blocks (6-7), as shown in Figure 14. The motion vectors of the seven neighboring blocks are collected and added to the single-prediction candidate list in the following order: single-prediction motion vector, L0 motion vector of bi-prediction motion vector, L1 motion vector of bi-prediction motion vector, and average motion vector of L0 and L1 motion vectors of bi-prediction motion vector. If the number of candidates is less than five, motion vector zero is added to the list. The motion candidates added to this TPM list are called TPM candidates, and the motion information derived from the spatial / temporal blocks is called regular motion candidates.

[0102] Specifically, the following steps are included:

[0103] 1) When adding regular motion candidates from spatially neighboring blocks, the full pruning operation Canonical motion candidates are obtained from A1, B1, B0, A0, B2, Col, and Col2 (corresponding to blocks 1-7 in Figure 14).

[0104] 2) Set the variable numCurrMergeCand=0.

[0105] 3) For each regular motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and numCurrMergeCand is less than 5, then if the regular motion candidate is single-predicted (from either List0 or List1), it is added directly to the merge list as a TPM candidate with numCurrMergeCand increased by 1. We name such a TPM candidate "originally single-predicted candidate." Full pruning applies.

[0106] 4) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and numCurrMergeCand is less than 5, if the regular motion candidate is bi-predictive, add the motion information from List 0 to the TPM merge list as a new TPM candidate (i.e., modified from List 0 to uni-predictive) and add 1 to numCurrMergeCand. We call such a TPM candidate a "shortened List 0 prediction candidate." Full pruning applies.

[0107] 5) For each motion candidate derived from A1, B1, B0, A0, B2, Col, and Col2, if it has not been pruned and if numCurrMergeCand is less than 5, if the regular motion candidate is bi-predictive, add the motion information from List1 to the TPM merge list (i.e., corrected to uni-predictive from List1) and increase numCurrMergeCand by 1. We call such a TPM candidate a "shortened List1-predictive candidate." Full pruning applies.

[0108] 6) For each motion candidate derived from A1, B1, B0, A0, B2, Col, Col2, if not pruned, numCurrMergeCand is less than 5 if the regular motion candidate is bi-predictive. - If the List0 reference picture slice QP is smaller than the List1 reference picture slice QP, first scale the motion information of List1 to the List0 reference picture, and add the average of the two MVs (one from the original List0 and the other from the scaled List1) to the TPM merge list. We call such a candidate the average single prediction from List0 motion candidate, and increase numCurrMergeCand by 1. - Otherwise, first scale the motion information of List0 to the List1 reference picture, and add the average of the two MVs (one from the original List1 and the other from the scaled MV from List0) to the TPM merge list. Such a TPM candidate is called the average uniprediction from List1 motion candidate, and increase numCurrMergeCand by 1. Full pruning applies.

[0109] 7) If numCurrMergeCand is less than 5, add a zero motion vector candidate.

[0110] If, when inserting a candidate into a list, it has to be compared with all the previously added candidates to see if it is the same as one of them, then such a process is called full pruning.

[0111] 2.3.2.2 Adaptive Weighting

[0112] After each triangular prediction unit is predicted, the diagonal edge between two triangular prediction units is subjected to adaptive weighting to derive the final prediction for the entire CU. Two sets of weighting coefficients are defined as follows: The first set of weighting factors uses {7 / 8,6 / 8,4 / 8,2 / 8,1 / 8} and {7 / 8,4 / 8,1 / 8} for luma and chroma samples, respectively. The second set of weighting factors uses {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} for luma and chroma samples respectively.

[0113] A set of weighting factors is selected based on a comparison of the motion vectors of the two triangular prediction units. The second set of weighting factors is used if the reference pictures of the two triangular prediction units are different or if the difference between their motion vectors is greater than 16 pixels. Otherwise, the first set of weighting factors is used.

[0114] 2.3.2.3 Triangle Prediction Mode (TPM) Signaling

[0115] A single bit flag may be signaled first to indicate whether a TPM is being used, followed by further signaling two split patterns (as shown in Figures 13A and 13B) and the selected merge index for each of the two splits.

[0116] 2.3.2.3.1 TPM Flag Signaling

[0117] The width and height of one luminance block are represented by W and H, respectively. If W*H<64, the triangular prediction mode is disabled.

[0118] When coding a block in affine mode, the triangular prediction mode is also disabled.

[0119] When a block is coded in merge mode, a one-bit flag can be signaled to indicate whether triangular prediction mode is enabled or disabled for this block.

[0120] This flag is coded in three contexts according to the following formula: Ctx index=((left block L available && L is coded with TPM?)1:0) +((Above block A available && A is coded with TPM?)1:0);

[0121] FIG. 15 shows an example of neighboring blocks (A and L) used for context selection in TPM flag encoding.

[0122] 2.3.2.3.2 Indication of two split patterns (as shown in Figure 13) and signaling the merge index selected for each of the two splits

[0123] Note that the split pattern and the merge indexes of the two splits are coded relative to each other. In the existing implementation, there is a restriction that two splits cannot use the same reference index. Therefore, there are 2 (split pattern) * N (maximum number of merge candidates) * (N-1) possibilities, where N is set to 5. One indication is coded, and the mapping between the split pattern, the two merge indexes, and the coded instructions is derived from the array defined below.

[0124] const uint8_t g_TriangleCombination[TRIANGLE_MAX_NUM_CANDS][3]={ {0,1,0},{1,0,1},{1,0,2},{0,0,1},{0,2,0},{1,0,3},{1,0,4},{1,1,0},{0,3,0},{0,4,0},{0,0,2},{0,1,2},{1,1,2},{0,0,4},{0,0,3},{0,1,3},{0,1,4},{1,1,4},{1,1,3},{1,2,1}, {1,2,0},{0,2,1},{0,4,3},{1,3,0},{1,3,2},{1,3,4},{1,4,0},{1,3,1},{1,2,3},{1,4,1},{0,4,1},{0,2,3},{1,4,2},{0,3,2},{1,4,3},{0,3,1},{0,2,4},{1,2,4},{0,4,2},{0,3,4}};

[0125] Division pattern (45 degrees or 135 degrees) = g_TriangleCombination[signaled indication][0]; Merge index of candidate A=g_TriangleCombination[signaled indication][1]; Merge index of candidate B=g_TriangleCombination[signaled indication][2];

[0126] After deriving two motion candidates A and B, the (PU1, PU2) motion information of the two partitions can be set from either A or B, and whether PU1 uses the motion information of merge candidate A or B depends on the prediction direction of the two motion candidates. Table 1 shows the relationship between the two derived motion candidates A and B with two partitions.

[0127] [Table 1]

[0128] 2.3.2.3.3 Entropy coding of representation (denoted by merge_triangle_idx)

[0129] merge_triangle_idx is in the range [0,39] inclusive. The K_th order Exponential Golomb (EG) code is used to binarize merge_triangle_idx (K is set to 1).

[0130] K-th orderEG

[0131] This can be generalized using a non-negative integer parameter k to encode larger numbers with fewer bits (at the expense of using more bits to encode smaller numbers). To encode a non-negative integer x with an exp-Golomb code of degree k, we do: 1. Use the order-0 exp-Golomb code above to find [x / 2 k ] is encoded. 2. x mod 2 k is encoded in binary.

[0132] [Table 2]

[0133] 2.3.3 Affine Motion Compensation Prediction

[0134] In HEVC, only a translational motion model is applied for motion compensated prediction (MCP). In the real world, there are various types of motion, such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions. In VVC, a four-parameter affine model and a six-parameter affine model are used to apply simple affine transformation motion compensated prediction. As shown in Figures 16A and 16B, the affine motion field of a block is represented by two control point motion vectors (CPMVs) in the case of the four-parameter affine model (Figure 16A), and by three CPMVs in the case of the six-parameter affine model (Figure 16B).

[0135] The motion vector field (MVF) of a block is expressed by the following equations using the four-parameter affine model in equation (1) (where the four parameters are defined as variables a, b, e, and f) and the six-parameter affine model in equation (2) (where the four parameters are defined as variables a, b, c, d, e, and f), respectively:

[0136]

number

[0137]

number

[0138] where (mv h 0,mv h 0) is the motion vector of the control point in the upper left corner, and (mv h 1,mv h 1) is the motion vector of the control point in the upper right corner, and (mv h 2,MV h 2) is the motion vector of the control point in the bottom left corner. All three motion vectors are called control point motion vectors (CPMVs). (x,y) represents the coordinates of the representative point relative to the top left sample in the current block. (mv h (x,y),mv v(x,y) is the motion vector derived for the sample located at (x,y). The CP motion vector may be signaled (as in Affine AMVP mode) or derived on the fly (as in Affine Merge mode). w and h are the width and height of the current block. In practice, this division is implemented by a right shift with rounding operation. In VTM, the representative point is taken as the center position of a sub-block. For example, if the coordinates of the top-left sample of the top-left corner of a sub-block in the current block are (xs,ys), the coordinates of the representative point are taken as (xs+2,ys+2). For each sub-block (i.e., 4x4 in VTM), the representative point is used to derive the motion vector for the entire sub-block.

[0139] To further simplify motion compensation prediction, sub-block-based affine transformation prediction is applied. To derive a motion vector for each M×N (in current VVC, both M and N are set to 4) sub-block, as shown in Figure 17, the motion vector of the center sample of each sub-block is calculated according to Equation (1) and Equation (2) and rounded to 1 / 16 fractional precision. Then, a 1 / 16-pixel motion compensation interpolation filter is applied, and the derived motion vector is used to generate a prediction for each sub-block. The 1 / 16-pixel interpolation filter is implemented in affine mode.

[0140] After MCP, the high-precision motion vectors for each sub-block are rounded and stored with the same precision as the normal motion vectors.

[0141] 2.3.3.1 Signaling Affine Prediction

[0142] Similar to the translational motion model, there are two modes for signaling side information with affine prediction: AFFINE_INTER mode and AFFINE_MERGE mode.

[0143] AF_INTER mode

[0144] For CUs where both width and height are greater than 8, the AF_INTER mode can be applied. A CU-level affine flag is signaled in the bitstream to indicate whether the AF_INTER mode is used or not.

[0145] In this embodiment, for each reference picture list (list 0 or list 1), three types of affine motion predictors are used to construct an affine AMVP candidate list in the following order, where each candidate contains an estimated CPMV for the current block: The best CPMV difference found on the encoder side (such as mv0mv1mv2 in Figure 20) and the estimated CPMV are signaled; and the index of the affine AMVP candidate from which the estimated CPMV is derived is signaled:

[0146] 1) Inherited affine motion predictor

[0147] The check order is similar to the check order of spatial MVP in HEVC AMVP list construction. First, the left inherited affine motion predictor is derived from the first block in {A1, A0} that has the same reference picture as the affine-coded current block. Next, the inherited affine motion predictor is derived from the block that is derived from the first block in {B1, B0, B2} that is affine-coded and has the same reference picture as the current block. Figure 19 shows five blocks A1, A0, B1, B0, and B2.

[0148] If a neighboring block is found to be coded in affine mode, the CPMV of the coding unit containing this neighboring block is used to derive the predictor of the CPMV of the current block. For example, if A1 is coded in non-affine mode and A0 is coded in 4-parameter affine mode, the inherited affine MV predictor on the left side is derived from A0. In this case, for the CPMV in the top left of Figure 21B, MV0 is used. N , CPMV and MV1 for CPMV in the upper right NUsing the CPMV of the CU containing A0, MV0 is calculated for the top left (coordinate (x0, y0)), top right (coordinate (x1, y1)) and bottom right (coordinate (x2, y2)) positions of the current block. C ,MV1 C ,MV2 C Derive the estimated CPMV of the current block, denoted as:

[0149] 2) Constructed affine motion predictor

[0150] The constructed affine motion predictor consists of control point motion vectors (CPMVs) derived from neighboring inter-coded blocks with the same reference picture, as shown in Figure 20. If the current affine motion model is 4-parameter affine, the number of CPMVs is 2; otherwise, if the current affine motion model is 6-parameter affine, the number of CPMVs is 3. The CPMV {mv0} in the upper left is derived by the MV of the first block in group {A, B, C} that is inter-coded and has the same reference picture as the current block. The CPMV {mv1} in the upper right is derived by the MV of the first block in group {D, E} that is inter-coded and has the same reference picture as the current block. The CPMV {mv2} in the lower left is derived by the MV of the first block in group {F, G} that is inter-coded and has the same reference picture as the current block.

[0151] - If the current affine motion model is a four-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if both m̂v0̂ and m̂v1̂ are established, i.e., m̂v0̂ and m̂v1̂ are used as the estimated CPMVs for the top-left (coordinates (x0, y0)) and top-right (coordinates (x1, y1)) positions of the current block.

[0152] - If the current affine motion model is a six-parameter affine, the constructed affine motion predictor is inserted into the candidate list only if m ̄v0 ̄, m ̄v1 ̄, and m ̄v2 ̄ are all established, i.e., m ̄v0 ̄, m ̄v1 ̄, and m ̄v2 ̄ are all used as the estimated CPMVs for the top-left (coordinates (x0, y0)), top-right (coordinates (x1, y1)), and bottom-right (coordinates (x2, y2)) of the current block position.

[0153] When inserting the constructed affine motion predictor into the candidate list, no pruning process is applied.

[0154] 3) Regular AMVP motion predictor

[0155] The following is applied until the number of affine motion predictors reaches a maximum: 1) If available, derive the affine motion predictor by setting all CPMVs equal to m̂v2̂. 2) If available, derive the affine motion predictor by setting all CPMVs equal to m̂v1̂. 3) If available, set all CPMVs to m̂v0̂ and derive the affine motion predictor. 4) Derive affine motion predictors by setting all CPMVs equal to HEVCTMVP, if available. 5) Derive the affine motion predictor by setting all CPMVs to zero MV.

[0156] In addition, m ̄v i is already derived in the constructed affine motion predictor.

[0157] Figure 18A shows an example of a four-parameter affine model, and Figure 18B shows an example of a six-parameter affine model.

[0158] FIG. 19 shows an example of the MVP of the inherited affine candidate AF_INTER.

[0159] FIG. 20 shows an example of the MVP of AF_INTER for the constructed affine candidates.

[0160] In AF_INTER mode, if a 4 / 6 parameter affine mode is used, 2 / 3 control points are needed, and therefore 2 / 3 MVDs need to be coded for these control points, as shown in Figure 18. In existing implementations, it is proposed to derive the MVs as follows: mvd1 and mvd2 are predicted from mvd0.

[0161] mv0=m ̄v0 ̄+mvd0 mv1=m ̄v1 ̄+mvd1+mvd0 mv2=m ̄v2 ̄+mvd2+mvd0

[0162] where m ̄v i  ̄、mvd i , mv1 are the predicted motion vector, the motion vector difference, and the motion vector of the upper left pixel (i=0), the upper right pixel (i=1), and the lower left pixel (i=2), respectively, as shown in Figure 18B. Note that the addition of two motion vectors (e.g., mvA(xA, yA) and mvB(xB, yB)) is equal to the sum of the two modules separately, that is, newMV=mvA+mvB, and the two modules of newMV are set to (xA+xB) and (yA+yB), respectively.

[0163] 2.3.3.3 AF_MERGE Mode

[0164] When applying CU in AF_MERGE mode, the CU obtains the first block coded in affine mode from the valid neighboring reconstructed blocks. The selection order of candidate blocks is from left, top, top right, bottom left to top left, as shown in the order of ABCDE in Figure 21A. For example, if the neighboring bottom left block is coded in affine mode as shown by A0 in Figure 21B, the control point (CP) motion vector mv0 of the top left corner, top right corner, and bottom left corner of the neighboring CU / PU containing block A is N , mv1 Nand mv2 N Then, take out mv0 N , mv1 N and mv2 N Based on the above, the upper left / upper right / lower left motion vector mv0 in the current CU / PU is calculated. C , mv1 C and mv2 C (only used for 6-parameter affine model). Note that in VTM-2.0, the sub-block located in the upper left corner (for example, a 4x4 block in VTM) stores MV0, and the sub-block located in the upper right corner stores mv1 if the current block is affine coded. If the current block is coded using a 6-parameter affine model, the sub-block located in the lower left corner stores mv2; otherwise (in a 4-parameter affine model), the LB stores mv2'. The other sub-blocks store the MV used for MC.

[0165] Current CU mv0 C , mv1 C , mv2 C After deriving the CPMV of,the current CU, we generate the MVF of the current CU according,to the simplified affine motion model equations (1) and (2).,To identify whether the current CU is coded in AF_MERGE mode,,we signal an affine flag in the bitstream if there is,at least one neighboring block coded in affine mode.

[0166] In existing implementations, the affine merge candidate list is constructed using the following steps.

[0167] 1) Insert inherited affine candidates

[0168] An inherited affine candidate means that the candidate is derived from the affine motion model of its valid neighboring affine coded blocks. Up to two inherited affine candidates are derived from the affine motion models of neighboring blocks and inserted into the candidate list. For the left predictor, the scan order is {A0,A1}, and for the above predictor, the scan order is {B0,B1,B2}.

[0169] 2) Insert the constructed affine candidates

[0170] If the number of candidates in the affine merge candidate list is less than MaxNumAffineCand (e.g., 5), insert the constructed affine candidate into the candidate list. The constructed affine candidate means that the candidate is constructed by combining the motion information of the neighborhood of each control point.

[0171] a) First, derive motion information of the control points from the identified spatial and temporal neighborhoods shown in Figure 22. CPk (k=1, 2, 3, 4) represents the kth control point. A0, A1, A2, B0, B1, B2, B3 are spatial positions for predicting CPk (k=1, 2, 3), and T is the temporal position for predicting CP4. The coordinates of CP1, CP2, CP3, and CP4 are (0, 0), (W, 0), (H, 0), and (W, H), respectively, where W and H are the width and height of the current block.

[0172] The motion information of each control point is acquired according to the following priority. - For CP1, the check priority is B2->B3->A2. If available, use B2. Otherwise, if B2 is available, use B3. If both B2 and B3 are unavailable, use A2. If all three candidates are unavailable, the motion information of CP1 cannot be obtained. - For CP2, the check priority is B1->B0. - For CP3, the check priority is A1->A0. - Use T for CP4.

[0173] b) Then, use combinations of these control points to construct affine merge candidates. I. To construct a six-parameter affine candidate, motion information for three control points is required. The three control points can be selected from one of the following four combinations: {CP1, CP2, CP4}, {CP1, CP2, CP3}, {CP2, CP3, CP4}, {CP1, CP3, CP4}. The combinations {CP1, CP2, CP3}, {CP2, CP3, CP4}, and {CP1, CP3, CP4} are converted into a six-parameter motion model represented by the top-left, top-right, and bottom-left control points. II. To construct a four-parameter affine candidate, we need motion information for two control points. The two control points may be selected from one of two pairs ({CP1,CP2}, {CP1,CP3}). We convert these pairs into a four-parameter motion model represented by the top-left and top-right control points. III. Insert the constructed affine candidate combinations into the candidate list in the following order: {CP1,CP2,CP3},{CP1,CP2,CP4},{CP1,CP3,CP4},{CP2,CP3,CP4},{CP1,CP2},{CP1,CP3} i. For each combination, check the reference index of list X for each CP, if they are all the same, this combination has a valid CPMV for list X. If this combination does not have a valid CPMV for both list 0 and list 1, this combination is marked as invalid. Otherwise, it is valid and the CPMV is put into the sub-block merge list.

[0174] 3) Padding with zero motion vectors

[0175] If the number of candidates in the affine merge candidate list is less than five, then zero motion vectors with a reference index of zero are inserted into the candidate list until the list is full.

[0176] Specifically, for the sub-block merge candidate list, MV is set to (0,0) and the prediction direction is set to uni-prediction from list 0 (for P slices) and bi-prediction (for B slices) for the four-parameter merge candidates.

[0177] 2.3.4 Referencing the Current Picture

[0178] Intra-block copy (also known as IBC, or intra-picture block compensation), also known as current picture referencing (CPR), was adopted in the HEVC Screen Content Coding Extension (SCC). This tool is highly efficient for coding screen content video, in that repeated patterns in text- and graphics-rich content frequently occur within the same picture. Using the previously reconstructed blocks with the same or similar patterns as predictors can effectively reduce prediction errors and thus improve coding efficiency. Figure 23 shows an example of intra-block compensation.

[0179] Similar to the design of the CRP in HEVC SCC and VVC, the use of IBC mode is signaled at both the sequence level and the picture level. If IBC mode is enabled in the sequence parameter set (SPS), it can be enabled at the picture level. When IBC mode is enabled at the picture level, the current reconstructed picture is treated as a reference picture. Therefore, no syntax changes at the block level are required on top of the existing VVC inter mode to signal the use of IBC mode.

[0180] Key Features: - It is treated as a normal inter mode. Therefore, merge mode and skip mode are available even in IBC mode. We unify the merge candidate list construction, which includes merge candidates from neighboring positions coded in IBC mode or HEVC inter mode. Based on the selected merge index, the current block under merge or skip mode is either merged with a neighbor coded in IBC mode or coded in normal inter mode with a different picture as a reference picture. - The block vector prediction and coding scheme for IBC mode reuses the scheme used for motion vector prediction and coding in HEVC inter mode (AMVP and MVD coding). - Motion vectors for IBC mode, also called block vectors, are coded with integer pixel precision, but after decoding they are stored in memory with 1 / 16 pixel precision because 1 / 4 pixel precision is required in the interpolation and deblocking stages. When used for motion vector prediction for IBC mode, the stored vector predictor is right-shifted by 4. - Search Range: Limited to the current CTU. - CPR is not allowed if affine mode / triangular mode / GBI / weighted prediction is enabled.

[0181] 2.3.5 Merge List Design in VVC

[0182] There are three different merge list building processes supported by VVC:

[0183] 1) Sub-block merge candidate list: Contains ATMVP and affine merge candidates. One merge list construction process is shared for both affine and ATMVP modes. Note that ATMVP and affine merge candidates may be added in order. The size of the sub-block merge list is signaled in the slice header, and its maximum value is 5.

[0184] 2) Single-Predictive TPM Merge List: In triangular prediction mode, two partitions share one merge list construction process, and the two partitions may select their own merge candidate indexes. When constructing this merge list, we check the spatial neighboring blocks and the two temporal blocks of this block. The motion information derived from the spatial neighbors and temporal blocks is called canonical motion candidates in our IDF. These canonical motion candidates are further used to derive multiple TPM candidates. Note that this transformation is performed at the whole block level, and the two partitions may use different motion vectors to generate their own prediction blocks. Single-prediction TPM merge list size is fixed at 5.

[0185] 3) Regular merge list: For the remaining coding blocks, one merge list construction process is shared. Spatial / temporal / HMVP, pairwise composite bi-predictive merge candidates, and zero motion candidates may be inserted in this order. The size of the regular merge list is signaled in the slice header, and its maximum value is 6.

[0186] 2.3.5.1 Sub-block Merge Candidate List

[0187] Note that in addition to the regular merge list of non-subblock merge candidates, it is recommended to put all subblock-related motion candidates into a separate merge list.

[0188] Sub-block related motion candidates are put into a separate merge list, called the "sub-block merge candidate list."

[0189] In one example, the sub-block merging candidate list includes affine merging candidates, ATMVP candidates, and / or sub-block based STMVP candidates.

[0190] 2.3.5.1.1 Alternative ATMVP Implementation

[0191] In this contribution, we move the ATMVP merge candidate in the regular merge list to the first position in the affine merge list. All merge candidates in the new list (i.e., the sub-block based merge candidate list) are based on the sub-block coding tool.

[0192] 2.3.5.1.2 ATMVP for VTM-3.0

[0193] In VTM-3.0, in addition to the normal merge candidate list, a special merge candidate list called the sub-block merge candidate list (also known as the affine merge candidate list) is added. The sub-block merge candidate list satisfies candidates in the following order: b. ATMVP candidates (may or may not be available) c. Inherited affine candidates d. Constructed affine candidates e. Padding as a zero MV4 parameter affine model

[0194] The maximum number of candidates in the subblock merge candidate list (denoted as ML) is derived as follows. 1) If the ATMVP usage flag (e.g., the flag may be named "sps_sbTMVp_enabled_flag") is on (equal to 1) but the affine usage flag (e.g., the flag may be named "sps_affine_enabled_flag") is off (equal to 0), then ML is set to 1. 2) If the ATMVP usage flag is off (equal to 0) and the affine usage flag is off (equal to 0), then ML is set equal to 0. In this case, the sub-block merge candidate list is not used. 3) Otherwise (use affine flag is on (equal to 1), use ATMVP flag is on or off), ML is signaled from the encoder to the decoder. Valid ML is 0<=ML<=5.

[0195] When constructing the sub-block merge candidate list, first check the ATMVP candidate: If any one of the following conditions is true, skip the ATMVP candidate and do not include it in the sub-block merge candidate list: 1) The ATMVP usage flag is OFF. 2) The TMVP usage flag (e.g., if signaled at the slice level, the flag may be named "slice_temporal_mvp_enabled_flag") is off. 3) The reference picture with reference index 0 in reference list 0 is the same as the current picture (is CPR).

[0196] ATMVP in VTM-3.0 is much simpler than ATMVP in JEM. When generating ATMVP merge candidates, the following process is applied: a. Check neighboring blocks A1, B1, B0, A0 to find the first block that is inter-coded but not CPR-coded, denoted as block X, as shown in FIG. b. Initialize TMV=(0,0). If block X has one MV (denoted as MV'), refer to the collocated reference picture (if signaled in the slice header) and set TMV equal to MV'. c. Let the center point of the current block be (x0,y0), then locate the corresponding position of (x0,y0) in the collocated picture as M=(x0+MV'x,y0+MV'y). Find the block Z that contains M. i. If Z is intra-coded, ATMVP is not available. ii. If Z is inter-coded, the two lists MVZ_0 and MVZ_1 of block Z are scaled and stored as MVdefault0 and MVdefault1 to (Reflist0 index0) and (Reflist1 index0). d. For each 8x8 sub-block, assume its center point is (x0S, y0S), then locate the corresponding position of (x0S, y0S) in the collocated picture as MS = (x0S + MV'x, y0S + MV'y). Find the block ZS that contains MS. i. If ZS is intra-coded, MVdefault0 and MVdefault1 are assigned to the sub-block. ii. If ZS is inter-coded, the two lists MVZS_0 and MVZS_1 of block ZS are scaled to (Reflist0 index0) and (Reflist1 index0) and assigned to the sub-blocks.

[0197] MV Clipping and Masking in ATMVP:

[0198] When specifying a corresponding position such as M or MS in a collocated picture, it is clipped to fit within a given area. The size of a CTU is S × S, where S = 128 in VTM-3.0. If the top left position of a collocated CTU is (xCTU, yCTU), the position of the corresponding position M or MS in (xN, yN) must be within the valid area xCTU<=xN. <xCTU+S+4;yCTU<=yN<yCTU+Sにクリッピングされる。

[0199] Besides clipping, (xN,yN) is also masked as xN=xN&MASK, yN=yN&MASK, where MASK is ~(2 N (xN - yN is an integer equal to 8 (-1), where N=3 and at least 3 bits are set to 0. Therefore, xN and yN must be multiples of 8. ("~" represents the bitwise complement operator.)

[0200] FIG. 24 shows an example of valid corresponding regions in a collocated picture.

[0201] 2.3.5.1.3 Syntax Design of Slice Header

[0202] [Table 3]

[0203] 2.3.5.2 Regular Merge List

[0204] Unlike the merge list design, VVC employs a history-based motion vector prediction (HMVP) method.

[0205] The HMVP stores the coded motion information. The motion information of the coded block is defined as an HMVP candidate. Multiple HMVP candidates are stored in a table called the HMVP table, which is maintained on-the-fly during the coding / decoding process. When starting coding / decoding of a new slice, the HMVP table is emptied. Whenever there is an inter-coded block, the associated motion information is added as a new HMVP candidate to the last entry of the table. The overall coding flow is shown in Figure 25.

[0206] HMVP candidates can be used in both the AMVP and merge candidate list construction processes. Figure 26 shows the modified merge candidate list construction process (highlighted in blue). After inserting the TMVP candidates, if the merge candidate list is not full, the HMVP candidates stored in the HMVP table can be used to fill the merge candidate list. Considering that a block usually has a high correlation with its nearest neighbors in terms of motion information, the HMVP candidates in the table are inserted in descending order of index. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy elimination is applied to the HMVP candidates. When the total number of available merge candidates reaches the signalable maximum number of mergeable merge candidates, the merge candidate list construction process terminates.

[0207] 2.4 Rounding of MV

[0208] In VVC, when MV is right-shifted, MV is rounded toward 0. In a formulated form, when MV(MVx, MVy) is right-shifted by N bits, the result MV'(MVx', MVy') is derived as follows:

[0209] MVx'=(MVx+((1<<N)> >1)-(MV_x>=0?1:0))>>N;

[0210] MVy'=(MVy+((1<<N)> >1)-(MVy>=0?1:0))>>N;

[0211] 2.5 Reference Picture Resampling (RPR) Implementation

[0212] ARC, also known as Reference Picture Resampling (RPR), is built into existing and upcoming video standards.

[0213] In some embodiments of RPR, TMVP is disabled if the collocated picture has a different resolution than the current picture, and BDOF and DMVR are disabled if the resolution of the reference picture is different from the current picture.

[0214] To handle normal MC when the resolution of the reference picture is different from that of the current picture, the interpolation section is defined as follows:

[0215] 8.5.6.3 Fractional Sample Interpolation

[0216] 8.5.6.3.1 Overview

[0217] The inputs to this process are:

[0218] - a luminance position (xSb, ySb) specifying the top-left sample of the current coding sub-block relative to the top-left luminance sample of the current picture; - the variable sbWidth, which defines the width of the current coding sub-block, - variable sbHeight, which specifies the height of the current coding sub-block; - motion vector offset mvOffset, - fine-tuned motion vectors refMvLX, - the selected reference picture sample array refPicLX, - 1 / 2 sample interpolation filter index hpelIfIdx, - bidirectional optical flow flag bdofFlag, - A variable cIdx that specifies the color component index of the current block.

[0219] The output of this process is: - predSamplesLX, a (sbWidth+brdExtSize)x(sbHeight+brdExtSize) array of predicted sample values.

[0220] The prediction block boundary extension size brdExtSize is derived as follows. brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb] && sps_affine_prof_enabled_flag))?2:0 (8-752)

[0221] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in luma samples.

[0222] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in luma samples.

[0223] The motion vector mvLX is set equal to (refMvLX-mvOffset). - If cIdx is equal to 0, the following applies: - The scaling factor and its fixed-point representation are defined as follows: hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL (8-753) vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754) - Let (xIntL, yIntL) be the luma position given in full sample units, and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are used in this section only to specify a fractional sample position within the reference sample array refPicLX. - Bounding block for reference sample padding (xSbInt L ,ySbInt L ) equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)). - Each luminance sample position (x L =0..sbWidth-1+brdExtSize,y L = 0..sbHeight-1+brdExtSize), the corresponding predicted luminance sample value predSamplesLX[x L ][y L ] is derived as follows: - (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position indicated by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. L , refx L , refySb L , refy L is derived as follows: refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755) refx L=((Sign(refxSb)*((Abs(refxSb)+128)>>8)+x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756) refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757) refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL*((vert_scale_fp+8)>>4))+32)>>6 (8-758) - variable xInt L , yInt L , xFrac L , and yFrac L is derived as follows: xInt L =refx L >>4 (8-759) yInt L =refy L >>4 (8-760) xFrac L =refx L &15 (8-761) yFrac L =refy L &15 (8-762) - If bdofFlag is equal to TRUE (sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[xSb][ySb] is equal to TRUE), and one or more of the following conditions are true, the predicted luminance sample values ​​predSamplesLX[x L ][y L ] is used as input (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) is derived by calling the luminance integer sample extraction process using refPicLX. 1.x Lis equal to 0. 2.x L is equal to sbWidth+1. 3.y L is equal to 0. 4.y L is equal to sbHeight+1. - Otherwise, (xIntL-(brdExtSize>0?1:0),yIntL-(brdExtSize>0?1:0)), (xFracL,yFracL), (xSbInt L ,ySbInt L ), refPicLX, hpelIfIdx, sbWidth, sbHeight, and (xSb, ySb) as inputs to invoke the luma sample 8-tap interpolation filtering process to derive the predicted luma sample values ​​predSamplesLX[xL][yL]. - Otherwise (cIdx is not equal to 0), the following applies: - Let (xIntC,yIntC) be the chroma position given in full sample units, and (xFracC,yFracC) be the offset given in 1 / 32 sample units. These variables are used in this section only to specify the position of a general fractional sample within the reference sample array refPicLX. - The top-left coordinate of the bounding block (xSbIntC, ySbIntC) for the reference sample padding is set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)). For each chroma sample position (xC=0..sbWidth-1, yC=0..sbHeight-1) in the predicted chroma sample array predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows: - (refxSb C ,refySb C ) and (refx C ,refy C) is the chroma position indicated by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. C , refySb C , refx C , refy C is derived as follows: refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763) refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9)+xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764) refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765) refyC=((Sign(refySb C )*((Abs(refySb C )+256)>>9)+yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766) - variable xInt C , yInt C , xFrac C , yFrac C is derived as follows: xInt C =refx C >>5 (8-767) yInt C =refy C >>5 (8-768) xFrac C =refy C &31 (8-769) yFrac C =refy C &31 (8-770) - The predicted sample values ​​predSamplesLX[xC][yC] are derived by invoking the process specified in Section 8.5.6.3.4 using (xIntC,yIntC), (xFracC,yFracC), (xSbIntC,ySbIntC), sbWidth, sbHeight, and refPicLX as input.

[0224] 8.5.6.3.2 Luminance Sample Interpolation Filtering Process

[0225] The inputs to this process are: - Full sample unit (xInt L ,yInt L ) luminance position, - Luminance position in fractional samples (xFrac L ,yFrac L ), - a full sample unit (xSbInt) that specifies the top-left sample of the border block for padding of the reference sample relative to the top-left luma sample of the reference picture L ,ySbInt L ) luminance position, - Luminance reference sample array refPicLX L , - 1 / 2 sample interpolation filter index hpelIfIdx, - variable sbWidth, which defines the width of the current subblock, - variable sbHeight, which defines the height of the current subblock, - a luminance position (xSb, ySb) defining the top left sample of the current sub-block relative to the top left luminance sample of the current picture,

[0226] The output of this process is the predicted luminance sample value predSampleLX L is.

[0227] The variables shift1, shift2, and shift3 are derived as follows: - Set variable shift1 to Min(4,BitDepthY Set variable shift2 equal to 6, and set variable shift3 equal to Max(2,14-BitDepth Y ) - The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples.

[0228] xFrac L or yFrac L The luminance interpolation filter coefficients f for each 1 / 16 fractional sample position p are equal to L [p] is derived as follows: - If MotionModelIdc[xSb][ySb] is greater than 0 and sbWidth and sbHeight are both equal to 4, the luminance interpolation filter coefficient f L [p] is specified in Table 8-12. - Otherwise, the luminance interpolation filter coefficients f based on hpelIfIdx L [p] is specified in Table 8-11.

[0229] For i=0..7, full sample units (xInt i ,yInt i ) is derived as follows: - If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3) (8-771) yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772) - Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies: xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag? ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3): (8-773) xInt L +i-3) yInt i =Clip3(0,picH-1,yInt L +i-3) (8-774)

[0230] For i=0..7, the luma position in full sample units is further modified as follows: xInt i =Clip3(xSbInt L -3,xSbInt L +sbWidth+4,xInt i ) (8-775) yInt i =Clip3(ySbInt L -3,ySbInt L +sbHeight+4,yInt i ) (8-776)

[0231] Predicted luminance sample value predSampleLX L is derived as follows: - Both xFrac L and yFrac L If is equal to 0, predSampleLX L The value of is derived as follows: predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3 (8-777) - No, xFrac Lis not equal to 0 and yFrac L If is equal to 0, predSampleLX L The value of is derived as follows: predSampleLX L =(Σ 7 i=0 f L [xFrac L ][i]*refPicLX L [xInt i ][yInt3])>>shift1 (8-778) - No, xFrac L is equal to 0, and yFrac L If is not equal to 0, predSampleLX L The value of is derived as follows: predSampleLX L =(Σ 7 i=0 f L [yFrac L ][i]*refPicLX L [xInt3][yInt i ])>>shift1 (8-779) - No, xFrac L is not equal to 0 and yFrac L If is not equal to 0, predSampleLX L The value of is derived as follows: - The sample array temp[n] for n=0..7 is derived as follows: temp[n]=(Σ 7 i=0 f L [xFrac L ][i]*refPicLX L [xInt i ][yInt n ])>>shift1 (8-780) - Predicted luminance sample value predSampleLX L is derived as follows: predSampleLX L =(Σ 7 i=0 f L[yFrac L ][i]*temp[i])>>shift2 (8-781)

[0232] [Table 4]

[0233] [Table 5]

[0234] 8.5.6.3.3 Luminance Integer Sample Extraction Process

[0235] The inputs to this process are: - Full sample unit (xInt L ,yInt L ) luminance position, - Luminance reference sample array refPicLX L ,

[0236] The output of this process is the predicted luminance sample value predSampleLX L is. This variable shift is Max(2,14-BitDepth Y ). The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples.

[0237] The luminance position in full sample units (xInt, yInt) is derived as follows: xInt=Clip3(0,picW-1,sps_ref_wraparound_enabled_flag? ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L ):xInt L )(8-782) yInt=Clip3(0,picH-1,yInt L ) (8-783)

[0238] Predicted luminance sample value predSampleLX L is derived as follows: predSampleLX L =refPicLX L [xInt][yInt]< <shift3 (8-784)

[0239] 8.5.6.3.4 Chroma Sample Interpolation

[0240] The inputs to this process are: - Full sample unit (xInt C ,yInt C ) chroma position, - Chroma position in 1 / 32 fractional samples (xFrac C ,yFrac C ), - a chroma position in full sample units (xSbIntC, ySbIntC) specifying the top-left sample of the border block for reference sample padding relative to the top-left chroma sample of the reference picture; - variable sbWidth, which defines the width of the current subblock, - variable sbHeight, which defines the height of the current subblock, - Chroma reference sample array refPicLX C .

[0241] The output of this process is the predicted chroma sample value predSampleLX. C is.

[0242] The variables shift1, shift2, and shift3 are derived as follows: - Set variable shift1 to Min(4,BitDepth C Set variable shift2 equal to -8, set variable shift3 equal to Max(2,14-BitDepthC ) - variable picW C is set equal to pic_width_in_luma_samples / SubWidthC, and the variable picH C is set equal to pic_height_in_luma_samples / SubHeightC.

[0243] Table 8-13 shows the xFrac C or yFrac C Chrominance interpolation filter coefficients f for each 1 / 32 fractional sample position p are equal to C Indicates [p].

[0244] The variable xOffset is set equal to (sps_ref_wraparound_offset_minus1+1)*MinCbSizeY) / SubWidthC.

[0245] For i=0..3, full sample units (xInt i ,yInt i ) is derived as follows: - If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInt i =Clip3(SubPicLeftBoundaryPos / SubWidthC,SubPicRightBoundaryPos / SubWidthC,xInt L +i) (8-785) yInt i =Clip3(SubPicTopBoundaryPos / SubHeightC,SubPicBotBoundaryPos / SubHeightC,yInt L +i) (8-786) - Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies: xInt i =Clip3(0,picW C-1,sps_ref_wraparound_enabled_flag?ClipH(xOffset,picW C ,xInt C +i-1): (8-787) xInt C +i-1) yInt i =Clip3(0,picH C -1,yInt C +i-1) (8-788)

[0246] Full sample unit (xInt i ,yInt i ) for i=0..3 is further modified as follows: xInt i =Clip3(xSbIntC-1,xSbIntC+sbWidth+2,xInt i ) (8-789) yInt i =Clip3(ySbIntC-1,ySbIntC+sbHeight+2,yInt i ) (8-790)

[0247] Predicted chroma sample value predSampleLX C is derived as follows: - xFrac C and yFrac C If both are equal to 0, then predSampleLX C The value of is derived as follows: predSampleLX C =refPicLX C [xInt1][yInt1]< <shift3 (8-791) - No, xFrac C is not equal to 0 and yFrac C If is equal to 0, predSampleLX C The value of is derived as follows: predSampleLX C =(Σ 3 i=0 fC [xFrac C ][i]*refPicLX C [xInt i ][yInt1])>>shift1 (8-792) - No, xFrac C is equal to 0, and yFrac C If is not equal to 0, predSampleLX C The value of is derived as follows: predSampleLXC=(Σ 3 i=0 f C [yFrac C ][i]*refPicLX C [xInt1][yInt i ])>>shift1 (8-793) - No, xFrac C is not equal to 0 and yFrac C If is not equal to 0, predSampleLX C The value of is derived as follows: - The sample array temp[n] for n=0..3 is derived as follows: temp[n]=(Σ 3 i=0 f C [xFrac C ][i]*refPicLX C [xInt i ][yInt n ])>>shift1 (8-794) - Predicted chroma sample value predSampleLX C is derived as follows: predSampleLX C =(f C [yFrac C ][0]*temp[0]+f C [yFrac C ][1]*temp[1]+f C [yFrac C ][2]*temp[2]+ (8-795) f C [yFrac C][3]*temp[3])>>shift2

[0248] [Table 6] [Table 7]

[0249] 2.6 Subpicture Implementation

[0250] In the current syntax design of sub-pictures in existing implementations, the position and dimensions of a sub-picture are derived as follows:

[0251] [Table 8]

[0252] If subpics_present_flag is 1, it indicates that the subpicture parameter is currently present in the SPSRBSP syntax. If subpics_present_flag is 0, it indicates that the subpicture parameter is not currently present in the SPSRBSP syntax. NOTE 2 - If the bitstream is the result of a sub-bitstream extraction process and contains only a subset of the subpictures of the input bitstream to the sub-bitstream extraction process, the value of subpics_present_flag may need to be set to 1 in the RBSP of the SPS.

[0253] max_subpics_minus1 plus 1 specifies the maximum number of subpictures that can exist in the CVS. max_subpics_minus1 must be in the range 0 to 254. The value 255 is reserved for future use by ITU-T|ISO / IEC.

[0254] subpic_grid_col_width_minus1 plus1 specifies the width of each element of the subpicture identifier grid in units of 4 samples. The syntax element is Ceil(Log2(pic_width_max_in_luma_samples / 4)) bits in length.

[0255] The variable NumSubPicGridCols is derived as follows: NumSubPicGridCols=(pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*4+3) / (subpic_grid_col_width_minus1*4+4) (7-5)

[0256] subpic_grid_row_height_minus1 plus1 specifies the height of each element of the subpicture identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / 4)) bits.

[0257] The variable NumSubPicGridRows is derived as follows: NumSubPicGridRows=(pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*4+3) / (subpic_grid_row_height_minus1*4+4) (7-6)

[0258] subpic_grid_idx[i][j] specifies the subpicture index at grid position (i, j). The length of the syntax element is Ceil(Log2(max_subpics_minus1+1)) bits.

[0259] The variables SubPicTop[subpic_grid_idx[i][j]], SubPicLeft[subpic_grid_idx[i][j]], SubPicWidth[subpic_grid_idx[i][j]], SubPicHeight[subpic_grid_idx[i][j]], and NumSubPics are derived as follows.

[0260] NumSubPics = 0 for (i = 0; i < NumSubPicGridRows; i++) { for (j = 0; j < NumSubPicGridCols; j++) { if (i == 0) SubPicTop[subpic_grid_idx[i][j]] = 0 else if (subpic_grid_idx[i][j] != subpic_grid_idx[i - 1][j]) { SubPicTop[subpic_grid_idx[i][j]] = i SubPicHeight[subpic_grid_idx[i - 1][j]] = i - SubPicTop[subpic_grid_idx[i - 1][j]] } if (j == 0) SubPicLeft[subpic_grid_idx[i][j]] = 0 (7 - 7) else if (subpic_grid_idx[i][j] != subpic_grid_idx[i][j - 1]) { t SubPicLeft[subpic_grid_idx[i][j]] = j SubPicWidth[subpic_grid_idx[i][j]] = j - SubPicLeft[subpic_grid_idx[i][j - 1]] } if (i == NumSubPicGridRows - 1) SubPicHeight[subpic_grid_idx[i][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]]+1 if(j==NumSubPicGridRows-1) SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j-1]]+1 if(subpic_grid_idx[i][j]>NumSubPics) NumSubPics=subpic_grid_idx[i][j] } }

[0261] subpic_treated_as_pic_flag[i], when set to 1, specifies that the i-th subpicture of each coded picture in the CVS is treated as a picture in the decoding process, except for in-loop filtering operations. subpic_treated_as_pic_flag[i], when set to 0, specifies that the i-th subpicture of each coded picture in the CVS is not treated as a picture in the decoding process, except for in-loop filtering operations. If not present, the value of subpic_treated_as_pic_flag[i] is inferred to be equal to 0.

[0262] 2.7 Combined Inter-Intra Prediction (CIIP)

[0263] Combined Inter-Intra Prediction (CIIP) is adopted as a special merge candidate for VVC. It can only be enabled for WxH blocks with W<=64 and H<=64.

[0264] 3. Shortcomings of existing implementations

[0265] In the current VVC design, ATMVP has the following problems: 1) Whether ATMVP is applied is inconsistent between the slice level and the CU level. 2) In the slice header, ATMVP may be enabled even if TMVP is disabled, but the ATMVP flag is signaled before the TMVP flag. 3) Masking is always done regardless of whether the MV is compressed or not. 4) The valid corresponding area may be too large. 5) The derivation of TMV is very complicated. 6) A better default MV is desirable, even if ATMVP is not available. 7) The MV scaling method in ATMVP may not be efficient. 8) ATMVP should be considered for CPR cases. 9) Even if affine prediction is disabled, the list may contain the default 0 affine merge candidates. 10) The current picture is treated as a long-term reference picture, and other pictures are treated as short-term reference pictures. For both ATMVP and TMVP candidates, motion information from temporal blocks in co-located pictures is scaled to reference pictures with a fixed reference index (i.e., 0 for each reference picture list in the current design). However, when CPR mode is enabled, the current picture is also treated as a reference picture, and the current picture may be added to reference picture list 0 (RefPicList0) with an index equal to 0. a. In TMVP, if the temporal block is coded in CPR mode and the reference picture in RefPicList0 is a short reference picture, the TMVP candidate is set to unavailable. b. If the reference picture in RefPicList0 with index 0 is the current picture and the current picture is an Intra Random Access Point (IRAP) picture, then the ATMVP candidate is set to unavailable. c. For ATMVP sub-blocks within a block, when deriving motion information of the sub-block from a temporal block, if this temporal block is coded in CPR mode, a default ATMVP candidate (derived from a temporal block identified by the starting TMV and the center position of the current block) is used to fill in the motion information of this sub-block. 11) MV is right-shifted to integer precision, but does not follow the rounding rules in VVC. 12) In ATMVP, MV(MVx,MVy) (e.g., TMV is 0) used to specify the location of corresponding blocks in different pictures is used as is to refer to the co-located picture. This is based on the assumption that all pictures have the same resolution. However, when RPR is enabled, different picture resolutions may be used. Similar issues exist with regard to identifying corresponding blocks in co-located pictures to derive sub-block motion information. 13) If the width or height of a block is greater than 32 and the maximum transform block size for a CIIP coded block is 32, the intra prediction signal is generated using the CU size, while the inter prediction signal is generated using the TU size (recursively dividing the current block into multiple 32x32 blocks). Using the CU to derive the intra prediction signal results in lower efficiency.

[0266] There are problems with the current design. First, if the reference picture in RefPicList0 with index 0 is the current picture and the current picture is not an IRAP picture, the ATMVP procedure is still invoked, but since none of the temporal motion vectors can be scaled to fit the current picture, the ATMVP procedure could not find any available ATMVP candidates.

[0267] 4. Examples of Embodiments and Techniques

[0268] The following list of techniques and embodiments should be considered as examples to illustrate the general concepts. These techniques should not be construed in a narrow sense. Furthermore, these techniques can be combined in any way in an encoder or decoder embodiment.

[0269] 1. Whether TMVP is allowed and / or whether CPR is used should be taken into consideration for determining / parsing the maximum number of candidates in the sub-block merge candidate list and / or for determining whether ATMVP candidates should be added to the candidate list. Let the maximum number in the sub-block merge candidate list be ML. a) In one example, ATMVP is inferred to be inapplicable when determining or parsing the maximum number of candidates in a sub-block merge candidate list if the ATMVP usage flag is off (equal to 0) or TMVP is disabled. i. In one example, if the ATMVP use flag is on (equal to 1) and TMVP is disabled, the ATMVP candidate is not added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), TMVP is disabled, and the affine usage flag is off (equal to 0), then ML is set equal to 0, which means that sub-block merging is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1), TMVP is enabled, and the affine usage flag is off (equal to 0), then ML is set equal to 1. b) In one example, ATMVP is inferred to be not applicable when determining or parsing the maximum number of candidates in a sub-block merge candidate list if the ATMVP usage flag is off (equal to 0) or the collocated reference picture of the current picture is the current picture itself. i. In one example, if the ATMVP usage flag is on (equal to 1) and the collocated reference picture of the current picture is the current picture itself, the ATMVP candidate is not added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), the collocated reference picture of the current picture is the current picture itself, and the affine usage flag is off (equal to 0), ML is set equal to 0, which means that sub-block merging is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1), the collocated reference picture of the current picture is not the current picture itself, and the affine usage flag is off (equal to 0), then ML is set equal to 1. c) In one example, if the ATMVP usage flag is off (equal to 0) or the reference picture with reference picture index 0 in reference list 0 is the current picture itself, ATMVP is inferred to be not applicable when determining or parsing the maximum number of candidates in the sub-block merge candidate list. i. In one example, the ATMVP usage flag is on (equal to 1) and the collocated reference picture with reference picture index 0 in reference list 0 is the current picture itself, and no ATMVP candidate is added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), the reference picture with reference picture index 0 in reference list 0 is the current picture itself, and the affine usage flag is off (equal to 0), ML is set equal to 0, which means that sub-block merging is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1), the reference picture with reference picture index 0 in reference list 0 is not the current picture itself, and the affine usage flag is off (equal to 0), then ML is set equal to 1. d) In one example, if the ATMVP usage flag is off (equal to 0) or the reference picture with reference picture index 0 in reference list 1 is the current picture itself, ATMVP is inferred to be not applicable when determining or parsing the maximum number of candidates in the sub-block merge candidate list. i. In one example, the ATMVP usage flag is on (equal to 1) and the collocated reference picture with reference picture index 0 in reference list 1 is the current picture itself, and no ATMVP candidate is added to the sub-block merge candidate list or the ATMVP candidate list. ii. In one example, if the ATMVP usage flag is on (equal to 1), the reference picture with reference picture index 0 in reference list 1 is the current picture itself, and the affine usage flag is off (equal to 0), ML is set equal to 0, which means that sub-block merging is not applicable. iii. In one example, if the ATMVP usage flag is on (equal to 1), the reference picture with reference picture index 0 in reference list 1 is not the current picture itself, and the affine usage flag is off (equal to 0), then ML is set equal to 1.

[0270] 2. If TMVP is disabled at the slice / tile / picture level, ATMVP is implicitly disabled and the ATMVP flag is not signaled. a) In one example, the ATMVP flag is signaled after the TMVP flag in the slice header / tile header / PPS. b) In one example, ATMVP or / and TMVP flags may not be signaled in the slice header / tile header / PPS, but only in the SPS header.

[0271] 3. Whether and how to mask the corresponding position in ATMVP depends on whether and how the MV is compressed. Let (xN, yN) be the corresponding position calculated using the coordinator of the current block / sub-block and the starting motion vector (e.g., TMV) in the collocated picture. a) In one example, if MV does not need to be compressed (e.g., sps_disable_motioncompression signaled in SPS is 1), then (xN, yN) is not masked. Otherwise (MV needs to be compressed), (xN, yN) is masked as xN=xN&MASK, yN=yN&MASK, where MASK is ~(2 M −1), where M is an integer such as 3 or 4. b) 2 each K ×2 K The MV compression method of the MV storage result in the block shares the same motion information and the mask in ATMVP processing is ~(2 M -1). K does not have to be equal to M, for example, M = K + 1. c) The MASKs used for ATMVP and TMVP may be the same or different.

[0272] 4. In one example, the MV compression method may be flexible. a) In one example, the MV compression method can be selected between no compression, 8x8 compression (M=3 in Bullet3.a), or 16x16 compression (M=4 in Bullet3.a). b) In one example, the MV compression method may be signaled in the VPS / SPS / PPS / slice header / tile group header. c) In one example, the MV compression method may be set differently for different standard profiles / levels / tiers.

[0273] 5. The valid corresponding range in ATMVP may be adaptive. a) For example, the valid corresponding region may depend on the width and height of the current block. b) For example, the valid corresponding region may depend on the MV compression method. i. In one example, if the MV compression method is not used, the valid corresponding region is smaller, and if the MV compression method is used, the valid corresponding region is larger.

[0274] 6. The valid corresponding region in ATMVP may be based on a basic region with size MxN smaller than the CTU region. For example, the size of the CTU in VTM-3.0 may be 128x128, and the size of the basic region may be 64x64. Let W and H be the width and height of the current block. a) In one example, if W<=M and H<=N, meaning the current block is inside one basic region, then the valid corresponding regions in ATMVP are the collocated basic region and its extension in the collocated picture. Figure 27 shows an example. i. For example, if the top left position of the placed basic region is (xBR, yBR), the corresponding position in (xN, yN) is the effective region xBR<=xN <xBR+M+4;yBR<=yN<yBR+Nにクリッピングされる。

[0275] FIG. 27 shows an exemplary embodiment of the proposed valid region when the current block is in a basic region (BR).

[0276] FIG. 28 shows an example embodiment of a valid region when the current block is not within a fundamental region. b) In one example, if W>M and H>N, it means that the current block is not within one basic region, and divide the current block into multiple parts. Each part has its own valid corresponding region in the ATMVP. For a position A in the current block, its corresponding position B in the collocated block should be within the valid corresponding region of the part in which position A is located. For example, divide the current block into non-overlapping basic regions. The valid region corresponding to one basic region is its co-located basic region and its extension in the co-located picture. Figure 28 shows an example. 1. For example, suppose the position A of the current block is in one basic region R. Let CR be the co-located basic region of R in the co-located picture. The corresponding position of A in the co-located block is position B, and the top-left position of CR is (xCR, yCR). Then, (xN, yN) of position B is within the valid region xCR<=xN <xCR+M+4;yCR<=yN<yCR+Nにクリッピングされる。

[0277] 7. The motion vectors used in ATMVP to define the positions of corresponding blocks in different pictures can be derived as follows (for example, TMV in 2.3.5.1.2): a) In one example, the TMV is always set equal to a default MV, such as (0,0). i. In one example, the default MV is signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. b) In one example, the TMV is set to one MV stored in the HMVP table in the following way: i. If the HMVP list is empty, the TMV is set equal to the default MV, for example (0,0). ii. Otherwise (HMVP list is not empty), 1. The TMV may be set equal to the first element stored in the HMVP table. 2. Alternatively, the TMV may be set equal to the last element stored in the HMVP table. 3. Alternatively, the TMV may be set equal to a specific MV stored in the HMVP table. In one example, a particular MV references reference list 0. b. In one example, a particular MV references Reference List 1. c. In one example, a particular MV references a particular reference picture in reference list 0, for example, the reference picture with index 0. d. In one example, a particular MV references a particular reference picture in reference list 1, for example the reference picture with index 0. e. In one example, a particular MV references a collocated picture. 4. Alternatively, if a particular MV stored in the HMVP table (eg, as listed in bullet 3.) is not found, the TMV may be set equal to the default MV. In one example, search only the first element stored in the HMVP table to find a particular MV. b. In one example, search only the last element stored in the HMVP table to find the particular MV. c. In one example, some or all of the elements stored in the HMVP table are searched to find a particular MV. 5. Alternatively, or in addition, the TMV obtained from HMVP cannot refer to the current picture itself. 6. Alternatively or additionally, the TMV obtained from the HMVP table may be scaled to fit the collocated picture if not referenced. c) In one example, the TMV is set to the MV of one particular neighboring block, and no other neighboring blocks are included. i. The particular neighboring blocks may be blocks A0, A1, B0, B1, and B2 in FIG. ii. The TMV may be set equal to the default MV if: 1. The block in a particular neighborhood does not exist. 2. Certain neighboring blocks are not inter-coded. iii. The TMV may be set equal to a particular MV stored in a particular nearby block. 1. In one example, a particular MV references reference list 0. 2. In one example, a particular MV references Reference List 1. 3. In one example, a particular MV references a particular reference picture in reference list 0, for example the reference picture with index 0. 4. In one example, a particular MV references a particular reference picture in reference list 1, for example the reference picture with index 0. 5. In one example, a particular MV references a collocated picture. 6. If a particular MV is not found stored in a particular nearby block, the TMV may be set equal to the default MV. iv. The TMV obtained from a particular neighboring block may be scaled to fit the collocated picture if it does not reference it. v. The TMV obtained from a particular neighboring block cannot refer to the current picture itself.

[0278] 8. As disclosed in 2.3.5.1.2, MVdefault0 and MVdefault1 used in ATMVP may be derived as follows: a) In one example, MVdefault0 and MVdefault1 are set equal to (0,0). b) In one example, MVdefaultX (X=0 or 1) is derived from HMVP. i. If the HMVP list is empty, MVdefaultX is set equal to a predefined default MV, such as (0,0). 1. A predefined default MV may be signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. ii. Otherwise (HMVP list is not empty), 1. MVdefaultX may be set equal to the first element stored in the HMVP table. 2. MVdefaultX may be set equal to the last element stored in the HMVP table. 3. MVdefaultX may be set equal only to a specific MV stored in the HMVP table. In one example, a particular MV references a reference list X. b. In one example, a particular MV references a particular reference picture in reference list X, for example the reference picture with index 0. 4. If a particular MV stored in the HMVP table is not found, MVdefaultX may be set equal to a predefined default MV. In one example, only the first element stored in the HMVP table is retrieved. b. In one example, retrieve only the last element stored in the HMVP table. c. In one example, retrieving some or all of the elements stored in the HMVP table. 5. MVdefaultX obtained from the HMVP table may be scaled to fit the collocated picture if not referenced. 6. MVdefaultX obtained from HMVP cannot refer to the current picture itself. c) In one example, MVdefaultX (X=0 or 1) is derived from neighboring blocks. i. The neighboring blocks may include blocks A0, A1, B0, B1, and B2 in FIG. 1. For example, derive MVdefaultX using only one of these blocks. 2. Alternatively, derive MVdefaultX using some or all of these blocks. a. These blocks are checked in order until a valid MVdefaultX is found. 3. If no valid MVdefaultX is found from one or more selected neighboring blocks, it is set equal to a predefined default MV, such as (0,0). a. Predefined default MVs may be signaled in the VPS / SPS / PPS / slice header / tile group header / CTU / CU. ii. If no valid MVdefaultX is found in a particular neighborhood block: 1. The block in a particular neighborhood does not exist. 2. Certain neighboring blocks are not inter-coded. iii. MVdefaultX may be set equal to only a particular MV stored in a particular nearby block. 1. In one example, a particular MV references a reference list X. 2. In one example, a particular MV references a particular reference picture in reference list X, for example the reference picture with index 0. iv. MVdefaultX obtained from a particular neighboring block may be scaled to a particular reference picture, for example the reference picture with index 0 in reference list X. v. MVdefaultX obtained from a particular neighboring block cannot refer to the current picture itself.

[0279] 9. For either sub-block or non-sub-block ATMVP candidates, if one temporal block for one sub-block / full block in a collocated picture is coded in CPR mode, one default motion candidate may be utilized instead. a) In one example, a default motion candidate may be defined as the motion candidate associated with the center position of the current block (e.g., MVdefault0 and / or MVdefault1 used in ATMVP as disclosed in 2.3.5.1.2). b) In one example, a default motion candidate may be defined as a (0,0) motion vector and a reference picture index equal to 0 for both reference picture lists, if available.

[0280] 10. Note that the default motion information in ATMVP processing (e.g., MVdefault0, MVdefault1 used in ATMVP as disclosed in 2.3.5.1.2) may be derived based on the location of the position used in the sub-block motion information derivation process. In this proposed method, the default motion information is directly assigned to the sub-block, so there is no need to further derive motion information. a) In one example, instead of using the center position of the current block, the center position of a sub-block (eg, a center sub-block) in the current block may be used. b) Examples of existing and proposed implementations are shown in Figures 29A and 29B, respectively.

[0281] 11. ATMVP candidates shall always be made available in the following manner: a) Let (x0,y0) be the center point of the current block, then let M = (x0 + MV'x,y0 + MV'y) be the corresponding position of (x0,y0) in the collocated picture. Find the block Z containing M. If Z is intra-coded, derive MVdefault0, MVdefault1 by any of the methods suggested in item 6. b) Alternatively, block Z is not located to obtain motion information, and some methods proposed in item 8 are directly applied to obtain MVdefault0 and MVdefault1. c) Alternatively, the default motion candidate used in ATMVP processing is always available. Based on the current design, if it is set to unavailable (e.g., the temporal block is intra-coded), other motion vectors may be used instead of the default motion candidate. i. In one example, the solution of International Application No. PCT / CN2018 / 124639, which is incorporated herein by reference, may be applied. d) Alternatively or additionally, whether ATMVP candidates are always available depends on other high-level syntactic information. i. In one example, an ATMVP candidate may be set to always be available only if the ATMVP enable flag in the slice / tile / picture header or other video unit is assumed to be true. ii. In one example, the above method may be applicable only when the ATMVP enable flag in the slice header / picture header or other video unit is set to true, and the current picture is not an IRAP picture, and the current picture has not been inserted into RefPicList0 with a reference index equal to 0. e) ATMVP candidates are assigned a fixed index or a fixed group index. If ATMVP candidates are not always available, the fixed index / group index may be inferred to other types of motion candidates (e.g., affine candidates).

[0282] 12. It is noted that whether to include zero motion affine merge candidates in the sub-block merge candidate list should depend on whether affine prediction is enabled. a) For example, if the affine enabled flag is off (sps_affine_enabled_flag is equal to 0), zero motion affine merge candidates are not included in the sub-block merge candidate list. b) Alternatively or additionally, add default motion vector candidates that are non-affine candidates instead.

[0283] 13. Note that non-affine padding candidates may be included in the sub-block merging candidate list. a) If the sub-block merging candidate list is not full, a zero motion non-affine padding candidate may be added. b) When selecting such a padding candidate, the affine_flag of the current block MUST be set to 0. c) Alternatively, if the sub-block merging candidate list is not filled and the use affine flag is off, include a zero motion non-affine padding candidate in the sub-block merging candidate list.

[0284] 14. Let MV0 and MV1 denote the MVs in Reference List 0 and Reference List 1 of the block containing the corresponding position (for example, MV0 and MV1 can be MVZ_0 and MVZ_1 or MVZS_0 and MVZS_1 as described in Section 2.3.5.1.2). Let MV0' and MV1' denote the MVs in Reference List 0 and Reference List 1 to be derived for the current block or sub-block. In that case, MV0' and MV1' should be derived by scaling. a) MV0, if the collocated picture is in reference list 1. b) MV1, if the collocated picture is in reference list 0.

[0285] 15. In reference picture list X (PicRefListX, e.g., X=0), if the current picture is treated as a reference picture with index set to M (e.g., 0), the ATMVP and / or TMVP allow / disable flag may be inferred to be false for a slice / tile or other type of video unit, where M may be equal to the object reference picture index that scales the motion information of the temporal block to PicRefListX in ATMVP / TMVP processing. a) Alternatively, the above method is only applicable if the current picture is an Intra Random Access Point (IRAP) picture. b) In one example, if the current picture in PicRefListX is treated as a reference picture with index set to M (e.g., 0) and / or if the index in PicRefListY is treated as a reference picture with index set to N (e.g., 0), then the ATMVP and / or TMVP allow / disable flags may be inferred to be false. The variables M and N represent the object reference picture indices used in TMVP or ATMVP processing. c) For ATMVP processing, the confirmation bitstream is restricted to follow the rule that the collocated picture from which the motion information of the current block is derived is not the current picture. d) Alternatively, if the above conditions are true, no ATMVP or TMVP processing is invoked.

[0286] 16. If the current picture is a reference picture whose index in the reference picture list X (PicRefListX, e.g., X=0) for the current block is set to M (e.g., 0), ATMVP can still be enabled for this block. a) In one example, the motion information of all sub-blocks points to the current picture. b) In one example, when obtaining motion information of a sub-block from a temporal block, the temporal block is coded with at least one reference picture that points to the current picture of the temporal block. c) In one example, when obtaining motion information of a sub-block from a temporal block, no scaling operation is applied.

[0287] 17. The coding method for sub-block merge indexes will be unified regardless of whether ATMVP is used. a) In one example, for the first L bins, they are context coded. For the remaining bins, they are bypass coded. In one example, L is set to 1. b) Alternatively, for all bins, they are context coded.

[0288] 18. In ATMVP, MVs (MVx, MVy) (e.g., TMV is 0) used to find corresponding blocks in different pictures may be right-shifted to integer precision (denoted as MVx', MVy') using a rounding method similar to the MV scaling process. a) Alternatively, the MVs used to locate corresponding blocks in different pictures in ATMVP (eg, TMV is 0) may be right-shifted to integer precision with the same rounding method as the MV averaging process. b) Alternatively, MVs used to locate corresponding blocks in different pictures in ATMVP (e.g., TMV is 0) may be right-shifted to integer precision with the same rounding method as in adaptive MV resolution (AMVR) processing.

[0289] 19. In ATMVP, MV(MVx,MVy) (e.g., TMV is 0) used to find corresponding blocks in different pictures may be right-shifted to integer precision (denoted as (MVx',MVy')) by rounding towards 0. a) For example, MVx'=(MVx+((1<<N)> >1)-(MVx>=0?1:0))>>N; N is an integer representing the resolution of the MV, e.g., N=4. i. For example, MVx'=(MVx+(MVx>=0?7:8))>>4. b) For example, MVy'=(MVy+((1<<N)> >1)-(MVy>=0?1:0))>>N; N is an integer representing the resolution of the MV, e.g., N=4. i. For example, MVy'=(MVy+(MVy>=0?7:8))>>4.

[0290] 20. In one example, the MV(MVx,MVy) in bullet 18 and bullet 19 is used to define the position of the corresponding block using the center position of the sub-block and the shifted MV, or using the top-left position of the current block and the shifted MV, to derive the default motion information used in ATMVP. a) In one example, MV(MVx,MVy) is used to derive motion information of a sub-block in a current block during ATMVP processing, for example, to define the position of the corresponding block using the center position of the sub-block and the shifted MV.

[0291] 21. The methods proposed in bullets 18, 19 and 20 may be applied to other coding tools that require the location of a reference block in a different picture or in the current picture to be defined by a motion vector.

[0292] 22. The MV(MVx,MVy) (e.g., TMV at 0) used to find corresponding blocks in different pictures in ATMVP may point to the co-located picture or may be scaled. a) In one example, if the width and / or height of the collocated picture (or conformance window therein) differs from the width and / or height of the current picture (or conformance window therein), the MV may be scaled. b) Let the width and height (of the conformance window) of the collocated picture be W1 and H1, respectively. Let the width and height (of the conformance window) of the current picture be W2 and H2, respectively. Then, MV(MVx,MVy) may be scaled as MVx'=MVx*W1 / W2 and MVy'=MVy*H1 / H2.

[0293] 23. The center point of the current block (eg, position (x0, y0) in 2.3.5.1.2) used to derive motion information in ATMVP processing may be further modified by scaling and / or adding an offset. a) In one example, if the width and / or height of the collocated picture (or conformance window therein) differs from the width and / or height of the current picture (or conformance window therein), the center point may be further modified. b) Let X1 and Y1 be the top-left position of the conformance window in the collocated picture. Let X2 and Y2 be the top-left position of the conformance window defined for the current picture. The width and height of the (conformance window of) the collocated picture are denoted as W1 and H1, respectively. Let W2 and H2 be the width and height of the (conformance window of) the current picture, respectively. Then (x0, y0) may be modified as x0' = (x0 - X2) * W1 / W2 + X1, y0' = (y0 - Y2) * H1 / H2 + Y1. i. Alternatively, x0'=x0*W1 / W2, y0'=y0*H1 / H2.

[0294] 24. The corresponding position (e.g., position M in 2.3.5.1.2) used to derive motion information in ATMVP processing may be further modified by scaling and / or adding an offset. a) In one example, if the width and / or height of the collocated picture (or the conformance window therein) differs from the width and / or height of the current picture (or the conformance window therein), the corresponding positions may be further modified. b) Let X1 and Y1 be the top-left position of the conformance window in the collocated picture. Let X2 and Y2 be the top-left position of the conformance window defined for the current picture. The width and height of the (conformance window of) the collocated picture are denoted as W1 and H1, respectively. Let W2 and H2 be the width and height of the (conformance window of) the current picture, respectively. Then M(x,y) may be modified as x'=(x-X2)*W1 / W2+X1 and y'=(y-Y2)*H1 / H2+Y1. i. Alternatively, x'=x*W1 / W2 and y'=y*H1 / H2.

[0295] Subpicture Related

[0296] 25. In one example, if positions (i, j) and (i, j-1) belong to different subpictures, the width of subpicture S ending in column (j-1) may be set equal to j - the leftmost column of subpicture S. a) The following highlights embodiments based on existing implementation examples. NumSubPics=0 for(i=0;i. <NumSubPicGridRows;i++){ for(j=0;j <NumSubPicGridCols;j++){ if(i==0) SubPicTop[subpic_grid_idx[i][j]]=0 else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i-1][j]){ SubPicTop[subpic_grid_idx[i][j]]=i SubPicHeight[subpic_grid_idx[i-1][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]] } if(j==0) SubPicLeft[subpic_grid_idx[i][j]]=0 (7-7) else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i][j-1]){ SubPicLeft[subpic_grid_idx[i][j]]=j SubPicWidth[subpic_grid_idx[i][j-1]]=j-SubPicLeft[subpic_grid_idx[i][j-1]] } if(i==NumSubPicGridRows-1) SubPicHeight[subpic_grid_idx[i][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]]+1 if(j==NumSubPicGridRows-1) SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j-1]]+1 if(subpic_grid_idx[i][j]>NumSubPics) NumSubPics=subpic_grid_idx[i][j] } }

[0297] 26. In one example, the height of a subpicture S that ends at (NumSubPicGridRows-1) rows may be set equal to (NumSubPicGridRows-1) - the top row of subpicture S + 1. a) The following highlights embodiments based on existing implementation examples. NumSubPics=0 for(i=0;i. <NumSubPicGridRows;i++){ for(j=0;j <NumSubPicGridCols;j++){ if(i==0) SubPicTop[subpic_grid_idx[i][j]]=0 else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i-1][j]){ SubPicTop[subpic_grid_idx[i][j]]=i SubPicHeight[subpic_grid_idx[i-1][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]] } if(j==0) SubPicLeft[subpic_grid_idx[i][j]]=0 (7-7) else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i][j-1]){ SubPicLeft[subpic_grid_idx[i][j]]=j SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j-1]] } if(i==NumSubPicGridRows-1) SubPicHeight[subpic_grid_idx[i][j]]=i-SubPicTop[subpic_grid_idx[i][j]]+1 if (j==NumSubPicGridRows-1) SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j-1]]+1 if(subpic_grid_idx[i][j]>NumSubPics) NumSubPics=subpic_grid_idx[i][j]

[0298] 27. In one example, the width of a subpicture S that ends at (NumSubPicGridColumns-1) columns may be set equal to (NumSubPicGridColumns-1) minus the leftmost column of subpicture S, plus one. a) The following highlights embodiments based on existing implementation examples. NumSubPics=0 for(i=0;i. <NumSubPicGridRows;i++){ for(j=0;j <NumSubPicGridCols;j++){ if(i==0) SubPicTop[subpic_grid_idx[i][j]]=0 else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i-1][j]){ SubPicTop[subpic_grid_idx[i][j]]=i SubPicHeight[subpic_grid_idx[i-1][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]] } if(j==0) SubPicLeft[subpic_grid_idx[i][j]]=0 (7-7) else if(subpic_grid_idx[i][j]!=subpic_grid_idx[i][j-1]){ SubPicLeft[subpic_grid_idx[i][j]]=j SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j-1]] } if(i==NumSubPicGridRows-1) SubPicHeight[subpic_grid_idx[i][j]]=i-SubPicTop[subpic_grid_idx[i-1][j]]+1 if(j==NumSubPicGridColumns-1) SubPicWidth[subpic_grid_idx[i][j]]=j-SubPicLeft[subpic_grid_idx[i][j]]+1 if(subpic_grid_idx[i][j]>NumSubPics) NumSubPics=subpic_grid_idx[i][j]

[0299] 28. The subpicture grid must be an integer multiple of the CTU size. a) The following highlights embodiments based on existing implementation examples. subpic_grid_col_width_minus1 plus1 specifies the width of each element of the subpicture identifier grid in units of CtbSizeY. The syntax element is Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)) bits in length. The variable NumSubPicGridCols is derived as follows: NumSubPicGridCols=(pic_width_max_in_luma_samples+subpic_grid_col_width_minus1*CtbSizeY+CtbSizeY-1) / (subpic_grid_col_width_minus1*CtbSizeY+CtbSizeY) (7-5)subpic_grid_row_height_minus1 plus1 specifies the height of each element of the subpicture identifier grid in units of 4 samples. The length of the syntax element is Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)) bits. The variable NumSubPicGridRows is derived as follows: NumSubPicGridRows=(pic_height_max_in_luma_samples+subpic_grid_row_height_minus1*CtbSizeY+CtbSizeY-1) / (subpic_grid_row_height_minus1*CtbSizeY+CtbSizeY) (7-6)

[0300] 29. Conformance constraints are added to ensure that sub-pictures do not overlap each other and that all sub-pictures must contain the entire picture. a) Example embodiments based on existing implementations are highlighted below. subpic_grid_idx[i][j] must be equal to idx if both of the following conditions are met: i>=SubPicTop[idx]and i <SubPicTop[idx]+SubPicHeight[idx]. j>=SubPicLeft[idx] and j <SubPicLeft[idx]+SubPicWidth[idx]. subpic_grid_idx[i][j] must be different from idx unless both of the following conditions are met: i>=SubPicTop[idx]and i <SubPicTop[idx]+SubPicHeight[idx]. j>=SubPicLeft[idx] and j <SubPicLeft[idx]+SubPicWidth[idx].

[0301] RPR related

[0302] 30. A syntax element (e.g., a flag) denoted as RPR_flag is signaled to indicate whether RPR is available in a video unit (e.g., a sequence). RPR_flag may be signaled in an SPS, VPS, or DPS. a) In one example, if RPR is signaled not to be used (e.g., RPR_flag is 0), all widths / heights signaled in the PPS must be the same as the maximum widths / heights signaled in the SPS. b) In one example, if RPR is signaled not to be used (e.g., RPR_flag is 0), all widths / heights of the PPS are not signaled and are inferred to be the maximum width / height signaled in the SPS. c) In one example, if RPR is not signaled to be used (e.g., RPR_flag is 0), conformance window information is not used in the decoding process. Otherwise (if RPR is signaled to be used), conformance window information may be used in the decoding process.

[0303] 31. The interpolation filter used to derive a prediction block for a current block in a motion compensation process may be selected based on whether the resolution of the reference picture is different from that of the current picture, or whether the width and / or height of the reference picture is greater than the resolution of the current picture. a. In one example, if condition A is met, and condition A depends on the dimensions of the current picture and / or the reference picture, an interpolation filter with fewer taps may be applied. i. In one example, condition A is that the resolution of the reference picture is different from that of the current picture. ii. In one example, condition A is that the width and / or height of the reference picture is greater than that of the current picture. iii. In one example, condition A is W1>a*W2 and / or H1>b*H2, where (W1, H1) represent the width and height of the reference picture, (W2, H2) represent the width and height of the current picture, and a and b are two factors, for example, a=b=1.5. iv. In one example, condition A may depend on whether bi-prediction is used. 1) Condition A is met if and only if bi-prediction is used for the current block. v. In one example, condition A may depend on M and N, where M and N represent the width and height of the current block. 1) For example, condition A is satisfied if and only if M*N<=T, where T is an integer such as 64. 2) For example, condition A is satisfied if and only if M<=T1 or N<=T2, where T1 and T2 are integers, e.g., T1=T2=4. 3) For example, condition A is satisfied if and only if M<=T1 and N<=T2, where T1 and T2 are integers, e.g., T1=T2=4. 4) For example, condition A is satisfied if and only if M*N=T, or M=T1 or N=T2, where T, T1, T2 are integers, e.g., T=64, T1=T2=4. 5) In one example, the smaller term in the above sub-bullet may be replaced with a larger one. vi. In one example, a 1-tap filter is applied, i.e., outputting unfiltered integer pixels as the interpolated result. vii. In one example, if the resolution of the reference picture is different from the current picture, a bilinear filter is applied. viii. In one example, if the resolution of the reference picture is different from that of the current picture, or if the width and / or height of the reference picture is greater than the resolution of the current picture, a 4-tap filter or a 6-tap filter is applied. 1) A 6-tap filter may be used for affine motion compensation. 2) A 4-tap filter may be used to interpolate the chroma samples. b. Whether and / or how to apply the method disclosed in bullet 31 may depend on the color component. i. For example, these methods only apply to the luminance component. c. Whether and / or how to apply the method disclosed in bullet31 may depend on the interpolation filtering direction. i. For example, this method only applies to horizontal filtering. ii. For example, this method only applies to vertical filtering.

[0304] CIIP related

[0305] 32. The intra prediction signal used in the CIIP process may be performed at the TU level instead of the CU level (eg, using reference samples outside the TU instead of the CU). a) In one example, if either the width or height of a CU is larger than the maximum transform block size, the CU may be split into multiple TUs, and intra / inter predictions may be generated for each TU, for example, using reference samples outside the TU. b) In one example, if the maximum transform size K is smaller than 64 (eg, K=32), the intra prediction used in CIIP is performed in a recursive manner as in a normal intra-coded block. c) For example, by dividing a KM×KN CIIP coding block into MN K×K blocks, where M and N are integers, intra prediction is performed for each K×K block. The intra prediction of a subsequently coded / decoded K×K block may depend on the reconstructed samples of said coded / decoded K×K block.

[0306] 5. Additional Exemplary Embodiments (Bold text indicates changes to the current version of the standard)

[0307] 5.1 Embodiment #1: Example of syntax design for SPS / PPS / slice header / tile group header

[0308] Changes compared to the VTM3.0.1rC1 standard software are highlighted in large bold font as follows:

[0309] [Table 9]

[0310] 5.2 Embodiment #2: Example of syntax design for SPS / PPS / slice header / tile group header

[0311] 7.3.2.1 Sequence Parameter Set RBSP Syntax

[0312] [Table 10]

[0313] When sps_sbtmvp_enabled_flag is equal to 1, it specifies that a sub-block based temporal motion vector predictor may be used and that pictures containing all slices with slice_type not equal to I in the CVS can be decoded. sps_sbtmvp_enabled_flag equal to 0 specifies that a sub-block based temporal motion vector predictor is not used in the CVS. If sps_sbtmvp_enabled_flag is not present, it is inferred to be equal to 0.

[0314] five_minus_max_num_subblock_merge_cand specifies the maximum number of subblock-based merge motion vector prediction (MVP) candidates supported in a slice subtracted from 5. If five_minus_max_num_subblock_merge_cand is not present, it is inferred to be equal to 5 - sps_sbtmvp_enabled_flag. The maximum number of subblock-based merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows:

[0315] MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand (7-45)

[0316] The value of MaxNumSubblockMergeCand is in the range of 0 to 5.

[0317] 8.3.4.2 Motion Vector and Reference Indices Derivation Process in Sub-Block Merge Mode

[0318] The inputs to this process are:

[0319] ..[No changes to the current VVC specification proposal]

[0320] The output of this process is:

[0321] ...[No changes to the current VVC specification proposal]

[0322] The variables numSbX, numSbY and the subblock merge candidate list subblockMergeCandList are derived by the following sequential steps.

[0323] If sps_sbtmvp_enabled_flag is equal to 1 and (current picture is IRAP and index 0 of reference picture list 0 is the current picture) is not true, the following applies:

[0324] The derivation process for merging candidates from neighboring coding units specified in Section 8.3.2.3 is called with the inputs (xCb, yCb) of the luma coding block, cbWidth, cbHeight, and the luma code block width, where X is either 0 or 1, and outputs the availability flags availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, and availableFlagB2. and availableFlagB2, reference indices refIdxLXA0, refIdxLXA1, refIdxLXB0, refIdxLXB1 and refIdxLXB2, prediction list usage flags predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1 and predFlagLXB2, and motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1 and mvLXB2.

[0325] The subblock-based temporal merge candidate derivation process specified in Section 8.3.4.3 is as follows: xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, where X is 0 or 1, and the luminance position (xCb, yCb), luminance coding block width cbWidth, luminance coding block height cbHeight, availability flags availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, reference indices refIdxLXA0, refIdxLXA1, refIdxLXB0, efIdxLXB1, prediction list usage It is called with flags predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1 and motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1 as input, and outputs the availability flag availableFlagSbCol, the number of luma coding sub-blocks in the horizontal direction numSbX and vertical direction numSbY, the reference index refIdxLXSbCol, the luma motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list usage flag predFlagLXSbCol[xSbIdx][ySbIdx].

[0326] If sps_affine_enabled_flag is equal to 1, the sample locations (xNbA0,yNbA0), (xNbA1,yNbA1), (xNbA2,yNbA2), (xNbB0,yNbB0), (xNbB1,yNbB1), (xNbB2,yNbB2), (xNbB3,yNbB3) and the variables numSbX and numSbY are derived as follows:

[0327] [No changes to the current VVC specification proposal]

[0328] 5.3 Example #3: MV Rounding The syntax changes are based on existing implementations.

[0329] 8.5.5.3 Sub-Block Based Temporal Merge Candidate Derivation Process … The position (xColSb, yColSb) of the collocated sub-block inside ColPic is derived as follows:

[0330] [ka]

[0331] 8.5.5.4 Sub-block based temporal merge-based motion data derivation process … The position (xColCb, yColCb) of the collocated block inside ColPic is derived as follows:

[0332] [ka]

[0333] 5.3 Implementation #3: MV Rounding Example The syntax changes are based on existing implementations.

[0334] 8.5.5.3 Sub-Block Based Temporal Merge Candidate Derivation Process …

[0335] [ka]

[0336] 8.5.5.4 Sub-block based temporal merge-based motion data derivation process …

[0337] [ka]

[0338] 5.4 Implementation #4: Second Example of MV Rounding

[0339] 8.5.5.3 Sub-Block Based Temporal Merge Candidate Derivation Process

[0340] The inputs to this process are: - the luminance position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - a variable cbWidth that specifies the width of the current coding block in luma samples, - A variable cbHeight that specifies the height of the current coding block in luma samples. - the availability flag of the neighboring coding unit availableFlagA1, - a reference index of a neighboring coding unit refIdxLXA1, where X is 0 or 1; - a prediction list usage flag for a neighboring coding unit, predFlagLXA1, where X is 0 or 1; a motion vector at 1 / 16 fractional sample precision mvLXA1 of a neighboring coding unit, where X is 0 or 1;

[0341] The output of this process is: - availability flag availableFlagSbCol, - the number of luminance-coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, - reference indices refIdxL0SbCol and refIdxL1SbCol, - Luma motion vectors in 1 / 16 fractional sample precision mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, - Prediction list usage flags predFlagL0SbCol[xSbIdx][ySbIdx] and predFlagL1SbCol[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1.

[0342] The availability flag availableFlagSbColl is derived as follows: - availableFlagSbCol is set to 0 if one or more of the following conditions are true: - slice_temporal_mvp_enabled_flag equals 0. - sps_sbtmvp_enabled_flag equals 0. - cbWidth is less than 8. - cbHeight is less than 8. - Otherwise, the following ordered steps apply:

[0343] 1. The position (xCtb, yCtb) of the top-left sample of the luma coding tree block including the current coding block and the position (xCtr, yCtr) of the bottom-right center sample of the current luma coding block are derived as follows: xCtb=(xCb>>CtuLog2Size)< <ctulog2size (8-542) yctb="(yCb">>CtuLog2Size)< <CtuLog2Size (8-543) xCtr=xCb+(cbWidth / 2) (8-544) yCtr=yCb+(cbHeight / 2) (8-545)

[0344] 2. The luma location (xColCtrCb, yColCtrCb) is set equal to the top-left sample of the co-located luma coding block containing the location given by (xCtr, yCtr) inside ColPic, relative to the top-left luma sample of the collocated picture specified by ColPic.

[0345] 3. The sub-block-based temporal merge-based motion data derivation process specified in Section 8.5.5.4 is invoked with inputs (xCtb, yCtb), location (xColCtrCb, yColCtrCb), availability flag A1, prediction list usage flag predFlagLXA1, and reference index refIdxLXA1, and motion vector mvLXA1, where X is 0 and 1, and outputs the motion vector ctrMVLX, prediction list usage flag ctrPredFlagLX of the collocated block, and temporal motion vector tempMv, where X is 0 and 1.

[0346] 4. The variable availableFlagSbCol is derived as follows: If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, then availableFlagSbCol is set equal to 0. - Otherwise, availableFlagSbCol is set equal to 1.

[0347] If availableFlagSbCol is equal to 1, the following applies: - The variables numSbX, numSbY, sbWidth, sbHeight, and refIdxLXSbCol are derived as follows: numSbX=cbWidth>>3 (8-546) numSbY=cbHeight>>3 (8-547) sbWidth=cbWidth / numSbX (8-548) sbHeight=cbHeight / numSbY (8-549) refIdxLXSbCol=0 (8-550) When xSbIdx=0..numSbX-1 and ySbIdx=0...numSbY-1, the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list usage flag predFlagsLXSbCol[xSbIdx] are derived as follows: The luminance position (xSb, ySb) that defines the top-left sample of the current coding sub-block relative to the top-left luminance sample of the current picture is derived as follows: xSb=xCb+xSbIdx*sbWidth+sbWidth / 2 (8-551) ySb=yCb+ySbIdx*sbHeight+sbHeight / 2 (8-552) The position (xColSb, yColSb) of the collocated sub-block inside ColPic is derived as follows:

[0348] [ka]

[0349] The variable currCb defines the luma coding block that contains the current coding sub-block in the current picture. The variable colCb defines the luminance coding block containing the modified position given by ((xColSb>>3)<<3, (yColSb>>3)<<3) inside ColPic. - The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the collocated picture specified by ColPic. - The co-located motion vector derivation process specified in clause 8.5.2.12 is called with inputs currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1, and the output is assigned to the motion vector of sub-block mvL0SbCol[xSbIdx][ySbIdx] and availableFlagL0SbCol. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked with inputs currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1, and the output is assigned to the motion vector of sub-block MVL1SbCol[xSbIdx][ySbIdx] and availableFlagL1SbCol. - If availableFlagL0SbCol and availableFlagL1SbCol are both equal to 0, then if X is 0 and 1, the following applies: mvLXSbCol[xSbIdx][ySbIdx]=ctrMvLX (8-556) predFlagLXSbCol[xSbIdx][ySbIdx]=ctrPredFlagLX (8-557)

[0350] 8.5.5.4 Sub-block based temporal merge-based motion data derivation process

[0351] The inputs to this process are: - the position (xCtb, yCtb) of the top-left sample of the luma coding tree block containing the current coding block, - The location (xColCtrCb, yColCtrCb) of the top-left sample of the co-located luma coding block with the bottom-right center sample. - the availability flag of the neighboring coding unit availableFlagA1, - the reference index of the neighboring coding unit refIdxLXA1, - prediction list usage flag predFlagLXA1 for neighboring coding units; - Motion vectors of neighboring coding units at 1 / 16 fractional sample precision mvLXA1.

[0352] The output of this process is: - the motion vectors ctrMvL0 and ctrMvL1, - Prediction list usage flags ctrPredFlagL0, ctrPredFlagL1, - Temporal motion vector tempMv.

[0353] The variable tempMv is set as follows: tempMv[0]=0 (8-558) tempMv[1]=0 (8-559)

[0354] The variable currPic defines the current picture.

[0355] If availableFlagA1 is equal to TRUE, the following applies: - tempMv is set equal to mvL0A1 if all of the following conditions are true: - predFlagL0A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A1]) is equal to 0, - Otherwise, if all of the following conditions are true, then tempMv is set equal to mvL1A1: - Slice type is the same as B, - predFlagL1A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[1][refIdxL1A1]) equals 0.

[0356] [ka]

[0357] The array colPredMode is set equal to the prediction mode array CuPredMode[0] of the collocated picture specified by ColPic.

[0358] The motion vectors ctrMvL0 and ctrMvL1 and the prediction list usage flags ctrPredFlagL0 and ctrPredFlagL1 are derived as follows. - If colPredMode[xColCb][yColCb] is equal to MODE_INTER, the following applies: The variable currCb defines the luma coding block containing (xCtrCb, yCtrCb) in the current picture. The variable colCb defines the luminance coding block containing the modified position given by ((xColCb>>3)<<3, (yColCb>>3)<<3) inside ColPic. - The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the collocated picture specified by ColPic. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked by taking as input currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1, and assigning the output to ctrMvL0 and ctrPredFlagL0. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked by taking as input currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1, and assigning the output to ctrMvL1 and ctrPredFlagL1. - Otherwise, the following applies: ctrPredFlagL0=0 (8-563) ctrPredFlagL1=0 (8-564)

[0359] 5.5 Implementation #5: Third Example of MV Rounding

[0360] 8.5.5.3 Sub-Block Based Temporal Merge Candidate Derivation Process

[0361] The inputs to this process are: - the luminance position (xCb, yCb) of the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture, - a variable cbWidth that specifies the width of the current coding block in luma samples, - A variable cbHeight that specifies the height of the current coding block in luma samples. - the availability flag of the neighboring coding unit availableFlagA1, - a reference index of a neighboring coding unit refIdxLXA1, where X is 0 or 1; - a prediction list usage flag for a neighboring coding unit, predFlagLXA1, where X is 0 or 1; a motion vector at 1 / 16 fractional sample precision mvLXA1 of a neighboring coding unit, where X is 0 or 1;

[0362] The output of this process is: - Availability flag availableFlagSbCol, - the number of luminance-coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, - reference indices refIdxL0SbCol and refIdxL1SbCol, - Luma motion vectors in 1 / 16 fractional sample precision mvL0SbCol[xSbIdx][ySbIdx] and mvL1SbCol[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1, - Prediction list usage flags predFlagL0SbCol[xSbIdx][ySbIdx] and predFlagL1SbCol[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1.

[0363] The availability flag availableFlagSbColl is derived as follows: - availableFlagSbCol is set to 0 if one or more of the following conditions are true: - slice_temporal_mvp_enabled_flag equals 0. - sps_sbtmvp_enabled_flag equals 0. - cbWidth is less than 8. - cbHeight is less than 8. - Otherwise, the following ordered steps apply:

[0364] 5. The position (xCtb, yCtb) of the top-left sample of the luma coding tree block including the current coding block and the position (xCtr, yCtr) of the bottom-right center sample of the current luma coding block are derived as follows: xCtb=(xCb>>CtuLog2Size)< <ctulog2size (8-542) yctb="(yCb">>CtuLog2Size)< <CtuLog2Size (8-543) xCtr=xCb+(cbWidth / 2) (8-544) yCtr=yCb+(cbHeight / 2) (8-545)

[0365] 6. The luma location (xColCtrCb, yColCtrCb) is set equal to the top-left sample of the co-located luma coding block that contains the location given by (xCtr, yCtr) inside ColPic, relative to the top-left luma sample of the collocated picture specified by ColPic.

[0366] 7. The sub-block-based temporal merge-based motion data derivation process specified in Section 8.5.5.4 is invoked with inputs (xCtb, yCtb), location (xColCtrCb, yColCtrCb), availability flag A1, prediction list usage flag predFlagLXA1, and reference index refIdxLXA1, and motion vector mvLXA1, where X is 0 and 1, and outputs the motion vector ctrMVLX, prediction list usage flag ctrPredFlagLX of the collocated block, and temporal motion vector tempMv, where X is 0 and 1.

[0367] 8. The variable availableFlagSbCol is derived as follows: If both ctrPredFlagL0 and ctrPredFlagL1 are equal to 0, then availableFlagSbCol is set equal to 0. - Otherwise, availableFlagSbCol is set equal to 1.

[0368] If availableFlagSbCol is equal to 1, the following applies: - The variables numSbX, numSbY, sbWidth, sbHeight, and refIdxLXSbCol are derived as follows: numSbX=cbWidth>>3 (8-546) numSbY=cbHeight>>3 (8-547) sbWidth=cbWidth / numSbX (8-548) sbHeight=cbHeight / numSbY (8-549) refIdxLXSbCol=0 (8-550) When xSbIdx=0..numSbX-1 and ySbIdx=0...numSbY-1, the motion vector mvLXSbCol[xSbIdx][ySbIdx] and the prediction list usage flag predFlagsLXSbCol[xSbIdx] are derived as follows: The luminance position (xSb, ySb) that defines the top-left sample of the current coding sub-block relative to the top-left luminance sample of the current picture is derived as follows: xSb=xCb+xSbIdx*sbWidth+sbWidth / 2 (8-551) ySb=yCb+ySbIdx*sbHeight+sbHeight / 2 (8-552) The position (xColSb, yColSb) of the collocated sub-block inside ColPic is derived as follows:

[0369] [ka]

[0370] The variable currCb defines the luma coding block that contains the current coding sub-block in the current picture. The variable colCb defines the luminance coding block containing the modified position given by ((xColSb>>3)<<3, (yColSb>>3)<<3) inside ColPic. - The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the collocated picture specified by ColPic. - The co-located motion vector derivation process specified in clause 8.5.2.12 is called with inputs currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1, and the output is assigned to the motion vector of sub-block mvL0SbCol[xSbIdx][ySbIdx] and availableFlagL0SbCol. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked with inputs currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1, and the output is assigned to the motion vector of sub-block MVL1SbCol[xSbIdx][ySbIdx] and availableFlagL1SbCol. - If availableFlagL0SbCol and availableFlagL1SbCol are both equal to 0, then if X is 0 and 1, the following applies: mvLXSbCol[xSbIdx][ySbIdx]=ctrMvLX (8-556) predFlagLXSbCol[xSbIdx][ySbIdx]=ctrPredFlagLX (8-557)

[0371] 8.5.5.4 Sub-block based temporal merge-based motion data derivation process

[0372] The inputs to this process are: - the position (xCtb, yCtb) of the top-left sample of the luma coding tree block containing the current coding block, - The location (xColCtrCb, yColCtrCb) of the top-left sample of the co-located luma coding block with the bottom-right center sample. - the availability flag of the neighboring coding unit availableFlagA1, - the reference index of the neighboring coding unit refIdxLXA1, - prediction list usage flag predFlagLXA1 for neighboring coding units; - Motion vectors of neighboring coding units at 1 / 16 fractional sample precision mvLXA1.

[0373] The output of this process is: - the motion vectors ctrMvL0 and ctrMvL1, - Prediction list usage flags ctrPredFlagL0, ctrPredFlagL1, - Temporal motion vector tempMv.

[0374] The variable tempMv is set as follows: tempMv[0]=0 (8-558) tempMv[1]=0 (8-559)

[0375] The variable currPic defines the current picture.

[0376] If availableFlagA1 is equal to TRUE, the following applies: - tempMv is set equal to mvL0A1 if all of the following conditions are true: - predFlagL0A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A1]) is equal to 0, - Otherwise, if all of the following conditions are true, then tempMv is set equal to mvL1A1: - Slice type is the same as B, - predFlagL1A1 is equal to 1, - DiffPicOrderCnt(ColPic,RefPicList[1][refIdxL1A1]) equals 0.

[0377] [ka]

[0378] The array colPredMode is set equal to the prediction mode array CuPredMode[0] of the collocated picture specified by ColPic.

[0379] The motion vectors ctrMvL0 and ctrMvL1 and the prediction list usage flags ctrPredFlagL0 and ctrPredFlagL1 are derived as follows. - If colPredMode[xColCb][yColCb] is equal to MODE_INTER, the following applies: The variable currCb defines the luma coding block containing (xCtrCb, yCtrCb) in the current picture. The variable colCb defines the luminance coding block containing the modified position given by ((xColCb>>3)<<3, (yColCb>>3)<<3) inside ColPic. - The luma position (xColCb, yColCb) is set equal to the top-left sample of the co-located luma coding block specified by colCb relative to the top-left luma sample of the collocated picture specified by ColPic. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked by taking as input currCb, colCb, (xColCb, yColCb), refIdxL0 set equal to 0, and sbFlag set equal to 1, and assigning the output to ctrMvL0 and ctrPredFlagL0. - The co-located motion vector derivation process specified in clause 8.5.2.12 is invoked by taking as input currCb, colCb, (xColCb, yColCb), refIdxL1 set equal to 0, and sbFlag set equal to 1, and assigning the output to ctrMvL1 and ctrPredFlagL1. - Otherwise, the following applies: ctrPredFlagL0=0 (8-563) ctrPredFlagL1=0 (8-564)

[0380] 8.5.6.3 Fractional Sample Interpolation

[0381] 8.5.6.3.1 General

[0382] The inputs to this process are: - a luminance position (xSb, ySb) defining the top left sample of the current coding sub-block relative to the top left luminance sample of the current picture, - variable sbWidth, which defines the width of the current coding sub-block; - variable sbHeight, which specifies the height of the current coding sub-block; - motion vector offset mvOffset, - fine-tuned motion vectors refMvLX, - the selected reference picture sample array refPicLX, - 1 / 2 sample interpolation filter index hpelIfIdx, - bidirectional optical flow flag bdofFlag, - A variable cIdx that specifies the color component index of the current block.

[0383] The output of this process is: - predSamplesLX, a (sbWidth+brdExtSize)x(sbHeight+brdExtSize) array of predicted sample values.

[0384] The prediction block boundary extension size brdExtSize is derived as follows. brdExtSize=(bdofFlag||(inter_affine_flag[xSb][ySb] && sps_affine_prof_enabled_flag))?2:0 (8-752)

[0385] The variable fRefWidth is set equal to the PicOutputWidthL of the reference picture in luma samples.

[0386] The variable fRefHeight is set equal to the PicOutputHeightL of the reference picture in luma samples.

[0387] The motion vector mvLX is set equal to (refMvLX-mvOffset). - If cIdx is equal to 0, the following applies: - The scaling factor and its fixed-point representation are defined as follows: hori_scale_fp=((fRefWidth<<14)+(PicOutputWidthL>>1)) / PicOutputWidthL (8-753) vert_scale_fp=((fRefHeight<<14)+(PicOutputHeightL>>1)) / PicOutputHeightL (8-754) - Let (xIntL, yIntL) be the luma position given in full sample units, and (xFracL, yFracL) be the offset given in 1 / 16 sample units. These variables are used in this section only to specify a fractional sample position within the reference sample array refPicLX. - Bounding block for reference sample padding (xSbInt L ,ySbInt L ) equal to (xSb+(mvLX[0]>>4),ySb+(mvLX[1]>>4)). - Predicted luminance sample array predSamples L Each luminance sample position in X (xL=0..sbWidth-1+brdExtSize,y L = 0..sbHeight-1+brdExtSize), the corresponding predicted luminance sample value predSamplesLX[x L ][y L ] is derived as follows: - (refxSb L ,refySb L ) and (refx L ,refy L ) is the luminance position indicated by the motion vector (refMvLX[0], refMvLX[1]) given in 1 / 16 sample units. L , refx L , refySb L , refy L is derived as follows: refxSb L =((xSb<<4)+refMvLX[0])*hori_scale_fp (8-755) refx L =((Sign(refxSb)*((Abs(refxSb)+128)>>8) +x L *((hori_scale_fp+8)>>4))+32)>>6 (8-756) refySb L =((ySb<<4)+refMvLX[1])*vert_scale_fp (8-757) refyL=((Sign(refySb)*((Abs(refySb)+128)>>8)+yL* ((vert_scale_fp+8)>>4))+32)>>6 (8-758) - variable xInt L , yInt L , xFrac L , and yFrac L is derived as follows: xInt L =refx L >>4 (8-759) yInt L =refy L >>4 (8-760) xFrac L =refx L &15 (8-761) yFrac L =refy L &15 (8-762)

[0388] [ka]

[0389] - If bdofFlag is equal to TRUE (sps_affine_prof_enabled_flag is equal to TRUE and inter_affine_flag[xSb][ySb] is equal to TRUE), and one or more of the following conditions are true, the predicted luminance sample values ​​predSamplesLX[x L ][y L ] is used as input (xInt L +(xFrac L >>3)-1),yInt L +(yFrac L >>3)-1) is derived by calling the luminance integer sample extraction process using refPicLX. 1.x L is equal to 0. 2.x L is equal to sbWidth+1. 3.y L is equal to 0. 4.y L is equal to sbHeight+1.

[0390] [ka]

[0391] - Otherwise (cIdx is not equal to 0), the following applies: - Let (xIntC,yIntC) be the chroma position given in full sample units, and (xFracC,yFracC) be the offset given in 1 / 32 sample units. These variables are used in this section only to specify the position of a general fractional sample within the reference sample array refPicLX. - The top-left coordinate of the bounding block (xSbIntC, ySbIntC) for the reference sample padding is set equal to ((xSb / SubWidthC)+(mvLX[0]>>5),(ySb / SubHeightC)+(mvLX[1]>>5)). For each chroma sample position (xC=0..sbWidth-1, yC=0..sbHeight-1) in the predicted chroma sample array predSamplesLX, the corresponding predicted chroma sample value predSamplesLX[xC][yC] is derived as follows: - (refxSb C ,refySb C ) and (refx C ,refy C ) is the chroma position indicated by the motion vector (mvLX[0], mvLX[1]) given in 1 / 32 sample units. C , refySb C , refx C , refy C is derived as follows: refxSb C =((xSb / SubWidthC<<5)+mvLX[0])*hori_scale_fp (8-763) refx C =((Sign(refxSb C )*((Abs(refxSb C )+256)>>9) +xC*((hori_scale_fp+8)>>4))+16)>>5 (8-764) refySb C =((ySb / SubHeightC<<5)+mvLX[1])*vert_scale_fp (8-765) refy C =((Sign(refySb C )*((Abs(refySb C )+256)>>9) +yC*((vert_scale_fp+8)>>4))+16)>>5 (8-766) - variable xInt C , yInt C , xFrac C , yFrac C is derived as follows: xInt C =refx C >>5 (8-767) yInt C =refy C >>5 (8-768) xFrac C =refy C &31 (8-769) yFrac C =refy C &31 (8-770) - The predicted sample values ​​predSamplesLX[xC][yC] are derived by invoking the process specified in 8.5.6.3.4 using (xIntC,yIntC), (xFracC,yFracC), (xSbIntC,ySbIntC), sbWidth, sbHeight, and refPicLX as input.

[0392] 8.5.6.3.2 Luminance Sample Interpolation Filtering Process

[0393] [ka]

[0394] The output of this process is the predicted luminance sample value predSampleLX L is.

[0395] The variables shift1, shift2, and shift3 are derived as follows: - Set variable shift1 to Min(4,BitDepth Y Set variable shift2 equal to 6, and set variable shift3 equal to Max(2,14-BitDepth Y ) - The variable picW is set equal to pic_width_in_luma_samples, and the variable picH is set equal to pic_height_in_luma_samples.

[0396] [ka]

[0397] For i=0..7, full sample units (xInt i ,yInt i ) is derived as follows: - If subpic_treated_as_pic_flag[SubPicIdx] is equal to 1, the following applies: xInt i =Clip3(SubPicLeftBoundaryPos,SubPicRightBoundaryPos,xInt L +i-3) (8-771) yInt i =Clip3(SubPicTopBoundaryPos,SubPicBotBoundaryPos,yInt L +i-3) (8-772) - Otherwise (subpic_treated_as_pic_flag[SubPicIdx] equals 0), the following applies: xInt i =Clip3(0,picW-1,sps_ref_wraparound_enabled_flag? ClipH((sps_ref_wraparound_offset_minus1+1)*MinCbSizeY,picW,xInt L +i-3): (8-773) xInt L +i-3) yInt i =Clip3(0,picH-1,yInt L +i-3) (8-774)

[0398] For i=0..7, the luma position in full sample units is further modified as follows: xInt i =Clip3(xSbInt L -3,xSbIntL+sbWidth+4,xInt i ) (8-775) yInt i =Clip3(ySbInt L -3,ySbInt L +sbHeight+4,yInt i ) (8-776)

[0399] Predicted luminance sample value predSampleLX L is derived as follows: - xFrac L and yFrac L If both are equal to 0, then predSampleLX L The value of is derived as follows: predSampleLX L =refPicLX L [xInt3][yInt3]< <shift3 (8-777) - No, xFrac L is not equal to 0 and yFrac L If is equal to 0, predSampleLX L The value of is derived as follows: predSampleLX L =(Σ 7 i=0 f L [xFrac L ][i]*refPicLX L [xInt i ][yInt3])>>shift1 (8-778) - No, xFrac L is equal to 0, and yFrac L If is not equal to 0, predSampleLX L The value of is derived as follows: predSampleLX L =(Σ 7 i=0 f L [yFrac L ][i]*refPicLX L [xInt3][yInt i ])>>shift1 (8-779) - No, xFrac L is not equal to 0 and yFrac L If is not equal to 0, predSampleLX L The value of is derived as follows: - The sample array temp[n] for n=0..7 is derived as follows: temp[n]=(Σ 7 i=0 f L [xFrac L ][i]*refPicLX L [xInt i ][yInt n ])>>shift1 (8-780) - Predicted luminance sample value predSampleLX L is derived as follows: predSampleLX L =(Σ 7 i=0 f L [yFrac L ][i]*temp [i])>>shift2 (8-781)

[0400] [Table 11]

[0401] [Table 12]

[0402] 30 is a flowchart of a video processing method 3000. The method 3000 includes, at step 3010, determining whether to add a maximum number of candidates (ML) in a sub-block-based merge candidate list and / or sub-block-based temporal motion vector prediction (SbTMVP) candidates to the sub-block-based merge candidate list for converting between a current block of video and a bitstream representation of the video based on whether temporal motion vector prediction (TMVP) is enabled for use during the conversion or whether the conversion uses a current picture reference (CPR) coding mode.

[0403] The method 3000 includes, at step 3020, performing the conversion based on the determining.

[0404] 31 is a flowchart of a video processing method 3100. The method 3100 includes, at step 3110, determining a maximum number of candidates (ML) in a sub-block-based merge candidate list based on whether temporal motion vector prediction (TMVP), sub-block-based temporal motion vector prediction (SbTMVP) tools, and affine coding mode are enabled for the conversion between a current block of video and a bitstream representation of the video.

[0405] The method 3100 includes, at step 3120, performing a transformation based on the determination.

[0406] 32 is a flowchart of a video processing method 3200. The method 3200 includes, at step 3210, determining that sub-block-based motion vector prediction (SbTMVP) mode conversion is disabled for converting between a current block of a first video segment of a video and a bitstream representation of the video because temporal motion vector prediction (TMVP) mode is disabled at the first video segment level.

[0407] The method 3200 includes, at step 3220, performing a conversion based on a determination that the bitstream representation conforms to a format that specifies whether an indication of SbTMVP mode is included and / or the position of the indication of SbTMVP mode relative to the indication of TMVP mode in the merge candidate list.

[0408] 33 is a flowchart of a video processing method 3300. The method 3300 includes, at step 3310, converting between a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool and a bitstream representation of the video, wherein coordinates of the current block or a corresponding position of the current block or a sub-block of the current block are selectively masked using a mask based on compression of motion vectors associated with the SbTMVP tool or the TMVP tool, and applying the mask includes a bitwise AND operation between values ​​of the coordinates and values ​​of the mask.

[0409] Figure 34 is a flowchart of a video processing method 3400. The method 3400 includes, at step 3410, determining, based on one or more characteristics of a current block of a video segment of a video, a valid corresponding region of a current block for applying a sub-block based motion vector prediction (SbTMVP) tool to the current block.

[0410] Based on this determination, the method 3400 includes, at step 3420, converting between the current block and a bitstream representation of the video.

[0411] Figure 35 is a flowchart of a video processing method 3500. The method 3500 includes, at step 3510, determining a default motion vector for a current block of video being coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool.

[0412] The method 3500 includes, at step 3520, converting between the current block and a bitstream representation of the video based on the determination, and determining a default motion vector if no motion vector is available from a block containing a corresponding position in a collocated picture associated with the center position of the current block.

[0413] 36 is a flowchart of a video processing method 3600. The method 3600 includes, at step 3610, inferring, for a current block of a video segment of a video, that a current picture of the current block is a reference picture with index M in a reference picture list X, where M and X are integers, and that a sub-block based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool is disabled if X=0 or X=1.

[0414] At step 3620, the method 3600 includes converting between the current block and a bitstream representation of the video based on the inference.

[0415] 37 is a flowchart of a video processing method 3700. The method 3700 includes, at step 3710, determining, for a current block of the video, that application of a sub-block-based temporal motion vector prediction (SbTMVP) tool is enabled if the current picture of the current block is a reference picture having an index set to M in a reference picture list X, where M and X are integers.

[0416] Based on this determination, the method 3700 includes converting between the current block and a bitstream representation of the video at step 3720 .

[0417] 38 is a flowchart of a video processing method 3800. The method 3800 includes, at step 3810, converting between a current block of video and a bitstream representation of the video, where the current block is coded using a sub-block-based coding tool, where converting encodes a sub-block merge index in a uniform manner using multiple bins (N) when a sub-block-based temporal motion vector prediction (SbTMVP) tool is enabled or disabled.

[0418] Figure 39 is a flowchart of a video processing method 3900. The method 3900 includes, at step 3910, determining, for a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool, a motion vector to define a corresponding block in a picture different from the current picture in which the SbTMVP tool includes the current block.

[0419] The method 3900 includes, at step 3920, converting between the current block and a bitstream representation of the video based on the determination.

[0420] 40 is a flowchart of a video processing method 4000. The method 4000 includes, at step 4010, determining whether to insert a zero motion affine merge candidate into a sub-block merge candidate list for converting between a current block of the video and a bitstream representation of the video based on whether affine prediction is enabled for converting the current block.

[0421] The method 4000 includes, at step 4020, performing the conversion based on the determination.

[0422] Figure 41 is a flowchart of a video processing method 4100. The method 4100 includes, at step 4110, inserting a zero motion non-affine padding candidate into a sub-block merging candidate list if the sub-block merging candidate list is not full for conversion between a current block of the video and a bitstream representation of the video using the sub-block merging candidate list.

[0423] The method 4100 includes, in step 4120, performing a transformation after the insertion.

[0424] 42 is a flowchart of a video processing method 4200. The method 4200 includes, at step 4210, determining a motion vector for conversion between a current block of video and a bitstream representation of the video using a rule that determines deriving the motion vector from one or more motion vectors of a block that includes a corresponding position in a co-located picture.

[0425] The method 4200 includes, at step 4220, performing a transformation based on the motion vector.

[0426] FIG. 43 is a block diagram of a video processing device 4300. The device 4300 may be used to implement one or more of the methods described herein. The device 4300 may be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 4300 may include one or more processing units 4302, one or more memories 4304, and video processing hardware 4306. The processing unit 4302 may be configured to implement one or more methods described herein. The memory(s) 4304 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 4306 may be used to implement the techniques described herein in a hardware circuit.

[0427] In some embodiments, the video encoding method may be performed using an apparatus implemented on a hardware platform, such as that described with reference to FIG.

[0428] Some embodiments of the disclosed technology include determining or deciding to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder uses or implements the tool or mode when processing a single video block, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the conversion from a block of video to a bitstream representation of video uses the video processing tool or mode if the video processing tool or mode is enabled based on the determination or decision. In another example, when a video processing tool or mode is enabled, a decoder processes the bitstream, recognizing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of video to a block of video is performed using the video processing tool or mode enabled based on the determination or decision.

[0429] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder does not use the tool or mode when converting blocks of video into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder processes the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was enabled based on the decision or determination.

[0430] FIG. 44 is a block diagram illustrating an example video processing system 4400 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the modules of system 4400. System 4400 may include an input unit 4402 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8- or 10-bit multi-module pixel values, or in a compressed or encoded format. Input unit 4402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, passive optical networks (PONs), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0431] The system 4400 may include an encoding module 4404 capable of implementing various encoding or coding methods described herein. The encoding module 4404 may reduce the average bit rate of the video from the input unit 4402 to the output of the encoding module 4404, generating a coded representation of the video. Accordingly, this encoding technique may be referred to as a video compression or video transcoding technique. The output of the encoding module 4404 may be stored or transmitted via a connected communication, as represented by module 4406. The bitstream (or coded) representation of the video received at the input unit 4402, stored, or communicated, may be used by module 4408 to generate pixel values ​​or displayable video that are transmitted to the display interface unit 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video unpacking. Furthermore, while certain video processing operations are referred to as “encoding” operations or tools, it will be understood that the encoding tools or operations are performed by an encoder and corresponding decoding tools or operations that reverse the results of the decoding are performed by a decoder.

[0432] Examples of peripheral bus interface units or display interface units may include Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be implemented in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0433] In some embodiments, the following technical solutions can be implemented.

[0434] A1. A video processing method comprising: determining whether to add a maximum number of candidates (ML) and / or temporal motion vector prediction (SbTMVP) in a sub-block based merge candidate list to a sub-block based merge candidate list for a conversion between a current block of a video and a bitstream representation of the video, based on whether temporal motion vector prediction (TMVP) is valid for use during the conversion or whether a current picture reference (CPR) coding mode is used for the conversion; and performing the conversion based on the determination.

[0435] A2. The method of Solution A1, wherein use of SbTMVP candidates is disabled based on a determination that TMVP tools are disabled or that SbTMVP tools are disabled.

[0436] A3. The method of solution A2, wherein determining ML includes excluding SbTMVP candidates from the sub-block-based merge candidate list based on whether the SbTMVP tool or TMVP tool is disabled.

[0437] A4. A video processing method including: determining a maximum number of candidates (ML) in a sub-block based merge candidate list based on whether a temporal motion vector prediction (TMVP), a sub-block based temporal motion vector prediction (SbTMVP) tool, and an affine coding mode are valid for use for a conversion between a current block of video and a bitstream representation of the video; and performing the conversion based on the determination.

[0438] A5. The method according to solution A4, wherein ML is set on the fly and signaled in the bitstream representation due to a determination that affine coding mode is enabled.

[0439] A6. The method of solution A4, wherein ML is predefined upon determining that affine coding mode is disabled.

[0440] A7. The method described in solution A2 or A6, wherein determining ML includes setting ML to 0 when it is determined that the TMVP tool is disabled, and the SbTMVP tool is enabled and the affine coding mode of the current block is disabled.

[0441] A8. The method described in solution A2 or A6, wherein determining ML includes setting ML to 1 when it is determined that the SbTMVP tool is enabled, and the TMVP tool is enabled and the affine coding mode of the current block is disabled.

[0442] A9. The method of solution A1, wherein the SbTMVP tool is disabled or the use of SbTMVP candidates is disabled by determining that the collocated reference picture of the current picture for the current block is the current picture.

[0443] A10. The method described in Solution A9, wherein determining ML includes excluding SbTMVP candidates from a sub-block-based merge candidate list based on whether the SbTMVP tool is disabled or whether the current picture's collocated reference picture is the current picture.

[0444] A11. The method described in solution A9, wherein determining ML includes setting ML to 0 based on determining that the collocated reference picture of the current picture is the current picture, and disabling affine coding of the current block.

[0445] A12. The method described in Solution A9, wherein determining ML includes setting ML to 1 if the SbTMVP tool is determined to be enabled, thereby causing the collocated reference picture of the current picture to be other than the current picture and disabling affine coding of the current block.

[0446] A13. The method described in Solution A1, in which the use of SbTMVP candidates is disabled by either disabling the SbTMVP tool or determining that the reference picture with reference picture index 0 in Reference Picture List 0 (L0) is the current picture for the current block.

[0447] A14. A method according to solution A13, wherein determining ML includes excluding SbTMVP candidates from the sub-block-based merge candidate list based on whether the SbTMVP tool is disabled or whether the reference picture having reference picture index 0 in L0 is the current picture.

[0448] A15. A method as described in solution A10 or A13, wherein determining ML includes setting ML to 0 when it is determined that the SbTMVP tool is enabled, the reference picture having reference picture index 0 in L0 is the current picture, and affine coding of the current block is disabled.

[0449] A16. The method described in A13, wherein determining ML includes setting ML to 1 by determining that the SbTMVP tool is enabled, the reference picture having reference picture index 0 in L0 is not the current picture, and affine coding of the current block is disabled.

[0450] A17. The method described in Solution A1, wherein the use of SbTMVP candidates is disabled due to a determination that the SbTMVP tool is disabled or the reference picture having reference picture index 0 in Reference Picture List 1 (L1) is the current picture of the current block.

[0451] A18. The method described in solution A17, wherein determining ML includes excluding SbTMVP candidates from the sub-block-based merge candidate list based on whether the SbTMVP tool is disabled or whether the reference picture with reference picture index 0 in L1 is the current picture.

[0452] A19. The method described in A17, wherein determining ML includes setting ML to 0 when it is determined that the SbTMVP tool is enabled, and disabling the reference picture in L1 whose reference picture index 0 is the current picture and affine coding of the current block.

[0453] A20. The method described in A17, wherein determining ML includes setting ML to 1 when it is determined that the SbTMVP tool is enabled, the reference picture with reference picture index 0 in L1 is not the current picture, and the affine coding of the current block is disabled.

[0454] A21. A video processing method comprising: determining, for a conversion between a current block of a first video segment of a video and a bitstream representation of the video, that one sub-block-based motion vector prediction (SbTMVP) mode is disabled for the conversion because temporal motion vector prediction (TMVP) mode is disabled at the first video segment level; and performing the conversion based on this determination, wherein the bitstream representation conforms to a format that specifies whether an indication of SbTMVP mode is included and / or the position of the indication of SbTMVP relative to the indication of TMVP mode in a merge candidate list.

[0455] A22. The method of solution A21, wherein the first video segment is a sequence, a slice, a tile, or a picture.

[0456] A23. The method of solution A21, wherein the format specifies omitting the indication of SbTMVP mode by including the indication of TMVP mode at the first video segment level.

[0457] A24. The method of solution A21, wherein the format specifies that the representation of SbTMVP mode is at the first video segment level in decoding order after the representation of TMVP mode.

[0458] A25. A method according to any one of solutions A21 to A24, wherein the format specifies that the display of SbTMVP mode is omitted because the TMVP mode is determined to be invalid.

[0459] A26. The method of solution A21, wherein the format specifies that the indication of SbTMVP mode is included at the video sequence level and omitted at the second video segment level.

[0460] A27. The method of solution A26, wherein the second video segment at the second video segment level is a slice, a tile, or a picture.

[0461] A28. The method according to any one of solutions A1 to A27, wherein the transformation generates the current block from a bitstream representation.

[0462] A29. A method according to any one of solutions A1 to A27, wherein the transformation generates a bitstream representation from the current block.

[0463] A30. A method according to any of Solutions A1-A27, wherein said converting comprises parsing said bitstream representation based on one or more decoding rules.

[0464] A31. An apparatus in a video system including a processing unit and a non-transitory memory loaded with instructions, the instructions executed by the processing unit causing the processing unit to implement a method described in any one of solutions A1 to A43.

[0465] A32. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for executing the method described in any one of solutions A1 to A43.

[0466] In some embodiments, the following technical solutions can be implemented.

[0467] B1. A video processing method including converting between a current block of video encoded using a sub-block-based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool and a bitstream representation of the video, selectively masking coordinates of positions corresponding to the current block or a sub-block of the current block using a mask based on compression of motion vectors associated with the SbTMVP tool or the TMVP tool, and applying the mask including a bitwise AND operation between the value of the coordinate and the value of the mask.

[0468] B2. The method of Solution B1, wherein the coordinates are (xN, yN), the mask (MASK) is an integer equal to ~(2M-1), where M is an integer, and applying the mask results in masked coordinates (xN', yN'), where xN'=xN&MASK, yN'=yN&MASK, and "~" is a bitwise NOT operation and "&" is a bitwise AND operation.

[0469] B3. The method according to solution B2, wherein M=3 or M=4.

[0470] B4. The method according to Solution B2 or B3, wherein multiple sub-blocks of size 2K×2K share the same motion information based on the compression of the motion vectors, and K is an integer not equal to M.

[0471] B5. The method according to solution B4, wherein M=K+1.

[0472] B6. The method of Solution 1, wherein the mask is not applied if it is determined that the SbTMVP tool or the motion vectors associated with the TMVP tool are not compressed.

[0473] B7. A method according to any one of solutions B1 to B6, wherein the mask for the SbTMVP tool is the same as the mask for the TMVP tool.

[0474] B8. The method according to any one of solutions B1 to B6, wherein the mask for the ATMVP tool is different from the mask for the TMVP tool.

[0475] B9. The method of solution B1, wherein one type of compression is no compression, 8x8 compression, or 16x16 compression.

[0476] B10. The method of solution B9, wherein the type of compression is signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, or a tile group header.

[0477] B11. The method of any one of Solutions B9 and B10, wherein the type of compression is based on a standard profile, level, or tier corresponding to the current block.

[0478] B12. A video processing method comprising: determining, based on one or more characteristics of a current block of a video segment of a video, a valid corresponding region of the current block for applying a sub-block based motion vector prediction (SbTMVP) tool based on the current block; and converting between the current block and a bitstream representation of the video based on the determination.

[0479] B13. The method of solution B12, wherein the one or more characteristics include a height or width of the current block.

[0480] B14. The method of solution B12, wherein the one or more characteristics include a type of compression of the motion vector associated with the current block.

[0481] B15. The method described in solution B14, wherein the valid corresponding area is a first size because the compression type is determined to not include compression, and the valid corresponding area is a second size larger than the first size because the compression type is determined to include KxK compression.

[0482] B16. The method according to solution B12, wherein the size of the valid corresponding region is based on a basic region of size M×N smaller than the size of the coding tree unit (CTU) region, and the size of the current block is W×H.

[0483] B17. The method described in solution B16, wherein the size of the CTU region is 128x128, M=64, and N=64.

[0484] B18. The method according to solution B16, wherein it is determined that W≦M and H≦N, so that the valid corresponding regions are the collocated basic region and extension in the collocated picture.

[0485] B19. A method according to solution B16, wherein when it is determined that W>M and H>N, the current block is divided into multiple parts, each of which includes a corresponding valid area for applying the SbTMVP tool.

[0486] B20. A video processing method comprising: determining a default motion vector for a current block of video coded using a sub-block based temporal motion vector prediction (SbTMVP) tool; and converting between the current block and a bitstream representation of the video based on the determination; wherein the default motion vector is determined because it is determined that a motion vector is not obtained from a block that includes a corresponding position in the collocated picture associated with the center position of the current block.

[0487] B21. The method of solution B20, wherein the default motion vector is set to (0,0).

[0488] B22. The method of solution B20, wherein the default motion vector is derived from a history-based motion vector prediction (HMVP) table.

[0489] B23. The method according to solution B22, wherein the default motion vector is set to (0,0) upon determining that the HMVP table is empty.

[0490] B24. The method described in solution B22, wherein the default motion vector is predefined and signaled to the video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, tile group header, coding tree unit (CTU), or coding unit (CU).

[0491] B25. The method of solution B22, setting the default motion vector to the first element stored in the HMVP table based on determining that the HMVP table is not empty.

[0492] B26. The method of solution B22, setting the default motion vector to the last element stored in the HMVP table based on determining that the HMVP table is not empty.

[0493] B27. The method of solution B22, setting the default motion vector to a particular motion vector stored in the HMVP table based on determining that the HMVP table is not empty.

[0494] B28. The method according to solution B27, wherein the specific motion vector references reference list 0.

[0495] B29. The method according to solution B27, wherein the specific motion vector refers to Reference List 1.

[0496] B30. The method of solution B27, wherein a particular motion vector references a particular reference picture in reference list 0.

[0497] B31. The method of Solution B27, wherein a particular motion vector references a particular reference picture in Reference List 1.

[0498] B32. The method according to solution B30 or B31, wherein the particular reference picture has index 0.

[0499] B33. The method of solution B27, wherein a particular motion vector references a collocated picture.

[0500] B34. The method of solution B22, wherein if the search process in the HMVP table determines that the specific motion vector is not found, the default motion vector is set to a predefined default motion vector.

[0501] B35. The method according to solution B34, wherein the search process searches only the first element or only the last element of the HMVP table.

[0502] B36. The method of solution B34, wherein the search process searches only a subset of the elements in the HMVP table.

[0503] B37. The method according to solution B22, wherein the default motion vector does not refer to the current picture for the current block.

[0504] B38. The method of solution B22, scaling a default motion vector to a collocated picture based on a determination that the default motion vector does not reference the collocated picture.

[0505] B39. The method of solution B20, wherein the default motion vector is derived from a neighboring block.

[0506] B40. The method described in solution B39, wherein the upper right corner of neighboring block (A0) is directly adjacent to the lower left corner of the current block, or the lower right corner of neighboring block (A1) is directly adjacent to the lower left corner of the current block, or the lower left corner of neighboring block (B0) is directly adjacent to the upper right corner of the current block, or the lower right corner of neighboring block (B1) is directly adjacent to the upper right corner of the current block, or the lower right corner of neighboring block (B2) is directly adjacent to the upper left corner of the current block.

[0507] B41. The method of Solution B40, wherein the default motion vector is derived from only one of the neighboring blocks A0, A1, B0, B1, B2.

[0508] B42. The method of Solution B40, wherein the default motion vector is derived from one or more of neighboring blocks A0, A1, B0, B1, B2.

[0509] B43. The method according to Solution B40, wherein if it is determined that no valid default motion vector is found for any of the neighboring blocks A0, A1, B0, B1, B2, the default motion vector is set to a predefined default motion vector.

[0510] B44. The method described in solution B43, wherein the predefined default motion vector is signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a coding tree unit (CTU), or a coding unit (CU).

[0511] B45. The method of solution B43 or B44, wherein the predefined default motion vector is (0,0).

[0512] B46. The method of Solution B39, wherein the default motion vector is set to a specific motion vector from a neighboring block.

[0513] B47. The method according to solution B46, wherein the specific motion vector references reference list 0.

[0514] B48. The method according to solution B46, wherein the specific motion vector refers to Reference List 1.

[0515] B49. The method of Solution B46, wherein a particular motion vector references a particular reference picture in Reference List 0.

[0516] B50. The method of Solution B46, wherein a particular motion vector references a particular reference picture in Reference List 1.

[0517] B51. The method of solution B49 or B50, wherein the particular reference picture has index 0.

[0518] B52. The method of solution B46, wherein a particular motion vector references a co-located picture.

[0519] B53. The method of Solution B20, wherein a default motion vector is used because the block containing the corresponding position in the co-located picture is determined to be intra-coded.

[0520] B54. The method of solution B20, wherein the derivation method is modified by determining that the block containing the corresponding position in the collocated picture is not located.

[0521] B55. The method of solution B20, wherein a default motion vector candidate is always available.

[0522] B56. The method of solution B20, wherein if it is determined that the default motion vector candidate is set to unavailable, an alternative default motion vector is derived.

[0523] B57. The method of solution B20, wherein the availability of the default motion vector is based on syntactic information in a bitstream representation associated with a video segment.

[0524] B58. The method of solution B57, wherein the syntax information includes an indication that the SbTMVP tool is enabled, and the video segment is a slice, a tile, or a picture.

[0525] B59. The method according to Solution B58, wherein the current picture of the current block is not an Intra Random Access Point (IRAP) reference index picture, and the current picture is not inserted into Reference Picture List 0 (L0) with reference index 0.

[0526] B60. A method according to solution B20, in which, upon determining that the SbTMVP tool is valid, a fixed index or set of fixed indexes is assigned to a candidate associated with the SbTMVP tool, and, upon determining that the SbTMVP tool is invalid, a fixed index or set of fixed indexes is assigned to a candidate associated with a coding tool other than the SbTMVP tool.

[0527] B61. A video processing method comprising: inferring, for a current block of a video segment of a video, that a sub-block based temporal motion vector prediction (SbTMVP) tool or a temporal motion vector prediction (TMVP) tool is disabled for the video segment based on determining that a current picture of the current block is a reference picture having an index set to M in a reference picture list X, where M and X are integers and X=0 or X=1; and converting between the current block and a bitstream representation of the video based on the inference.

[0528] B62. The method described in solution B61, wherein for the SbTMVP tool or a reference picture list X for the TMVP tool, M corresponds to a target reference picture index for scaling the motion information of the temporal block.

[0529] B63. The method of solution B61, wherein the current picture is an Intra Random Access Point (IRAP) picture.

[0530] B64. A video processing method comprising: determining, for a current block of a video, that application of a sub-block based temporal motion vector prediction (SbTMVP) tool is valid by determining that the current picture of the current block is a reference picture having an index set to M in a reference picture list X, where M and X are integers; and converting between the current block and a bitstream representation of the video based on the determination.

[0531] B65. The method of solution B64, wherein the motion information corresponding to each sub-block of the current block references the current picture.

[0532] B66. The method described in solution B64, wherein motion information for sub-blocks of the current block is derived from one temporal block, and this temporal block is coded with at least one reference picture that references the current picture of this temporal block.

[0533] B67. The method of solution B66, wherein the transformation excludes a scaling operation.

[0534] B68. A video processing method including converting between a current block of video and a bitstream representation of the video, the current block being coded using a sub-block based coding tool, and the converting including coding a sub-block merge index in a uniform manner using a plurality of bins (N) based on determining that the sub-block based temporal motion vector prediction (SbTMVP) tool is enabled or disabled.

[0535] B69. The method of solution B68, wherein a first number of bins (L) of the plurality of bins are context coded and a second number of bins (NL) are bypass coded.

[0536] B70. The method of solution B69, wherein L=1.

[0537] B71. The method of solution B68, wherein each of the plurality of bins is context coded.

[0538] B72. The method according to any one of solutions B1 to B71, wherein the transformation generates the current block from a bitstream representation.

[0539] B73. The method according to any one of solutions B1 to B71, wherein the transformation generates a bitstream representation from the current block.

[0540] B74. The method of any of Solutions B1-B71, wherein performing the transformation includes parsing the bitstream representation based on one or more decoding rules.

[0541] B75. An apparatus in a video system, comprising a processing unit and a non-transitory memory loaded with instructions, the instructions executed by the processing unit causing the processing unit to implement a method described in any one of solutions B1 to B11.

[0542] B76. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for performing the method according to any one of solutions B1 to B11.

[0543] In some embodiments, the following technical solutions can be implemented.

[0544] C1. A video processing method including: determining, for a current block of video coded using a sub-block-based temporal motion vector prediction (SbTMVP) tool, a motion vector that the SbTMVP tool uses to locate a corresponding block in a picture different from the current picture containing the current block; and converting between the current block and a bitstream representation of the video based on the determination.

[0545] C2. The method of Solution 1, wherein the motion vector is set to a default motion vector.

[0546] C3. The method according to solution C2, wherein the default motion vector is (0,0).

[0547] C4. The method of step C2, wherein the default motion vector is signaled in a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), a slice header, a tile group header, a coding tree unit (CTU), or a coding unit (CU).

[0548] C5. The method of solution C1, wherein the motion vector is set to a motion vector stored in a history-based motion vector prediction (HMVP) table.

[0549] C6. The method of solution C5, setting the motion vector to a default motion vector based on determining that the HMVP table is empty.

[0550] C7. The method according to solution C6, wherein the default motion vector is (0,0).

[0551] C8. The method of solution C5, setting the motion vector to the first motion vector stored in the HMVP table based on determining that the HMVP table is not empty.

[0552] C9. The method of solution C5, wherein the motion vector is set to the last motion vector stored in the HMVP table upon determining that the HMVP table is not empty.

[0553] C10. The method of solution C5, setting the motion vector to a particular motion vector stored in the HMVP table based on determining that the HMVP table is not empty.

[0554] C11. The method according to solution C10, wherein the particular motion vector references reference list 0.

[0555] C12. The method according to solution C10, wherein the specific motion vector refers to Reference List 1.

[0556] C13. The method of solution C10, wherein a particular motion vector references a particular reference picture in reference list 0.

[0557] C14. The method of solution C10, wherein a particular motion vector references a particular reference picture in reference list 1.

[0558] C15. The method according to solution C13 or C14, wherein the particular reference picture has index 0.

[0559] C16. The method of solution C10, wherein the particular motion vector references a collocated picture.

[0560] C17. The method of solution C5, wherein if the search process in the HMVP table determines that the specific motion vector is not found, the motion vector is set to a default motion vector.

[0561] C18. The method according to solution C17, wherein the search process searches only the first element or only the last element of the HMVP table.

[0562] C19. The method of solution C17, wherein the search process searches only a subset of the elements in the HMVP table.

[0563] C20. The method according to solution C5, wherein the motion vectors stored in the HMVP table do not refer to the current picture.

[0564] C21. The method of solution C5, wherein the motion vector stored in the HMVP table is scaled to the collocated picture because it is determined that the motion vector stored in the HMVP table does not reference the collocated picture.

[0565] C22. The method of solution C1, wherein the motion vector is set to a specific motion vector of a specific neighboring block.

[0566] C23. The method according to solution C22, wherein the upper right corner of a particular neighboring block (A0) is directly adjacent to the lower left corner of the current block, or the lower right corner of a particular neighboring block (A1) is directly adjacent to the lower left corner of the current block, or the lower left corner of a particular neighboring block (B0) is directly adjacent to the upper right corner of the current block, or the lower right corner of a particular neighboring block (B1) is directly adjacent to the upper right corner of the current block, or the lower right corner of a particular neighboring block (B2) is directly adjacent to the upper left corner of the current block, or is directly adjacent to the upper left corner of the current block.

[0567] C24. The method of solution C1, wherein the absence of a particular neighboring block determines that the motion vector is set to a default motion vector.

[0568] C25. The method of solution C1, wherein a motion vector is set to a default motion vector upon determining that a particular neighboring block is not inter-coded.

[0569] C26. The method of solution C22, wherein the specific motion vector references reference list 0.

[0570] C27. The method according to solution C22, wherein the specific motion vector refers to Reference List 1.

[0571] C28. The method of solution C22, wherein a particular motion vector references a particular reference picture in reference list 0.

[0572] C29. The method of solution C22, wherein a particular motion vector references a particular reference picture in reference list 1.

[0573] C30. The method according to solution C28 or C29, wherein the specific reference picture has index 0.

[0574] C31. The method of solution C22 or C23, wherein the particular motion vector references a collocated picture.

[0575] C32. The method of solution C22 or C23, wherein a motion vector is set to a default motion vector upon determining that a particular neighboring block does not reference a collocated picture.

[0576] C33. The method according to any one of solutions C24 to C32, wherein the default motion vector is (0,0).

[0577] C34. The method of solution C1, wherein if it is determined that a particular motion vector stored for a particular neighboring block cannot be found, the motion vector is set to a default motion vector.

[0578] C35. The method of solution C22, wherein a particular motion vector is scaled to one collocated picture by determining that the particular motion vector does not reference the collocated picture.

[0579] C36. The method of solution C22, wherein the particular motion vector does not refer to the current picture.

[0580] C37. The method according to any one of solutions C1 to C36, wherein the transformation generates the current block from a bitstream representation.

[0581] C38. The method according to any one of solutions C1 to C36, wherein the transformation generates a bitstream representation from the current block.

[0582] C39. The method of any of Solutions C1-C36, wherein performing the conversion includes parsing the bitstream representation based on one or more decoding rules.

[0583] C40. An apparatus in a video system including a processing unit and a non-transitory memory loaded with instructions, the instructions executed by the processing unit causing the processing unit to implement a method described in any one of solutions C1 to C25.

[0584] C41. A computer program product stored on a non-transitory computer-readable medium, the computer program product comprising program code for executing the method according to any one of solutions C1 to C25.

[0585] In some embodiments, the following technical solutions can be implemented.

[0586] D1. A video processing method including: determining whether to insert a zero motion affine merge candidate into a sub-block merge candidate list for a transformation between a current block of video and a bitstream representation of the video based on whether affine prediction is enabled for the transformation of the current block; and performing the transformation based on the determination.

[0587] D2. The method of solution D1, wherein the zero motion affine merge candidate is not inserted into the sub-block merge candidate list because the affine use flag in the bitstream representation is determined to be off.

[0588] D3. The method of solution D2, further comprising inserting a default motion vector candidate that is a non-affine candidate into the sub-block merging candidate list upon determining that the affine use flag is off.

[0589] D4. A video processing method comprising: inserting a zero motion non-affine padding candidate into a sub-block merging candidate list when it is determined that the sub-block merging candidate list is not filled for conversion between a current block of the video and a bitstream representation of the video using the sub-block merging candidate list; and performing the conversion following the insertion.

[0590] D5. The method of solution D4, further comprising setting the affine usage flag of the current block to 0.

[0591] D6. The method of solution D4, wherein the inserting step is further based on whether an affine usage flag in the bitstream representation is off.

[0592] D7. A video processing method including determining a motion vector for conversion between a current block of video and a bitstream representation of the video using a rule that determines that the motion vector is derived from one or more motion vectors of blocks containing corresponding positions in a co-located picture, and performing the conversion based on the motion vector.

[0593] D8. The method described in solution D7, wherein the one or more motion vectors comprise MV0 and MV1 representing motion vectors in reference list 0 and reference list 1, respectively, and the motion vector to be derived comprises MV0' and MV1' representing motion vectors in reference list 0 and reference list 1.

[0594] D9. The method according to solution D8, wherein MV0′ and MV1′ are derived based on MV0 by determining that one collocated picture is in reference list 0.

[0595] D10. The method of solution D8, wherein MV0′ and MV1′ are derived based on MV1 based on determining that one collocated picture is in reference list 1.

[0596] D11. The method according to any of solutions D1 to D10, wherein the transformation generates the current block from a bitstream representation.

[0597] D12. The method according to any one of solutions D1 to D16, wherein the transformation generates a bitstream representation from the current block.

[0598] D13. The method of any of solutions D1-D10, wherein performing the conversion includes parsing the bitstream representation based on one or more decoding rules.

[0599] D14. An apparatus in a video system including a processing unit and a non-transitory memory loaded with instructions, wherein the instructions executed by the processing unit cause the processing unit to implement a method described in any one of solutions D1 to D18.

[0600] D15. A computer program product stored on a non-transitory computer-readable medium, the computer program product including program code for performing the method according to any one of solutions D1-D18.

[0601] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein, including the structures disclosed herein and their structural equivalents, may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, a programmable processing apparatus, a computer, or multiple processing apparatuses or computers. In addition to hardware, the apparatus may include code that creates an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiving device.

[0602] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), may be stored in a single file dedicated to the program, or may be stored in multiple coordinating files (e.g., files containing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer located at a single site or on multiple computers distributed across multiple sites and interconnected by a communications network.

[0603] The processes and logic flows described herein may be performed by one or more programmable processing devices executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and devices may be implemented as, special purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0604] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will include one or more mass storage devices, e.g., magnetic, magneto-optical, or optical disks, for storing data, or will be operatively coupled to receive data from or transfer data to these mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, EPROMs, EEPROMs, flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and semiconductor memory devices such as CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0605] While this patent specification contains many details, these should not be construed as limiting the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent specification in the context of separate embodiments may also be implemented in combination in a single example. Conversely, various features described in the context of a single example may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be extracted from the combination, and the claimed combination may be directed to a subcombination or variations of the subcombination.

[0606] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve desired results. Also, the separation of various system modules in the embodiments described in this patent specification should not be understood as requiring such separation in all embodiments.

[0607] Only a few implementations and examples have been described; other embodiments, extensions and variations are possible based on what is described and illustrated in this patent specification.

Claims

1. 1. A method for coding video data, comprising: a first initialization of temporal motion information to default motion information for a conversion between a current block of a video coded using sub-block-based temporal motion vector prediction and a bitstream of the video, the motion vector associated with the default motion information being (0, 0); Determining whether a reference picture of a particular neighboring block A1 is a collocated picture of the current block that is different from the current picture that constitutes the current block; If the reference picture of the specific neighboring block A1 is the collocated picture of the current block, setting the temporal motion information to specific motion information related to the specific neighboring block A1 without checking other neighboring blocks; if the reference picture of the particular neighboring block A1 is not the collocated picture of the current block, the temporal motion information remains equal to the default motion information; the specific neighboring block A1 is adjacent to the current block at the lower left corner, the specific neighboring block A1 covers a luminance position (xCb-1, yCb+cbHeight-1), where (xCb, yCb) is the luminance position of the top-left sample of the current block relative to the top-left luminance sample of the current picture, and cbHeight is the height of the current block; generating rounded temporal motion information by rounding the temporal motion information to integer precision; identifying, in the co-located picture, at least one video region of the current block based on the rounded temporal motion information; deriving at least one sub-block motion information associated with the current block based on the at least one image region; constructing a sub-block motion candidate list based on the at least one sub-block motion information; determining a candidate from the sub-block motion candidate list based on a sub-block merge index included in the bitstream; performing the conversion based on the determined candidates; Including, Deriving at least one sub-block motion information associated with the current block based on the at least one image region includes: determining whether the at least one video region is coded in inter mode; deriving the at least one sub-block motion information using the at least one video region if the at least one video region is coded in the inter mode; and refraining from using the at least one video region to derive the at least one sub-block motion information if the at least one video region is coded in an intra block copy mode. method.

2. a number of bins (N) are used to encode the sub-block merge index; The method of claim 1.

3. a first number (L) of bins of the plurality of bins are context coded and a second number (N-L) of bins are bypass coded; The method of claim 2.

4. L=1, The method of claim 3.

5. the converting includes decoding the current block from the bitstream.

5. The method according to any one of claims 1 to 4.

6. the transforming includes encoding the current block into the bitstream.

5. The method according to any one of claims 1 to 4.

7. the transforming includes generating the bitstream from the current block; The method further includes storing the bitstream on a non-transitory computer-readable recording medium.

5. The method according to any one of claims 1 to 4.

8. 1. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions stored thereon, The instructions, when executed by the processing unit, cause the processing unit to: a first initialization of temporal motion information to default motion information for a conversion between a current block of a video coded using sub-block-based temporal motion vector prediction and a bitstream of the video, the motion vector associated with the default motion information being (0, 0); Determining whether a reference picture of a particular neighboring block A1 is a collocated picture of the current block that is different from the current picture that constitutes the current block; If the reference picture of the specific neighboring block A1 is the collocated picture of the current block, setting the temporal motion information to specific motion information related to the specific neighboring block A1 without checking other neighboring blocks; if the reference picture of the particular neighboring block A1 is not the collocated picture of the current block, the temporal motion information remains equal to the default motion information; the specific neighboring block A1 is adjacent to the current block at the lower left corner, the specific neighboring block A1 covers a luminance position (xCb-1, yCb+cbHeight-1), where (xCb, yCb) is the luminance position of the top-left sample of the current block relative to the top-left luminance sample of the current picture, and cbHeight is the height of the current block; generating rounded temporal motion information by rounding the temporal motion information to integer precision; identifying, in the co-located picture, at least one video region of the current block based on the rounded temporal motion information; deriving at least one sub-block motion information associated with the current block based on the at least one image region; constructing a sub-block motion candidate list based on the at least one sub-block motion information; determining a candidate from the sub-block motion candidate list based on a sub-block merge index included in the bitstream; performing the conversion based on the determined candidates; Execute Deriving at least one sub-block motion information associated with the current block based on the at least one image region includes: determining whether the at least one video region is coded in inter mode; deriving the at least one sub-block motion information using the at least one video region if the at least one video region is coded in the inter mode; and refraining from using the at least one video region to derive the at least one sub-block motion information if the at least one video region is coded in an intra block copy mode. Device.

9. 1. A method for storing a video bitstream, comprising: a first initialization of temporal motion information for a current block of video coded using sub-block-based temporal motion vector prediction to default motion information, wherein a motion vector associated with the default motion information is (0,0); Determining whether a reference picture of a particular neighboring block A1 is a collocated picture of the current block that is different from the current picture that constitutes the current block; If the reference picture of the specific neighboring block A1 is the collocated picture of the current block, setting the temporal motion information to specific motion information related to the specific neighboring block A1 without checking other neighboring blocks; if the reference picture of the particular neighboring block A1 is not the collocated picture of the current block, the temporal motion information remains equal to the default motion information; the specific neighboring block A1 is adjacent to the current block at the lower left corner, the specific neighboring block A1 covers a luminance position (xCb-1, yCb+cbHeight-1), where (xCb, yCb) is the luminance position of the top-left sample of the current block relative to the top-left luminance sample of the current picture, and cbHeight is the height of the current block; generating rounded temporal motion information by rounding the temporal motion information to integer precision; identifying, in the co-located picture, at least one video region of the current block based on the rounded temporal motion information; deriving at least one sub-block motion information associated with the current block based on the at least one image region; constructing a sub-block motion candidate list based on the at least one sub-block motion information; determining a candidate from the sub-block motion candidate list based on a sub-block merge index included in the bitstream; generating the bitstream based on the determined candidates; storing the bitstream on a non-transitory computer-readable recording medium; Including, Deriving at least one sub-block motion information associated with the current block based on the at least one image region includes: determining whether the at least one video region is coded in inter mode; deriving the at least one sub-block motion information using the at least one video region if the at least one video region is coded in the inter mode; and refraining from using the at least one video region to derive the at least one sub-block motion information if the at least one video region is coded in an intra block copy mode. method.

10. A non-transitory computer-readable storage medium storing instructions, comprising: The instructions may include: a first initialization of temporal motion information to default motion information for a conversion between a current block of a video coded using sub-block-based temporal motion vector prediction and a bitstream of the video, the motion vector associated with the default motion information being (0, 0); Determining whether a reference picture of a particular neighboring block A1 is a collocated picture of the current block that is different from the current picture that constitutes the current block; If the reference picture of the specific neighboring block A1 is the collocated picture of the current block, setting the temporal motion information to specific motion information related to the specific neighboring block A1 without checking other neighboring blocks; if the reference picture of the particular neighboring block A1 is not the collocated picture of the current block, the temporal motion information remains equal to the default motion information; the specific neighboring block A1 is adjacent to the current block at the lower left corner, the specific neighboring block A1 covers a luminance position (xCb-1, yCb+cbHeight-1), where (xCb, yCb) is the luminance position of the top-left sample of the current block relative to the top-left luminance sample of the current picture, and cbHeight is the height of the current block; generating rounded temporal motion information by rounding the temporal motion information to integer precision; identifying, in the co-located picture, at least one video region of the current block based on the rounded temporal motion information; deriving at least one sub-block motion information associated with the current block based on the at least one image region; constructing a sub-block motion candidate list based on the at least one sub-block motion information; determining a candidate from the sub-block motion candidate list based on a sub-block merge index included in the bitstream; performing the conversion based on the determined candidates; Execute Deriving at least one sub-block motion information associated with the current block based on the at least one image region includes: determining whether the at least one video region is coded in inter mode; deriving the at least one sub-block motion information using the at least one video region if the at least one video region is coded in the inter mode; and refraining from using the at least one video region to derive the at least one sub-block motion information if the at least one video region is coded in an intra block copy mode. A non-transitory computer-readable storage medium.