Conditional Execution of the Motion Candidate List Construction Process
By employing non-rectangular partitioning modes and advanced motion vector prediction techniques, the challenges of increasing bandwidth demands in video coding are addressed, resulting in improved compression efficiency and video quality.
Patent Information
- Application Number
- JP2023182535
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-04
- Filing Date
- 2023-10-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-06-04
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding digital video due to increasing bandwidth demands and the need for improved compression techniques.
The implementation of non-rectangular partitioning modes, such as triangular partitioning, in video decoders and encoders, along with unified pruning processes and advanced motion vector prediction techniques, to enhance encoding efficiency and adapt to varying video block dimensions.
This approach allows for improved compression efficiency, reduced bandwidth requirements, and enhanced video quality by effectively utilizing non-rectangular prediction partitions and advanced motion vector prediction methods.
Smart Images

Figure 0007683985000024 
Figure 0007683985000025 
Figure 0007683985000026
Abstract
Description
Technical Field
[0001] This application is based on the national patent application No. 2021-571916 transferred domestically on December 3, 2021, which is based on the international patent application No. PCT / CN2020 / 094306 filed on June 4, 2020. The international patent application claims the priority and benefits of the international patent application No. PCT / CN2019 / 089970 filed on June 4, 2019. All of the above patent applications are incorporated herein by reference in their entirety.
[0002] This document relates to video and image encoding and decoding technologies.
Background Art
[0003] Digital video occupies the largest bandwidth usage in the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, it is expected that the bandwidth demand for digital video usage will continue to increase.
Summary of the Invention
[0004] The disclosed technology can be used by embodiments of a video or image decoder or encoder to perform encoding or decoding of a video bitstream using non-rectangular partitioning such as, for example, triangular partitioning mode.
[0005] In one exemplary embodiment, a method for visual media processing is disclosed. The method includes determining that a first video block of visual media data uses a Geometric Partitioning Mode (GPM) and a second video block of visual media data uses a non-GPM mode; constructing, based on a unified pruning process, a first merge list for the first video block and a second merge list for the second video block, wherein the first merge list and the second merge list have merge candidates, and the unified pruning process includes adding a new merge candidate to the merge list based on a comparison between motion information of the new merge candidate and motion information of at least one merge candidate in the merge list; and the GPM includes dividing the first video block into a plurality of prediction partitions to which motion prediction is separately applied, and at least one partition has a non-rectangular shape.
[0006] In another exemplary embodiment, another method for visual media processing is disclosed. The method includes determining, based on a rule, an initial value of motion information applicable to a current video block of visual media data for conversion between the current video block and a bitstream representation of the visual media data, the rule defining inspecting whether a Sub-block based Temporal Motion Vector Predictor (SbTMVP) mode is available for the current video block based on a reference list (denoted as list X) of adjacent blocks of the current video block, where X is an integer and the value of X depends at least on encoding conditions of the adjacent blocks; and performing the conversion based on the determination.
[0007] In yet another further exemplary aspect, another method for visual media processing is disclosed. The method includes: for the conversion between the current video block of visual media data and the bitstream representation of the visual media data, a step of derivating one or more collocated motion vectors for one or more sub-blocks of the current video block based on rules, wherein the rules define using a unified derivation process for derivating the one or more collocated motion vectors regardless of the encoding tools used to encode the current video block into the bitstream representation; and a step of performing the conversion using a merge list having the one or more collocated motion vectors.
[0008] In yet another further exemplary aspect, another method for visual media processing is disclosed. The method includes: a step of identifying one or more conditions related to the dimensions of the current video block of visual media data, wherein an intra-block copy (IBC) mode is applied to the current video block; a step of determining a motion candidate list construction process for the current video block based on whether the one or more conditions related to the dimensions of the current video block are satisfied; and a step of performing a conversion between the current video block and the bitstream representation of the current video block based on the motion candidate list.
[0009] In yet another further exemplary aspect, another method for visual media processing is disclosed. The method includes: for the conversion between the current video block of visual media data and the bitstream representation of the visual media data, a step of determining that an encoding technique is disabled for the conversion, wherein the bitstream representation is configured to include a field indicating that the maximum number of merge candidates for the encoding technique is zero; and a step of performing the conversion based on the determination that the encoding technique is disabled.
[0010] In yet another exemplary aspect, another method for visual media processing is disclosed. The method includes determining, using rules, for conversion between a current video block and a bitstream representation of the current video block, the rules specifying that a first syntax element in the bitstream representation is conditionally included based on a second syntax element in the bitstream representation that indicates a maximum number of merge candidates related to at least one coding technique applied to the current video block; and performing the conversion between the current video block and the bitstream representation of the current video block based on the determining.
[0011] In another exemplary aspect, the above method may be implemented by a video decoder device having a processor.
[0012] In another exemplary aspect, the above method may be implemented by a video encoder device having a processor.
[0013] In yet another exemplary aspect, these methods may be embodied in the form of processor-executable instructions and stored in a computer-readable program medium.
[0014] These and other aspects are further described in this document.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
DETAILED DESCRIPTION OF THE INVENTION
[0016] This document provides various techniques that can be used by a decoder of an image or video bitstream to improve the quality of a decompressed or decoded digital video or image. For the sake of brevity, the term "video" is used herein to include both a sequence of pictures (traditionally called video) and individual images. Also, a video encoder may implement these techniques in the encoding process to reconstruct decoded frames used for further encoding.
[0017] In this document, section headings are used for ease of understanding, but they do not limit the embodiments and techniques to the corresponding sections. Thus, embodiments from one section can be combined with embodiments from other sections.
[0018] 1. Overview This document relates to video coding technology. Specifically, it relates to merge coding including triangular prediction mode. This may be applied to existing video coding standards such as HEVC, or to emerging standards (Versatile Video Coding). This is also applicable to future video coding standards or video codecs.
[0019] 2. Introduction Video coding standards have mainly evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, ISO / IEC created MPEG-1 and MPEG-4 Visual, and the two organizations jointly created H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes transform coding in addition to temporal prediction. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the JVET (Joint Video Exploration Team) in 2015. Since then, numerous new methods have been adopted by the JVET and incorporated into the reference software named the Joint Exploration Model (JEM). In April 2018, the JVET (Joint Video Expert Team) was launched between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard aiming for a 50% bitrate reduction compared to HEVC.
[0020] The latest version of the VVC draft, namely, Versatile Video Coding (draft 5) is: http: / / phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N1001-v7.zip and can be found at.
[0021] The latest reference software for VVC, called VTM, can be found at: https: / / vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0 and can be found at
[0022] 2.1 Inter Prediction in HEVC / H.265 For an encoded coding unit (CU) to be inter-coded, it can be encoded with one prediction unit (PU) or two PUs according to the partition mode. Each PU by inter prediction has motion parameters regarding one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists can also be signaled using inter_pred_idc. The motion vector can be explicitly encoded as a delta relative to the predictor.
[0023] When a CU is encoded in skip mode, one PU is associated with that CU, there are no significant residual coefficients, and there are no encoded motion vector deltas or reference picture indices. The merge mode is defined, according to which the motion parameters for the current PU are obtained from adjacent PUs, including spatial and temporal candidates. The merge mode can be applied not only to the skip mode but also to any inter-predicted PU. An alternative to the merge mode is the explicit transmission of the motion parameters, in which the motion vector (more precisely, the motion vector difference (MVD) compared to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the reference picture list usage are explicitly signaled for each PU. Such a mode is referred to in this disclosure as Advanced motion vector prediction (AMVP).
[0024] When signaling indicates that one of the two reference picture lists is used, the PU is generated from samples of one block. This is called "uni-prediction". Uni-prediction is available for both P slices and B slices.
[0025] When signaling indicates that both of those reference picture lists are used, the PU is generated from samples of two blocks. This is called "bi-prediction". Bi-prediction is available only for B slices.
[0026] The following text provides details about the inter prediction modes defined in HEVC. The description will start from the merge mode.
[0027] 2.1.1 Reference Picture Lists In HEVC, the term inter prediction is used to denote a prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the currently decoded picture. Similar to H.264 / AVC, a picture can be predicted from multiple reference pictures. The reference pictures used for inter prediction are organized into one or more reference picture lists. A reference index specifies which of the reference pictures in the list should be used to create the prediction signal.
[0028] For P slices, a single reference picture list called list 0 is used, and for B slices, two reference picture lists called list 0 and list 1 are used. Note that the reference pictures included in list 0 / 1 can be from past and future pictures with respect to the capture / display order.
[0029] 2.1.2 Merge Mode 2.1.2.1 Derivation of Candidates for Merge Mode When PU is predicted using the merge mode, an index pointing to an entry in the merge candidate list is syntax-analyzed from the bitstream and used to extract motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following series of steps: · Step 1: Initial candidate derivation - Step 1.1: Spatial candidate derivation - Step 1.2: Redundancy check for spatial candidates - Step 1.3: Temporal candidate derivation · Step 2: Additional candidate insertion - Step 2.1: Creation of bi-prediction candidates - Step 2.2: Insertion of zero motion candidates.
[0030] Figure 1 also schematically shows these steps. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since a fixed number of candidates are assumed for each PU at the decoder, if the number of candidates obtained from Step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header, additional candidates are generated. Since the number of candidates is fixed, the index of the best merge candidate is encoded using truncated unary binarization (TU). When the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list that is the same as the merge candidate list of the 2N×2N prediction unit.
[0031] The processing related to the above steps is detailed below.
[0032] 2.1.2.2 Spatial candidate derivation In the derivation of spatial merge candidates, up to four merge candidates are selected from candidates located at the positions shown in Figure 2. The order of derivation is A 1 , B 1 , B 0 , A 0 , and B 2is. Position B 2 is position A 1 , B 1 , B 0 , A 0 is considered only when any of the PUs of A, B, B, A is not available (for example, because it belongs to another slice or tile) or is intra-coded. After the candidate for position A 1 is added, the addition of the remaining candidates is subjected to a redundancy check that ensures that candidates having the same motion information are excluded from the list so as to improve the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the above-mentioned redundancy check. Instead, only the pairs connected by the arrows in Figure 3 are considered, and a candidate is added to the list only when the corresponding candidates used for the redundancy check do not have the same motion information. Another source of overlapping motion information is the "second PU" associated with a different partitioning than 2N×2N. As an example, Figure 4 shows the second PUs for the cases of N×2N and 2N×N, respectively. When the current PU is partitioned into N×2N, the candidate at A 1 is not considered for list construction. In fact, adding this candidate would lead to two prediction units having the same motion information, which is redundant for having only one PU within the coding unit. Similarly, when the current PU is partitioned into 2N×N, position B 1 is not considered.
[0033] 2.1.2.3 Temporal Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, scaled motion vectors are derived based on the collocated PUs in the collocated picture. The reference picture list used for the derivation of the collocated PUs is explicitly signaled in the slice header. The scaled motion vectors for the temporal merge candidate are obtained as shown by the dotted line in Figure 5, which are scaled from the motion vectors of the collocated PUs (col_PU) using the POC distances tb and td, where tb is defined as the POC difference between the reference picture (curr_ref) of the current picture (curr_pic) and the current picture, and td is defined as the POC difference between the reference picture (col_ref) of the collocated picture (col_pic) and the collocated picture. The reference picture index of the temporal merge candidate is set equal to zero. The actual implementation of the scaling process is defined in the HEVC specification. In a B slice, two motion vectors, one related to reference picture list 0 and the other related to reference picture list 1, are obtained, and these are combined to form the bi-prediction merge candidate.
[0034] 2.1.2.4 Collocated Picture and Collocated PU When TMVP is enabled (i.e., slice_temporal_mvp_enabled_flag is equal to 1), the variable ColPic representing the collocated picture is derived as follows: - If the current slice is a B slice and the signaled collocated_from_10_flag is equal to 0, ColPic is set equal to RefPicList1[collocated_ref_idx]; - Otherwise (slice_type is equal to B and collocated_from_l0_flag is equal to 1, or slice_type is equal to P), ColPic is set equal to RefPicList0[collocated_ref_idx]; However, collocated_ref_idx and collocated_from_10_flag are two syntax elements that can be signaled within the slice header.
[0035] For the collocated PU (Y) belonging to the reference frame, as shown in Figure 6, the position regarding the temporal candidate is between candidates C 0 and C 1 is selected. If the PU at position C 0 is not available, or is intra-coded, or is outside the current coding tree unit (CTU, also known as the largest coding unit LCU) row, then position C 1 is used. Otherwise, position C 0 is used for the derivation of the temporal merge candidate.
[0036] The related syntax elements are described as follows:
Table 1
[0037] 2.1.2.5 Derivation of the MV for TMVP Candidates More specifically, to derive the TMVP candidates, the following steps are executed: 1) Set the reference picture list X = 0 and set the target reference picture to the reference picture having an index equal to 0 within list X (i.e., curr_ref); 2) If the current slice is a B slice, set the reference picture list X = 1 and set the target reference picture to the reference picture having an index equal to 0 within list X (i.e., curr_ref). To obtain the MV regarding list X pointing to curr_ref, call the derivation process of the collocated motion vector.
[0038] The derivation process of the collocated motion vector is described in the following Subsection 2.1.2.5.1.
[0039] 2.1.2.5.1 Derivation Process of Collocated Motion Vectors For a collocated block, it can be intra-coded or inter-coded with single prediction or dual prediction. If it is intra-coded, the TMVP candidates are set to unavailable.
[0040] If it is single prediction from list A, the motion vector of list A is scaled to the target reference picture list X.
[0041] If it is dual prediction and the target reference picture list is X, the motion vector of list A is scaled to the target reference picture list X, and A is determined according to the following rules: - If none of the reference pictures have POC values greater than that of the current picture, A is set equal to X; - Otherwise, A is set equal to collocated_from_10_flag.
[0042] The related working draft of JCTVC-W1005-v4 is described as follows.
[0043] 8.5.3.2.9 Derivation Process of Collocated Motion Vectors The inputs to this process are as follows: - The variable currPb that defines the current prediction block, - The variable colPb that defines the collocated prediction block inside the collocated picture defined by ColPic, - The luma position (xColPb, yColPb) that defines the top-left sample of the collocated luma prediction block defined by ColPb with respect to the top-left luma sample of the collocated picture defined by ColPic, - The reference index refIdxLX with X being 0 or 1.
[0044] The outputs of this process are as follows: - The motion vector prediction mvLXCol - Use the availability flag availableFlagLXCol.
[0045] The variable currPic defines the current picture.
[0046] The arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] are set equal to PredFlagL0[x][y], MvL0[x][y], and RefIdxL0[x][y] respectively of the collocated picture defined by ColPic, and the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] are set equal to PredFlagL1[x][y], MvL1[x][y], and RefIdxL1[x][y] respectively of the collocated picture defined by ColPic.
[0047] The variables mvLXCol and availableFlagLXCol are derived as follows: - If colPb is encoded in the intra prediction mode, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0. - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows: - If predFlagL0Col[xColPb][yColPb] is equal to 0, mvCol, refIdxCol, and listCol are set equal to mvL1Col[xColPb][yColPb], refIdxL1Col[xColPb][yColPb], and L1 respectively; - Otherwise, when predFlagL0Col[xColPb][yColPb] is equal to 1 and predFlagL1Col[xColPb][yColPb] is equal to 0, mvCol, refIdxCol, and listCol are set equal to mvL0Col[xColPb][yColPb], refIdxL0Col[xColPb][yColPb], and L0, respectively; - In other cases (when predFlagL0Col[xColPb][yColPb] is equal to 1 and predFlagL1Col[xColPb][yColPb] is equal to 1), the following assignment is made: - When NoBackwardPredFlag is equal to 1, mvCol, refIdxCol, and listCol are set equal to mvLXCol[xColPb][yColPb], refIdxLXCol[xColPb][yColPb], and LX, respectively; - In other cases, mvCol, refIdxCol, and listCol are set equal to mvLNCol[xColPb][yColPb], refIdxLNCol[xColPb][yColPb], and LN, respectively, where N is the value of collocated_from_l0_flag; And mvLXCol and availableFlagLXCol are derived as follows: - When LongTermRefPic(currPic,currPb,refIdxLX,LX) is not equal to LongTermRefPic(ColPic,colPb,refIdxCol,listCol), both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0; - Otherwise, the variable availableFlagLXCol is set equal to 1, and refPicListCol[refIdxCol] is set to be the picture having the reference index refIdxCol in the reference picture list listCol of the slice containing the prediction block colPb in the collocated picture defined by ColPic, and the following applies: colPocDiff = DiffPicOrderCnt(ColPic, refPicListCol[refIdxCol]) (2-1) currPocDiff = DiffPicOrderCnt(currPic, RefPicListX[refIdxLX]) (2-2) - If RefPicListX[refIdxLX] is a long-term reference picture or colPocDiff is equal to currPocDiff, mvLXCol is derived as follows: mvLXCol = mvCol (2-3) - Otherwise, mvLXCol is derived as a scaled version of the motion vector mvCol as follows: tx = (16384+(Abs(td)>>1)) / td (2-4) distScaleFactor = Clip3(-4096, 4095, (tb*tx+32)>>6) (2-5) mvLXCol = Clip3(-32768, 32767, Sign(distScaleFactor*mvCol)* ((Abs(distScaleFactor*mvCol)+127)>>8)) (2-6) Here, td and tb are derived as follows: td = Clip3(-128, 127, colPocDiff) (2-7) tb = Clip3(-128, 127, currPocDiff) (2-8)
[0048] The definition of NoBackwardPredFlag is as follows: The variable NoBackwardPredFlag is derived as follows: - If DiffPicOrderCnt(aPic, CurrPic) is less than or equal to 0 for each picture aPic in RefPicList0 or RefPicList1 of the current slice, NoBackwardPredFlag is set to 1; - Otherwise, NoBackwardPredFlag is set equal to 0.
[0049] 2.1.1.4 Additional Candidate Insertion In addition to the spatial and temporal merge candidates, there are two additional types of merge candidates: combined bi-prediction merge candidates and zero merge candidates. The combined bi-prediction merge candidates are generated by utilizing the spatial and temporal merge candidates. The combined bi-prediction merge candidates are used only for B slices. The combined bi-prediction candidates are generated by combining the motion parameters of the first reference picture list of the initial candidates with those of the second reference picture list of another. If these two tuples provide different motion hypotheses, they form a new bi-prediction candidate. As an example, Figure 7 shows the case where combined bi-prediction merge candidates are created to be added to the final list (right side) using two candidates within the original list (left side) that have mvL0 and refIdxL0, or mvL1 and refIdxL1. There are numerous rules regarding the combinations considered to generate these additional merge candidates.
[0050] Zero motion candidates are inserted to fill the remaining entries of the merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have a zero spatial displacement and a reference picture index that starts from zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0051] 2.1.3 AMVP The AMVP utilizes the spatio-temporal correlation of motion vectors with adjacent PUs and it is used for explicit transmission of motion parameters. For each reference picture list, first, the availability of the temporally adjacent PU position in the upper left is examined, redundant candidates are removed, and a motion vector candidate list is constructed by adding zero vectors to make the candidate list a fixed length. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similarly, in merge index signaling, the index of the best motion vector candidate is encoded using truncated unary. The maximum value encoded in this case is 2 (see Figure 8). The following section provides details of the derivation process of motion vector prediction candidates.
[0052] 2.1.3.1 Derivation of AMVP Candidates Figure 8 summarizes the derivation process for motion vector prediction candidates.
[0053] In motion vector prediction, two types of motion vector candidates, namely spatial motion vector candidates and temporal motion vector candidates, are considered. In spatial motion vector candidate derivation, based on the motion vectors of each PU at five different positions as shown in Figure 2, finally two motion vector candidates are derived.
[0054] In temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different collocated positions. After the first list of spatio-temporal candidates is created, duplicate motion vector candidates within the list are removed. If the number of possible candidates is more than 2, the motion vector candidates whose reference picture index is greater than 1 within the relevant reference picture list are removed from the list. If the number of spatio-temporal motion vector candidates is less than 2, additional zero motion vector candidates are added to the list.
[0055] 2.1.3.2 Spatial Motion Vector Candidates In the derivation of the spatial motion vector candidates, the maximum of two candidates out of the five possible candidates derived from the PU at the position as shown in Figure 2 are considered, and their positions are the same as those of the motion merge. The derivation order for the left side of the current PU is A 0 、A 1 、and scaled A 0 、scaled A 1 as defined. The derivation order for the upper side of the current PU is B 0 、B 1 、B 2 、scaled B 0 、scaled B 1 、scaled B 2 as defined. Therefore, for each side, there are four cases that can be used as motion vector candidates, two cases do not require the use of spatial scaling, and spatial scaling is used in two cases. These four different cases are summarized as follows: · Without spatial scaling - (1) The same reference picture list and the same reference picture index (the same POC) - (2) Different reference picture lists, but the same reference picture (the same POC) · Spatial scaling - (3) The same reference picture list, but different reference pictures (different POCs) - (4) Different reference picture lists and different reference pictures (different POCs).
[0056] The cases without spatial scaling are inspected first, followed by spatial scaling. Spatial scaling is considered when the POC is different between the reference picture of the adjacent PU and the reference picture of the current PU regardless of the reference picture list. When all the PUs of the left candidate are not available or are intra-coded, scaling for the upper motion vector is enabled to assist in the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling is not enabled for the upper motion vector.
[0057] In the spatial scaling process, as shown in Figure 9, the motion vectors of adjacent PUs are scaled in a similar way as for temporal scaling. The main difference is that the reference picture list and the index of the current PU are given as input, and the actual scaling process is the same as that of temporal scaling.
[0058] 2.1.3.3 Temporal Motion Vector Candidates Apart from the derivation of the reference picture index, all processes for the derivation of temporal merge candidates are the same as those for the derivation of spatial motion vector candidates (see Figure 6). The reference picture index is signaled to the decoder.
[0059] 2.2 Inter Prediction Methods in VVC There are several new coding tools for inter prediction improvement, such as Adaptive Motion Vector difference Resolution (AMVR) for signaling MVD, Merge with Motion Vector Differences (MMVD), Triangular prediction mode (TPM), Combined intra-inter prediction (CIIP), Advanced TMVP (also known as ATMVP, SbTMVP), Affine prediction mode, Generalized Bi-Prediction (GBI), Decoder-side Motion Vector Refinement (DMVR), and Bi-directional Optical flow (also known as BIO, BDOF).
[0060] Three different merge list construction processes are supported in VVC: 1) Sub-block merge candidate list: This includes ATMVP candidates and affine merge candidates. One merge list construction process is shared in both the affine mode and the ATMVP mode. Here, ATMVP candidates and affine merge candidates can be added in sequence. The sub-block merge list size is signaled in the slice header, and the maximum value is 5; 2) Normal merge list: In the inter-coded block, one merge list construction process is shared. Here, spatial / temporal merge candidates, HMVP, pairwise merge candidates, and zero motion candidates can be inserted in sequence. The normal merge list size is signaled in the slice header, and the maximum value is 6. MMVD, TPM, CIIP rely on the normal merge list; 3) IBC merge list: This is done in the same way as the normal merge list.
[0061] Similarly, three AMVP lists are supported in VVC: 1) Affine AMVP candidate list; 2) Normal AMVP candidate list; 3) IBC AMVP candidate list: The same construction process as the IBC merge list by the adoption of JVET-N0843.
[0062] 2.2.1 Encoded block structure in VVC In VVC, a Quad-Tree / Binary-Tree / Ternary-Tree (QT / BT / TT) structure is adopted to divide a picture into square or rectangular blocks.
[0063] In addition to QT / BT / TT, VVC also adopts a separate tree (also known as a dual coding tree) for I-frames. In the separate tree, the encoded block structure is signaled separately for the luma component and the chroma component.
[0064] Furthermore, except for blocks encoded with specific encoding methods of 2 to 3 (for example, intra sub - partition prediction where PU is equal to TU but smaller than CU, and sub - block transformation for inter - encoded blocks where PU is equal to CU but TU is smaller than PU), the CU is set equal to the PU and TU.
[0065] 2.2.2 Affine prediction mode In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). On the other hand, in the real world, there are many types of motions such as zoom - in / out, rotation, perspective motion, and other irregular motions. In VVC, simplified affine - transform motion compensation prediction is applied using the 4 - parameter affine model and the 6 - parameter affine model. As shown in Figure 10, the affine motion field of a block is described by two control - point motion vectors (CPMV) in the 4 - parameter affine model and three CPMV in the 6 - parameter affine model.
[0066] The motion vector field (MVF) of a block is given by Equation (1) in the 4 - parameter affine model (where the four parameters are defined as variables a, b, e, and f) and Equation (2) in the 6 - parameter affine model (where the six parameters are defined as variables a, b, c, d, e, and f):
Equation
[0067] To further simplify motion compensation prediction, sub-block-based affine transform prediction is applied. To derive the motion vector for each M×N sub-block (where M and N are both set to 4 in the current VVC), as shown in Figure 11, the motion vector of the center sample of each sub-block is calculated according to Equations (1) and (2) and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter for 1 / 16 pel is applied to generate the prediction for each sub-block with the derived motion vector. The interpolation filter for 1 / 16 pel is introduced by the affine mode.
[0068] After MCP, the high-precision motion vectors of each sub-block are rounded and saved with the same precision as the normal motion vectors.
[0069] 2.2.3 Merge for the Entire Block 2.2.3.1 Construction of the Merge List for the Translational Normal Merge Mode 2.2.3.1.1 History-Based Motion Vector Prediction (HMVP) Unlike the merge list design, in VVC, the History-based Motion Vector Prediction (HMVP) method is adopted.
[0070] In HMVP, the previously encoded motion information is stored. The motion information of the previously encoded block is defined as an HMVP candidate. Multiple HMVP candidates are stored in a table named the HMVP table, and this table is maintained on-the-fly during the encoding / decoding process. The HMVP table is emptied when starting to encode / decrypt a new tile / LCU row / slice. Whenever there are inter-encoded blocks and non-TPM modes of non-sub-blocks, the relevant motion information is added to the last entry of the table as a new HMVP candidate. The overall encoding flow is shown in Figure 12.
[0071] 2.2.3.1.2 Normal Merge List Construction Process The construction of the regular (for translational motion) merge list can be summarized according to the following series of steps: · Step 1: Derivation of spatial candidates; · Step 2: Insertion of HMVP candidates; · Step 3: Insertion of pairwise average candidates; · Step 4: Default motion candidates.
[0072] The HMVP candidates can be used in both the AMVP candidate list construction process and the merge candidate list construction process. FIG. 13 shows the modified merge candidate list construction process (shown using boxed dots). If the merge candidate list is not full after the insertion of TMVP candidates, the merge candidate list can be filled using the HMVP candidates stored in the HMVP table. The blocks are usually inserted into the HMVP table in descending order of index considering that they have a higher correlation with the closest neighboring block with respect to motion information. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy removal is applied to the HMVP candidates. When the total number of available merge candidates reaches the maximum number of merge candidates allowed to be signaled, the merge candidate list construction process ends.
[0073] Note that all spatial / temporal / HMVP candidates are encoded in non-IBC mode. Otherwise, it is not allowed to be added to the normal merge candidate list.
[0074] The HMVP table contains a maximum of five normal motion candidates, each of which is unique.
[0075] 2.2.3.1.2.1 Pruning Process A candidate is added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information. Such a comparison process is called a pruning process.
[0076] The pruning process among spatial candidates depends on the use of TPM for the current block.
[0077] When the current block is encoded without using TPM mode (e.g., normal merge, MMVD, CIIP), the HEVC pruning process (i.e., 5 pruning) for spatial merge candidates is utilized.
[0078] 2.2.4 Triangular Prediction Mode (TPM) In VVC, the triangular partition mode is supported for inter prediction. The triangular partition mode is applied only to CUs that are larger than 8×8 and are encoded in merge mode rather than the MMVD or CIIP mode. For CUs that meet these conditions, a CU-level flag is signaled to indicate whether the triangular partition mode is applied.
[0079] When this mode is used, the CU is equally divided into two triangular partitions using either a diagonal split or an anti-diagonal split, as shown in FIG. 14. Each triangular partition within the CU is inter predicted using its own motion, and only uni-prediction is allowed for each partition, i.e., each partition has one motion vector and one reference index. This uni-prediction motion constraint is applied to ensure the same as the conventional bi-prediction where only two motion compensation predictions are required for each CU.
[0080] If the CU-level flag indicates that the current CU is encoded using the triangular partition mode, a flag indicating the direction of the triangular partition (diagonal or anti-diagonal), and two merge indices (one for each partition) are further signaled. After predicting each of the triangular partitions, the sample values along the diagonal edge or anti-diagonal edge are adjusted using a mixing process with adaptive weights. This is the prediction signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangular partition mode is stored in 4×4 units.
[0081] The normal merge candidate list is reused for triangular partition merge prediction without additional motion vector pruning. For each merge candidate in the normal merge candidate list, only one of its L0 or L1 motion vectors and only that one is used for triangular prediction. In addition, the order of selecting the motion vector between L0 and L1 is based on its merge index parity. Using this scheme, the normal merge list can be used directly.
[0082] 2.2.4.1 TPM-oriented Merge List Construction Process Basically, the normal merge list construction process is applied as proposed in JVET-N0340. However, some changes are made.
[0083] Specifically, the following are applied: 1) How the pruning process is performed depends on the use of TPM for the current block, - If the current block is not encoded with TPM, the HEVC 5 pruning applied to the spatial merge candidates is invoked; - Otherwise (if the current block is encoded with TPM), full pruning is applied when adding new spatial merge candidates. That is, B1 is compared with A1, B0 is compared with A1 and B1, A0 is compared with A1, B1 and B0, and B2 is compared with A1, B1, A0 and B0; 2) The condition regarding whether to examine the motion information from B2 depends on the use of TPM for the current block, - If the current block is not encoded with TPM, B2 is accessed and examined only if there are less than four spatial merge candidates before examining B2; - Otherwise (if the current block is encoded with TPM), B2 is always accessed and examined regardless of the number of available spatial merge candidates before adding B2.
[0084] 2.2.4.2 Adaptive Weighting Process After predicting each triangular prediction unit, an adaptive weighting process is applied to the diagonal edges between two triangular prediction units to derive the final prediction for the entire CU. Two groups of weighting coefficients are defined as follows: · The first group of weighting coefficients: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luminance samples and chrominance samples, respectively; · Second weighting coefficient group: For luminance samples and chrominance samples, {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used respectively.
[0085] The weighting coefficient group is selected based on the comparison of the motion vectors of two triangular prediction units. The second weighting coefficient group is used when any one of the following conditions is true: - The reference pictures of the two triangular prediction units are different from each other; - The absolute value of the difference in the horizontal values of the two motion vectors is greater than 16 pixels; - The absolute value of the difference in the vertical values of the two motion vectors is greater than 16 pixels.
[0086] Otherwise, the first weighting coefficient group is used. An example is shown in FIG. 15.
[0087] 2.2.4.3 Motion Vector Storage The motion vectors of the triangular prediction units (Mv1 and Mv2 in FIG. 16) are stored in a 4×4 grid. For each 4×4 grid, according to the position of the 4×4 grid within the CU, either a single prediction or a dual prediction motion vector is stored. As shown in FIG. 16, for a 4×4 grid located in the non-weighted area (i.e., not located on the diagonal edge), a single prediction motion vector, which is either Mv1 or Mv2, is stored. On the other hand, for a 4×4 grid located in the weighted area, a dual prediction motion vector is stored. The dual prediction motion vector is derived from Mv1 and Mv2 according to the following rules: 1) When Mv1 and Mv2 have motion vectors from different directions (L0 or L1), simply combine Mv1 and Mv2 to form a dual prediction motion vector; 2) When both Mv1 and Mv2 are from the same L0 (or L1) direction, - If the reference picture of Mv2 is the same as the picture in the L1 (or L0) reference picture list, Mv2 is scaled to that picture. Combine Mv1 and the scaled Mv2 to form a bi-predicted motion vector; - If the reference picture of Mv1 is the same as the picture in the L1 (or L0) reference picture list, Mv1 is scaled to that picture. Combine the scaled Mv1 and Mv2 to form a bi-predicted motion vector; - Otherwise, only Mv1 is stored for the weighted region.
[0088] 2.2.4.4 Syntax Tables, Semantics, and Decoding Process for Merge Mode
Table 2
Table 3
Table 4
[0089] 7.4.6.1 General Slice Header Semantics six_minus_max_num_merge_cand defines the result of subtracting the maximum number of merge motion vector prediction (MVP) candidates supported within a slice from 6. The maximum number of merge MVP candidates, MaxNumMergeCand, is derived as follows: MaxNumMergeCand = 6 - six_minus_max_num_merge_cand (7-57) The value of MaxNumMergeCand shall be in the range of 1 to 6, inclusive at both ends.
[0090] five_minus_max_num_subblock_merge_cand defines the result of subtracting the maximum number of sub-block based merge motion vector prediction (MVP) candidates supported within a slice from 5. When five_minus_max_num_subblock_merge_cand does not exist, it is assumed to be equal to 5 - sps_sbtmvp_enabled_flag. The maximum number of sub-block based merge MVP candidates MaxNumSubblockMergeCand is derived as follows: MaxNumSubblockMergeCand = 5 - five_minus_max_num_subblock_merge_cand (7-58) The value of MaxNumSubblockMergeCand is in the range from 1 to 5, inclusive. of.
[0091] 7.4.8.5 Encoding Unit Semantics A pred_mode_flag equal to 0 specifies that the current coding unit is coded in the inter prediction mode. A pred_mode_flag equal to 1 specifies that the current coding unit is coded in the intra prediction mode.
[0092] When pred_mode_flag does not exist, it is assumed as follows: - If cbWidth is equal to 4 and cbHeight is equal to 4, pred_mode_flag is assumed to be equal to 1; - Otherwise, pred_mode_flag is assumed to be equal to 1 when decoding an I slice and equal to 0 when decoding a P or B slice.
[0093] The variable CuPredMode[x][y] is derived as follows for x = x0..x0 + cbWidth - 1 and y = y0..y0 + cbHeight - 1: - When pred_mode_flag is equal to 0, CuPredMode[x][y] is set to be equal to MODE_INTER; - Otherwise (when pred_mode_flag is equal to 1), CuPredMode[x][y] is set to be equal to MODE_INTRA.
[0094] pred_mode_ibc_flag equal to 1 specifies that the current coding unit is coded in the IBC prediction mode. pred_mode_ibc_flag equal to 0 specifies that the current coding unit is not coded in the IBC prediction mode.
[0095] When pred_mode_ibc_flag does not exist, it is estimated as follows: - When cu_skip_flag[x0][y0] is equal to 1, cbWidth is equal to 4, and cbHeight is equal to 4, pred_mode_ibc_flag is estimated to be equal to 1; - Otherwise, when both cbWidth and cbHeight are equal to 128, pred_mode_ibc_flag is estimated to be equal to 0; - Otherwise, pred_mode_ibc_flag is assumed to be equal to the value of sps_ibc_enabled_flag when decoding an I slice, and assumed to be equal to 0 when decoding a P or B slice.
[0096] When pred_mode_ibc_flag is equal to 1, the variable CuPredMode[x][y] is set to be equal to MODE_IBC for x = x0..x0 + cbWidth - 1 and y = y0..y0 + cbHeight - 1.
[0097] general_merge_flag[x0][y0] defines whether the inter-prediction parameters for the current coding unit are inferred from adjacent inter-prediction partitions. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture.
[0098] When general_merge_flag[x0][y0] does not exist, it is inferred as follows: - If cu_skip_flag[x0][y0] is equal to 1, general_merge_flag[x0][y0] is inferred to be equal to 1; - Otherwise, general_merge_flag[x0][y0] is inferred to be equal to 0.
[0099] mvp_l0_flag[x0][y0] defines the motion vector predictor index for list 0, and x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When mvp_l0_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0100] mvp_l1_flag[x0][y0] has the same meaning as mvp_l0_flag, with l0 and list 0 replaced by l1 and list 1 respectively.
[0101] inter_pred_idc[x0][y0] defines, according to Table 7-10, whether list 0 is used for the current coding unit, whether list 1 is used, or whether bi-prediction is used. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture.
Table 5
[0102] When inter_pred_idc[x0][y0] does not exist, it is presumed to be equal to PRED_L0.
[0103] 7.4.8.7 Merge Data Semantics A regular_merge_flag[x0][y0] equal to 1 specifies that the normal merge mode is used to generate the inter prediction parameters of the current coding unit. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0104] When regular_merge_flag[x0][y0] does not exist, it is presumed as follows: - If all of the following conditions are true, regular_merge_flag[x0][y0] is presumed to be equal to 1: - sps_mmvd_enabled_flag is equal to 0; - general_merge_flag[x0][y0] is equal to 1; - cbWidth*cbHeight is equal to 32; - Otherwise, regular_merge_flag[x0][y0] is presumed to be equal to 0.
[0105] A mmvd_merge_flag[x0][y0] equal to 1 specifies that the motion vector difference merge mode is used to generate the inter prediction parameters of the current coding unit. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0106] When mmvd_merge_flag[x0][y0] does not exist, it is presumed as follows: - If all of the following conditions are true, mmvd_merge_flag[x0][y0] is presumed to be equal to 1: - The sps_mmvd_enabled_flag is equal to 1; - The general_merge_flag[x0][y0] is equal to 1; - The cbWidth * cbHeight is equal to 32; - The regular_merge_flag[x0][y0] is equal to 0; - Otherwise, the mmvd_merge_flag[x0][y0] is presumed to be equal to 0.
[0107] The mmvd_cand_flag[x0][y0], together with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0], defines whether the first candidate (0) or the second candidate (1) within the merge candidate list is used. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When mmvd_cand_flag[x0][y0] does not exist, it is presumed to be equal to 0.
[0108] The mmvd_distance_idx[x0][y0] defines the index used to derive MmvdDistance[x0][y0] as specified in Table 7-12. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture.
Table 6
Table 7
[0109] Both components of the merge + MVD offset MmvdOffset[x0][y0] are derived as follows: MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] (7-124) MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] (7-125)
[0110] merge_subblock_flag[x0][y0] defines whether the sub-block based inter-prediction parameter for the current coding unit is estimated from adjacent blocks. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When merge_subblock_flag[x0][y0] does not exist, it is assumed to be equal to 0.
[0111] merge_subblock_idx[x0][y0] defines the merge candidate index in the sub-block based merge candidate list, and x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When merge_subblock_idx[x0][y0] does not exist, it is assumed to be equal to 0.
[0112] ciip_flag[x0][y0] defines whether combined inter-picture merge + intra-picture prediction is applied to the current coding unit. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When ciip_flag[x0][y0] does not exist, it is assumed to be equal to 0.
[0113] When ciip_flag[x0][y0] is equal to 1, the variable IntraPredModeY[x][y] at x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1 is set to be equal to INTRA_PLANAR. The variable MergeTriangleFlag[x0][y0], which defines whether motion compensation based on a triangular shape is used to generate prediction samples for the current coded unit, is derived as follows when decoding a B slice: - MergeTriangleFlag[x0][y0] is set equal to 1 if all of the following conditions are satisfied: - sps_triangle_enabled_flag is equal to 1; - slice_type is equal to B; - general_merge_flag[x0][y0] is equal to 1; - MaxNumTriangleMergeCand is 2 or more; - cbWidth * cbHeight is 64 or more; - regular_merge_flag[x0][y0] is equal to 0; - mmvd_merge_flag[x0][y0] is equal to 0; - merge_subblock_flag[x0][y0] is equal to 0; - cip_flag[x0][y0] is equal to 0; - Otherwise, MergeTriangleFlag[x0][y0] is set equal to 0.
[0114] merge_triangle_split_dir[x0][y0] defines the splitting direction in the merge triangle mode. The array indices x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When merge_triangle_split_dir[x0][y0] does not exist, it is assumed to be equal to 0.
[0115] merge_triangle_idx0[x0][y0] defines the first merge candidate index in the motion compensation candidate list based on the triangle shape, and x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When merge_triangle_idx0[x0][y0] does not exist, it is assumed to be equal to 0.
[0116] merge_triangle_idx1[x0][y0] defines the second merge candidate index in the motion compensation candidate list based on the triangle shape, and x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture. When merge_triangle_idx1[x0][y0] does not exist, it is assumed to be equal to 0.
[0117] merge_idx[x0][y0] defines the merge candidate index in the merge candidate list, and x0, y0 define the position (x0, y0) of the top-left luma sample of the coding block under consideration with respect to the top-left luma sample of the picture.
[0118] When merge_idx[x0][y0] does not exist, it is assumed as follows: - If mmvd_merge_flag[x0][y0] is equal to 1, merge_idx[x0][y0] is assumed to be equal to mmvd_cand_flag[x0][y0]; - Otherwise (mmvd_merge_flag[x0][y0] is equal to 0), merge_idx[x0][y0] is assumed to be equal to 0.
[0119] 2.2.4.4.1 Decoding Process The decoding process provided in JVET-N0340 is defined as follows.
[0120] 8.5.2.2 Derivation Process of Luma Motion Vectors for Merge Mode This process is called only when general_merge_flag[xCb][yCb] is equal to 1, where (xCb,yCb) defines the top-left sample of the current luma coded block with respect to the top-left luma sample of the current picture.
[0121] The inputs to this process are as follows: - The luma position (xCb,yCb) of the top-left sample of the current luma coded block with respect to the top-left luma sample of the current picture; - The variable cbWidth that defines the width of the current coded block within the luma sample; - The variable cbHeight that defines the height of the current coded block within the luma sample.
[0122] The outputs of this process are as follows: - Luma motion vectors mvL0[0][0] and mvL1[0][0] with 1 / 16 fraction sample accuracy; - Reference indices refIdxL0 and refIdxL1; - Prediction list utilization flags predFlagL0[0][0] and predFlagL1[0][0]; - Bi-prediction weight index bcwIdx; - Merge candidate list mergeCandList.
[0123] The bi-prediction weight index bcwIdx is set equal to 0.
[0124] The motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, and the prediction usage flags predFlagL0[0][0] and predFlagL1[0][0] are derived by the following ordered steps: 1. The spatial merge candidate derivation process from the neighboring coded units defined in Section 8.5.2.3 is called with the luma coding block position (xCb, yCb), the luma coding block width cbWidth, and the luma coding block height cbHeight as inputs, and its output is the availability flag availableFlagA where X is 0 or 1 0 、availableFlagA 1 、availableFlagB 0 、availableFlagB 1 、and availableFlagB 2 、the reference index refIdxLXA 0 、refIdxLXA 1 、refIdxLXB 0 、refIdxLXB 1 、and refIdxLXB 2 、the prediction list usage flag predFlagLXA 0 、predFlagLXA 1 、predFlagLXB 0 、predFlagLXB 1 、and predFlagLXB 2 、and the motion vector mvLXA 0 、mvLXA 1 、mvLXB 0 、mvLXB 1 、and mvLXB 2 and the bi-prediction weight index bcwIdxA 0 、bcwIdxA 1 、bcwIdxB 0 、bcwIdxB 1 、bcwIdxB 2 ; 2. The reference index refIdxLXCol where X is 0 or 1 and the bi-prediction weight index bcwIdxCol for the temporal merge candidate Col are set equal to 0; 3. The derivation process of the temporal luma motion vector prediction defined in item 8.5.2.11 is called with the luma position (xCb, yCb), luma coded block width cbWidth, luma coded block height cbHeight, and variable refIdxL0Col as inputs, and its outputs are the availability flag availableFlagL0Col and the temporal motion vector mvL0Col. The variables availableFlagCol, predFlagL0Col, and predFlagL1Col are derived as follows: availableFlagCol = availableFlagL0Col (8-263) predFlagL0Col = availableFlagL0Col (8-264) predFlagL1Col = 0 (8265) 4. When slice_type is equal to B, the derivation process of the temporal luma motion vector prediction defined in item 8.5.2.11 is called with the luma position (xCb, yCb), luma coded block width cbWidth, luma coded block height cbHeight, and variable refIdxL1Col as inputs, and its outputs are the availability flag availableFlagL1Col and the temporal motion vector mvL1Col. The variables availableFlagCol and predFlagL1Col are derived as follows: availableFlagCol = availableFlagL0Col || availableFlagL1Col (8-266) predFlagL1Col = availableFlagL1Col (8-267) 5. The merge candidate list mergeCandList is constructed as follows: i = 0 if(availableFlagA 1 ) mergeCandList[i++] = A 1 if(availableFlagB 1 ) mergeCandList[i++] = B1 if(availableFlagB 0 ) mergeCandList[i++]=B 0 (8-268) if(availableFlagA 0 ) mergeCandList[i++]=A 0 if(availableFlagB 2 ) mergeCandList[i++]=B 2 if(availableFlagCol) mergeCandList[i++]=Col 6. The variables numCurrMergeCand and numOrigMergeCand are set equal to the number of merge candidates in mergeCandList; 7. When numCurrMergeCand is less than (MaxNumMergeCand - 1) and numHmvpCand is greater than 0, the following applies: - The history-based merge candidate derivation process defined in Section 8.5.2.6 is called with mergeCandList and numCurrMergeCand as inputs and the modified mergeCandList and numCurrMergeCand as outputs; - numOrigMergeCand is set equal to numCurrMergeCand; 8. When numCurrMergeCand is less than MaxNumMergeCand and greater than 1, the following applies: - The process for deriving pairwise average merge candidates defined in item 8.5.2.4 is called with mergeCandList, the reference indices refIdxL0N and refIdxL1N, prediction list usage flags predFlagL0N and predFlagL1N, motion vectors mvL0N and mvL1N, and numCurrMergeCand for all candidates N in the mergeCandList as inputs, and its output is assigned to mergeCandList, numCurrMergeCand, the reference indices refIdxL0avgCand and refIdxL1avgCand, prediction list usage flags predFlagL0avgCand and predFlagL1avgCand, and motion vectors mvL0avgCand and mvL1avgCand of the candidate avgCand added to the mergeCandList. The bi-prediction weight index bcwIdx of the candidate avgCand added to the mergeCandList is set equal to 0; - numOrigMergeCand is set equal to numCurrMergeCand; 9. The process for deriving zero motion vector merge candidates defined in item 8.5.2.5 is called with mergeCandList, the reference indices refIdxL0N and refIdxL1N, prediction list usage flags predFlagL0N and predFlagL1N, motion vectors mvL0N and mvL1N, and numCurrMergeCand for all candidates N in the mergeCandList as inputs, and its output is the reference indices refIdxL0zeroCand m and refIdxL1zeroCand m for all new candidates zeroCand added to mergeCandList, numCurrMergeCand, m the prediction list usage flags predFlagL0zeroCand m and predFlagL1zeroCand m and the motion vectors mvL0zeroCand m and mvL1zeroCand mis assigned. All new candidates zeroCand added to mergeCandList m The bi-prediction weight index bcwIdx of m is set equal to 0. The number of candidates added numZeroMergeCand is set equal to (numCurrMergeCand - numOrigMergeCand). When numZeroMergeCand is greater than 0, m ranges from 0 to numZeroMergeCand - 1 inclusive at both ends.
[0125] 10. Let N be the candidate at position merge_idx[xCb][yCb] in the merge candidate list mergeCandList (N = mergeCandList[merge_idx[xCb][yCb])), and X is replaced by 0 or 1, and the following assignments are made: refIdxLX = refIdxLXN (8 - 269) predFlagLX[0][0] = predFlagLXN (8 - 270) mvLX[0][0][0] = mvLXN[0] (8 - 271) mvLX[0][0][1] = mvLXN[1] (8 - 272) bcwIdx = bcwIdxN (8 - 273) 11. When mmvd_merge_flag[xCb][yCb] is equal to 1, the following applies: - The process for deriving the merge motion vector difference defined in 8.5.2.7 is called with the luma position (xCb, yCb), reference indices refIdxL0 and refIdxL1, and prediction list usage flags predFlagL0[0][0] and predFlagL1[0][0] as inputs and the motion vector differences mMvdL0 and mMvdL1 as outputs; - For X which is 0 or 1, the motion vector difference mMvdLX is added to the merge motion vector mvLX as follows: mvLX[0][0][0] += mMvdLX[0] (8 - 274) mvLX[0][0][1] += mMvdLX[1] (8 - 275) mvLX[0][0][0]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][0]) (8-276) mvLX[0][0][1]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][1]) (8-277)
[0126] 8.5.2.3 Derivation Process of Spatial Merge Candidates The inputs to this process are as follows: - The luma position (xCb, yCb) of the top-left sample of the current luma coding block with respect to the top-left luma sample of the current picture; - A variable cbWidth that defines the width of the current coding block within the luma sample; - A variable cbHeight that defines the height of the current coding block within the luma sample.
[0127] The output of this process is X, which is either 0 or 1, and is as follows: - The availability flag availableFlagA of the adjacent coding unit 0 , availableFlagA 1 , availableFlagB 0 , availableFlagB 1 , and availableFlagB 2 ; - The reference index refIdxLXA of the adjacent coding unit 0 , refIdxLXA 1 , refIdxLXB 0 , refIdxLXB 1 , and refIdxLXB 2 ; - The prediction list usage flag predFlagLXA of the adjacent coding unit 0 , predFlagLXA 1 , predFlagLXB 0 , predFlagLXB 1 , and predFlagLXB 2 ; - Motion vector mvLXA at 1 / 16 fractional sample accuracy of the adjacent symbolization unit 0 , mvLXA 1 , mvLXB 0 , mvLXB 1 , and mvLXB 2 ; - Dual-prediction weight index gbiIdxA 0 , gbiIdxA 1 , gbiIdxB 0 , gbiIdxB 1 , and gbiIdxB 2 .
[0128] availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 For the derivation of, the following applies: - The luma position (xNbA 1 , yNbA 1 ) inside the adjacent luma coded block is set equal to (xCb - 1, yCb + cbHeight - 1); - The block availability derivation process specified in Section 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbA 1 , yNbA 1 ) as input, and its output is assigned to the block availability flag availableA 1 ; - The variables availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 are derived as follows: - If availableA 1 is equal to FALSE, X which is 0 or 1, availableFlagA 1 is set equal to 0, and mvLXA1 Both components are set equal to 0, and refIdxLXA 1 is set equal to -1, predFlagLXA 1 is set equal to 0, and gbiIdxA 1 is set equal to 0; - Otherwise, availableFlagA 1 is set equal to 1 and the following assignments are made: mvLXA 1 =MvLX[xNbA 1 [yNbA 1 (8 - 294) refIdxLXA 1 =RefIdxLX[xNbA 1 [yNbA 1 (8 - 295) predFlagLXA 1 =PredFlagLX[xNbA 1 [yNbA 1 (8 - 296) gbiIdxA 1 =GbiIdx[xNbA 1 [yNbA 1 (8 - 297)
[0129] availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 For the derivation of, the following applies: - The luma position (xNbB 1 , yNbB 1 ) inside the adjacent luma - coded block is set equal to (xCb + cbWidth - 1, yCb - 1); - The block availability derivation process defined in Section 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is applied to the current luma position (xCurr, yCurr) set equal to (xCb, yCb) and the adjacent luma position (xNbB 1 , yNbB1 ) is called with the input, and its output is assigned to the block availability flag availableB 1 ; - The variable availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 are derived as follows: - For X which is 0 or 1 when one or more of the following conditions are true, availableFlagB 1 is set equal to 0, both components of mvLXB 1 are set equal to 0, refIdxLXB 1 is set equal to -1, predFlagLXB 1 is set equal to 0, and gbiIdxB 1 is set equal to 0; - availableB 1 is equal to FALSE; - availableA 1 is equal to TRUE, and the luma positions (xNbA 1 , yNbA 1 ) and (xNbB 1 , yNbB 1 ) have the same motion vector and the same reference index; - Otherwise, availableFlagB 1 is set equal to 1 and the following assignments are made: mvLXB 1 = MvLX[xNbB 1 [yNbB 1 (8 - 298) refIdxLXB 1 = RefIdxLX[xNbB 1 [yNbB 1 (8 - 299) predFlagLXB 1 = PredFlagLX[xNbB 1 [yNbB 1 (8 - 300) gbiIdxB1 =GbiIdx[xNbB 1 [yNbB 1 (8-301)
[0130] availableFlagB 0 、refIdxLXB 0 、predFlagLXB 0 、and mvLXB 0 For the derivation of, the following applies: - The luma position (xNbB 0 , yNbB 0 ) inside the adjacent luma coded block is set equal to (xCb + cbWidth, yCb - 1); - The block availability derivation process defined in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbB 0 , yNbB 0 ) as input, and its output is assigned to the block availability flag availableB 0 ; - The variables availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 are derived as follows: - If one or more of the following conditions are true, then X, which is 0 or 1, sets availableFlagB 0 equal to 0, sets both components of mvLXB 0 equal to 0, sets refIdxLXB 0 equal to -1, sets predFlagLXB 0 equal to 0, and sets gbiIdxB 0 equal to 0; - availableB 0 is equal to FALSE; - availableB1 is equal to TRUE, and the luma position (xNbB 1 , yNbB 1 ) and (xNbB 0 , yNbB 0 ) have the same motion vector and the same reference index; - availableA 1 is equal to TRUE, the luma positions (xNbA 1 , yNbA 1 ) and (xNbB 0 , yNbB 0 ) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1; - Otherwise, availableFlagB 0 is set equal to 1, and the following assignments are made: mvLXB 0 = MvLX[xNbB 0 [yNbB 0 (8-302) refIdxLXB 0 = RefIdxLX[xNbB 0 [yNbB 0 (8-303) predFlagLXB 0 = PredFlagLX[xNbB 0 [yNbB 0 (8-304) gbiIdxB 0 = GbiIdx[xNbB 0 [yNbB 0 (8-305)
[0131] availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 For the derivation of, the following applies: - The luma position (xNbA 0 , yNbA 0 ) inside the adjacent luma coded block is set equal to (xCb - 1, yCb + cbWidth); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) and the neighbouring luma position (xNbA 0 , yNbA 0 ) set equal to (xCb, yCb) as input, and its output is assigned to the block availability flag availableA 0 ; - The variables availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 are derived as follows: - If one or more of the following conditions are true, X, which is 0 or 1, has availableFlagA 0 set equal to 0, both components of mvLXA 0 set equal to 0, refIdxLXA 0 set equal to -1, predFlagLXA 0 set equal to 0, and gbiIdxA 0 set equal to 0; - availableA 0 is equal to FALSE; - availableA 1 is equal to TRUE, and the luma positions (xNbA 1 , yNbA 1 ) and (xNbA 0 , yNbA 0 ) have the same motion vector and the same reference index; - availableB 1 is equal to TRUE, the luma positions (xNbB 1 , yNbB 1 ) and (xNbA 0 , yNbA 0 ) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1; - availableB 0 is equal to TRUE, and the luma positions (xNbB 0 , yNbB 0 ) and (xNbA 0 , yNbA 0 ) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1; - Otherwise, availableFlagA 0 is set equal to 1, and the following assignments are made: mvLXA 0 = MvLX[xNbA 0 [yNbA 0 (8-306) refIdxLXA 0 = RefIdxLX[xNbA 0 [yNbA 0 (8-307) predFlagLXA 0 = PredFlagLX[xNbA 0 [yNbA 0 (8-308) gbiIdxA 0 = GbiIdx[xNbA 0 [yNbA 0 (8-309)
[0132] availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 For the derivation of, the following applies: - The luma position (xNbB 2 , yNbB 2 ) inside the adjacent luma coded block is set equal to (xCb-1, yCb-1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) and the neighbouring luma position (xNbB 2 , yNbB 2 ) set equal to (xCb, yCb), and its output is assigned to the block availability flag availableB 2 ; - The variables availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 are derived as follows: - If one or more of the following conditions are true, X, which is 0 or 1, is such that availableFlagB 2 is set equal to 0, both components of mvLXB 2 are set equal to 0, refIdxLXB 2 is set equal to -1, predFlagLXB 2 is set equal to 0, and gbiIdxB 2 is set equal to 0; - availableB 2 is equal to FALSE; - availableA 1 is equal to TRUE and the luma positions (xNbA 1 , yNbA 1 ) and (xNbB 2 , yNbB 2 ) have the same motion vector and the same reference index; - availableB 1 is equal to TRUE and the luma positions (xNbB 1 , yNbB 1 ) and (xNbB 2 , yNbB 2 ) have the same motion vector and the same reference index; - availableB 0is equal to TRUE, and the luma positions (xNbB 0 , yNbB 0 ) and (xNbB 2 , yNbB 2 ) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1; - availableA 0 is equal to TRUE, and the luma positions (xNbA 0 , yNbA 0 ) and (xNbB 2 , yNbB 2 ) have the same motion vector and the same reference index, and merge_triangle_flag[xCb][yCb] is equal to 1; - availableFlagA 0 + availableFlagA 1 + availableFlagB 0 + availableFlagB 1 is equal to 4, and merge_triangle_flag[xCb][yCb] is equal to 0; - Otherwise, availableFlagB 2 is set to 1, and the following assignments are made: mvLXB 2 = MvLX[xNbB 2 [yNbB 2 (8 - 310) refIdxLXB 2 = RefIdxLX[xNbB 2 [yNbB 2 (8 - 311) predFlagLXB 2 = PredFlagLX[xNbB 2 [yNbB 2 (8 - 312) gbiIdxB 2 = GbiIdx[xNbB 2 [yNbB 2 (8 - 313)
[0133] 2.2.5 MMVD In JVET-L0054, the Ultimate Motion Vector Representation (UMVE, also known as MMVD) is presented. UMVE is used in either skip mode or merge mode together with the proposed motion vector representation method.
[0134] UMVE reuses the same merge candidates as those included in the normal merge candidate list in VVC. Among the merge candidates, a base candidate can be selected, and it is further extended by the proposed motion vector representation method.
[0135] UMVE provides a new Motion Vector Difference (MVD) representation method that represents the MVD using a starting point, a magnitude of motion, and a direction of motion.
[0136] This proposed technique uses the merge candidate list as it is. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE extension.
[0137] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list as follows. [Table 8]
[0138] When the number of base candidates is equal to 1, the base candidate IDX is not signaled.
[0139] The distance index is information on the magnitude of motion. The distance index indicates a predetermined distance from the starting point information. The predetermined distance is as follows. [Table 9]
[0140] The direction index represents the direction of the MVD with respect to the starting point. The direction index can represent the following four directions.
Table 10
[0141] The UMVE flag is sent immediately after sending the skip flag or the merge flag. If the skip flag or the merge flag is true, the UMVE flag is parsed. If the UMVE flag is equal to 1, the UMVE syntax is parsed. However, if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, it is in the AFFINE mode. However, if it is not 1, the skip / merge index is parsed for the skip / merge mode of VTM.
[0142] No additional line buffer is required for UMVE candidates. This is because the software skip / merge candidates are directly used as the base candidates. Using the input UMVE index, the replenishment of the MV is determined immediately before motion compensation. For this, there is no need to hold a long line buffer.
[0143] Under current common test conditions, either the first merge candidate or the second merge candidate in the merge candidate list can be selected as the base candidate.
[0144] UMVE is also known as Merge with MV Differences (MMVD).
[0145] 2.2.6 Combined Intra-Inter Prediction (CIIP) Multi-hypothesis prediction has been proposed in JVET-L0100, in which combined intra-inter prediction is one method for generating multiple hypotheses.
[0146] When multiple hypothesis prediction is applied to improve the intra mode, the multiple hypothesis prediction combines one intra prediction and one merge index prediction. Within a merge CU, one flag is signaled for the merge mode, and when the flag is true, the intra mode is selected from the intra candidate list. For the luma component, the intra candidate list is derived from only one intra prediction mode, i.e., the planar mode. The weights applied to the prediction blocks from intra and inter predictions are determined by the coding modes (intra or non-intra) of two adjacent blocks (A1 and B1).
[0147] 2.2.7 Merge in sub-block based techniques It is proposed to put all sub-block related motion candidates into a separate merge list in addition to the normal merge list for non-sub-block merge candidates.
[0148] Sub-block related motion candidates are put into a separate merge list named'sub-block merge candidate list'.
[0149] In one example, the sub-block merge candidate list includes ATMVP candidates and affine merge candidates.
[0150] The sub-block merge candidate list is filled with candidates in the following order: a. ATMVP candidates (which may or may not be available); b. The affine merge list (including inherited affine candidates and constructed affine candidates); c. Padding as a zero MV 4-parameter affine model.
[0151] 2.2.7.1.1 ATMVP (also known as sub-block temporal motion vector predictor SbTMVP) The basic idea of ATMVP is to derive multiple sets of temporal motion vector predictors for one block. A set of motion information is assigned to each sub-block. When ATMVP merge candidates are generated, motion compensation is performed at the 8×8 level instead of the whole block level.
[0152] In the current design, ATMVP predicts the motion vectors of sub-CUs within a CU in two steps described in the following two subsections 2.2.7.1.1.1 and 2.2.7.1.1.2 respectively.
[0153] 2.2.7.1.1.1 Derivation of initialization motion vector The initialization motion vector is denoted by tempMv. When block A1 is available and not intra-coded (i.e., coded in inter mode or IBC mode), the following is applied to derive the initialization motion vector: - If all of the following conditions are true, tempMv is set equal to the motion vector of block A1 in list 1 denoted by mvL1A 1 : - The reference picture index in list 1 is available (not equal to -1) and it has the same POC value as the collocated picture (i.e., DiffPicOrderCnt(ColPic,RefPicList[1][refIdxL1A 1 ) is equal to 0); - All reference pictures do not have a POC greater than the current picture (i.e., for all pictures aPic in all reference picture lists of the current slice, DiffPicOrderCnt(aPic,currPic) is less than or equal to 0); - The current slice is equal to a B slice; - collocated_from_l0_flag is equal to 0; - Otherwise, if all of the following conditions are true, tempMv is set equal to the motion vector of block A1 in list 0 denoted by mvL0A 1 : - The reference picture index of list 0 is available (not equal to -1); - It has the same POC value as the collocated picture (i.e., DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A 1 ) is equal to 0); - Otherwise, a zero motion vector is used as the initial MV.
[0154] The corresponding block (the MV rounded to the center position of the current block and clipped to be within a certain range if necessary) is identified within the collocated picture signaled in the slice header together with the initialization motion vector.
[0155] If the block is inter-coded, proceed to the second step. Otherwise, the ATMVP candidate is set to NOT available.
[0156] 2.2.7.1.1.2 Sub-CU Motion Derivation The second step is to divide the current CU into a plurality of sub-CUs and obtain the motion information of each sub-CU from the blocks corresponding to each sub-CU in the collocated picture.
[0157] If the corresponding block of the sub-CU is coded in inter mode, use this motion information and call a derivation process for the collocated MV that is the same as the processing in the conventional TMVP process to derive the final motion information of the current sub-CU. Basically, when the corresponding block is predicted from the target list X for single prediction or bi-prediction, the motion vector is used; otherwise, when it is predicted from the target list Y (Y = 1 - X) for single prediction or bi-prediction and NoBackwardPredFlag is equal to 1, the MV related to list Y is used. Otherwise, no motion candidate may be found.
[0158] If the block in the collocated picture identified by the initialization MV and the position of the current sub-CU is intra-coded or IBC-coded, or if no motion candidate is found as described above, the following is further applied.
[0159] The motion vector used to fetch the motion field within the collocated picture R col is denoted as MV. col To minimize the impact of MV scaling, the MVs in the spatial candidate list used to derive MV col are selected by the following method: If the reference picture of the candidate MV is the collocated picture, this MV is selected and used as MV col without any scaling. Otherwise, the MV with the reference picture closest to the collocated picture is selected and MV col is derived using scaling.
[0160] The decoding process related to the collocated motion vector derivation process in JVET-N1001 is described below, with the parts related to ATMVP emphasized in bold underlined italic font: 8.5.2.12 Collocated Motion Vector Derivation Process The inputs to this process are as follows: - The variable currCb that defines the current coded block; - The variable colCb that defines the collocated coded block inside the collocated picture defined by ColPic; - The luma position (xColCb, yColCb) that defines the top-left sample of the collocated coded block defined by colCb with respect to the top-left luma sample of the collocated picture defined by ColPic; - The reference index refIdxLX at X which is 0 or 1; - The flag sbFlag that indicates the sub-block temporal merge candidate.
[0161] The outputs of this process are as follows: - Motion vector prediction mvLXCol at 1 / 16 fraction sample accuracy; - Availability flag availableFlagLXCol.
[0162] Variable currPic defines the current picture.
[0163] Arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] are set equal to PredFlagL0[x][y], MvDmvrL0[x][y], and RefIdxL0[x][y] of the collocated picture defined by ColPic, respectively, and arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] are set equal to PredFlagL1[x][y], MvDmvrL1[x][y], and RefIdxL1[x][y] of the collocated picture defined by ColPic, respectively.
[0164] Variables mvLXCol and availableFlagLXCol are derived as follows: - If colCb is encoded in an intra or IBC prediction mode, both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0; - Otherwise, motion vector mvCol, reference index refIdxCol, and reference list identifier listCol are derived as follows: - If sbFlag is equal to 0, availableFlagLXCol is set to 1 and the following applies: - If predFlagL0Col[xColCb][yColCb] is equal to 0, mvCol, refIdxCol, and listCol are set equal to mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively; - Otherwise, when predFlagL0Col[xColCb][yColCb] is equal to 1 and predFlagL1Col[xColCb][yColCb] is equal to 0, mvCol, refIdxCol, and listCol are set equal to mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively; - In other cases (when predFlagL0Col[xColCb][yColCb] is equal to 1 and predFlagL1Col[xColCb][yColCb] is equal to 1), the following assignments are made: - When NoBackwardPredFlag is equal to 1, mvCol, refIdxCol, and listCol are set equal to mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively; - In other cases, mvCol, refIdxCol, and listCol are set equal to mvLNCol[xColCb][yColCb], refIdxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_l0_flag; (Outer 1) TIFF0007683985000012.tif80170 - When availableFlagLXCol is equal to TRUE, mvLXCol and availableFlagLXCol are derived as follows: - When LongTermRefPic(currPic,currCb,refIdxLX,LX) is not equal to LongTermRefPic(ColPic,colCb,refIdxCol,listCol), both components of mvLXCol are set equal to 0 and availableFlagLXCol is set equal to 0; - Otherwise, the variable availableFlagLXCol is set equal to 1 and refPicList[listCol][refIdxCol] is set to be the picture having the reference index refIdxCol within the reference picture list listCol of the slice containing the coded block colCb within the collocated picture defined by ColPic, and the following applies: colPocDiff = DiffPicOrderCnt(ColPic, refPicList[listCol][refIdxCol]) (8-402) currPocDiff = DiffPicOrderCnt(currPic, RefPicList[X][refIdxLX]) (8-403) - The temporal motion buffer compression process for the collocated motion vector defined in Section 8.5.2.15 is called with mvCol as input and the modified mvCol as output; - If RefPicList[X][refIdxLX] is a long-term reference picture or colPocDiff is equal to currPocDiff, mvLXCol is derived as follows: mvLXCol = mvCol (8-404) - Otherwise, mvLXCol is derived as a scaled version of the motion vector mvCol as follows: tx = (16384+(Abs(td)>>1)) / td (8-405) distScaleFactor = Clip3(-4096, 4095, (tb*tx+32)>>6) (8-406) mvLXCol = Clip3(-131072, 131071, (distScaleFactor*mvCol+ 128-(distScaleFactor*mvCol>=0))>>8)) (8-407) where td and tb are derived as follows: td = Clip3(-128, 127, colPocDiff) (8-408) tb = Clip3(-128, 127, currPocDiff) (8 - 409)
[0165] 2.2.8 Normal Inter Mode (AMVP) 2.2.8.1 AMVP Motion Candidate List Similar to the AMVP design in HEVC, up to two AMVP candidates can be derived. However, an HMVP candidate can also be added after the TMVP candidate. The HMVP candidates in the HMVP table are considered in ascending order of index (i.e., starting from the index equal to 0 which is the oldest one). Up to four HMVP candidates can be examined to find whether its reference picture is the same as the target reference picture (i.e., the same POC value).
[0166] 2.2.8.2 AMVR In HEVC, when use_integer_mv_flag is equal to 0 in the slice header, the motion vector difference (MVD) (between the motion vector and the predicted motion vector of the PU) is signaled in units of 1 / 4 luma samples. In VVC, locally adaptive motion vector resolution (AMVR) is introduced. In VVC, the MVD can be encoded in units of 1 / 4 luma samples, integer luma samples, 4 luma samples (i.e., 1 / 4 pel, 1 pel, 4 pel). This MVD resolution is controlled at the coding unit (CU) level, and an MVD resolution flag is conditionally signaled for each CU having at least one non - zero MVD component.
[0167] For a CU having at least one non - zero MVD component, a first flag is signaled to indicate whether 1 / 4 luma sample MV accuracy is used in that CU. If the first flag (equal to 1) indicates that 1 / 4 luma sample MV accuracy is not used, another flag is signaled to indicate whether integer luma sample MV accuracy or 4 luma sample MV accuracy is used.
[0168] If the first MVD resolution flag of the CU is zero or the CU is not coded (which means that all MVDs within the CU are zero), 1 / 4 luma sample MV resolution is used for that CU. When the CU uses integer luma sample MV precision or 4 luma sample MV precision, the MVPs within the AMVP candidate list for that CU are rounded to the corresponding precision.
[0169] 2.2.8.3 Symmetric Motion Vector Difference in JVET-N1001-v2 In JVET-N1001-v2, symmetric motion vector difference (SMVD) is applied to motion information coding in bi-prediction.
[0170] First, at the slice level, variables RefIdxSymL0 and RefIdxSymL1, which respectively indicate the reference picture indices of list 0 / 1 used in the SMVD mode, are derived using the following steps defined in N1001-v2. When at least one of these two variables is equal to -1, the SMVD mode is disabled.
[0171] 2.2.9 Refinement of Motion Information 2.2.9.1 Decoder-Side Motion Vector Refinement (DMVR) In bi-prediction operation, for the prediction of one block region, two prediction blocks formed using the motion vectors (MVs) of list 0 and list 1 are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of bi-prediction are further refined.
[0172] In DMVR in VVC, MVD mirroring between list 0 and list 1 is assumed as shown in FIG. 19, and bilateral matching is performed to refine the MV, i.e., to find the best MVD among several MVD candidates. The MVs of the two reference picture lists are denoted by MVL0 (L0X, L0Y) and MVL1 (L1X, L1Y). The MVD denoted by (MvdX, MvdY) of list 0 that can minimize the cost function (e.g., SAD) is defined as the best MVD. In the case of the SAD function, it is defined as the SAD between the reference block of list 0 derived from the motion vector (L0X + MvdX, L0Y + MvdY) in the list 0 reference picture and the reference block of list 1 derived from the motion vector (L1X - MvdX, L1Y - MvdY) in the list 1 reference picture.
[0173] This motion vector refinement process can be repeated twice. In each iteration, as shown in FIG. 20, up to six (integer pixel precision) MVDs can be examined in two steps. In the first step, MVD (0, 0), (-1, 0), (1, 0), (0, -1), (0, 1) are examined. In the second step, one of MVD (-1, -1), (-1, 1), (1, -1) or (1, 1) is selected and further examined. Assume that the function Sad(x, y) returns the SAD value of MVD (x, y). The MVD denoted by (MvdX, MvdY) examined in the second step is determined as follows: MvdX = -1; MvdY = -1; If (Sad(1, 0) < Sad(-1, 0)) MvdX = 1; If (Sad(0, 1) < Sad(0, -1)) MvdY = 1;
[0174] In the first iteration, the starting point is the MV being signaled, and in the second iteration, the starting point is the MV being signaled plus the best MVD selected in the first iteration. DMVR is only applied when one reference picture is a previous picture, the other reference picture is a subsequent picture, and these two reference pictures have the same picture order count distance from the current picture.
[0175] To further simplify the DMVR process, JVET-M0147 proposed several changes to the design in JEM. More specifically, the DMVR design adopted in VTM-4.0 (to be released soon) has the following main features: · Early termination when the (0,0) position SAD between list0 and list1 is smaller than the threshold; · Early termination when the SAD between list0 and list1 is zero at some position; · Block size for DMVR: W*H>=64&&H>=8, where W and H are the width and height of the block; · Split the CU into multiple 16×16 sub-blocks for DMVR when the CU size>16×16. If only the width or height of the CU is larger than 16, it is split only in the vertical or horizontal direction; · Reference block size (W+7)*(H+7) (for luma); · 25-point SAD-based integer per search (i.e., (±)2 refinement search range, single stage); · Bicubic interpolation-based DMVR; · Sub-pel refinement based on the "parametric error surface equation". This procedure is only executed when the minimum SAD cost is not equal to zero and the best MVD is (0,0) in the last MV refinement iteration; · Luma / chroma MC with (optional) reference block padding; · Refined MV used only for MC and TMVP.
[0176] 2.2.9.1.1 Usage of DMVR DMVR can be enabled when all of the following conditions are true: - The DMVR enable flag in the SPS (i.e., sps_dmvr_enabled_flag) is equal to 1; - The TPM flag, inter-affinity flag, sub-block merge flag (either ATMVP or affinity merge), and MMVD flag are all equal to 0; - The merge flag is equal to 1; - The current block is bi-predicted and the POC distance between the current picture and the reference picture in list 1 is equal to the POC distance between the reference picture in list 0 and the current picture; - The height of the current CU is 8 or more; - The number of luma samples (width * height of the CU) is 64 or more.
[0177] 2.2.9.1.2 "Parametric Error Surface Equation" The method can be summarized as follows: 1. Parametric error surface fitting is calculated only when the center position is the best cost position in a given iteration; 2. Using the center position cost and the costs at positions (-1,0), (0,-1), (1,0), and (0,1) from the center, in the following form: E(x,y)=A(x - x 0 ) 2 +B(y - y 0 ) 2 +C fit the 2D parabolic error surface equation, where (x 0 ,y 0 ) corresponds to the position of the minimum cost and C corresponds to the minimum cost value. By solving five equations for five unknowns, (x 0 ,y 0 ) is: x 0 =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y 0=(E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0)) is calculated as. (x 0 , y 0 ) can be calculated to any required sub - pixel accuracy by adjusting the precision at which the division is performed (i.e., how many bits of quotient are calculated). For 1 / 16 pel accuracy, only 4 bits in the absolute value of the quotient need to be calculated, which is conducive to a fast shift - subtraction - based implementation of the two divisions required per CU; 3. Add the calculated (x 0 , y 0 ) to the integer - distance refined MV to obtain a sub - pixel - accuracy delta MV.
[0178] 2.3 Intra - block Copy In High Efficiency Video Coding (HEVC) Screen Content Coding Extension (HEVC-SCC) and the current Versatile Video Coding (VVC) Test Model (VTM-4.0), Intra Block Copy (IBC), also known as current picture reference, is adopted. IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. As shown in Fig. 21, when IBC is applied, the current block is predicted by a reference block within the same picture. Before the current block is encoded or decoded, the samples within the reference block must have been already reconstructed. IBC is not very efficient for most camera capture sequences, but shows significant coding gain for screen content. The reason is that there are many repeating patterns such as icons and text characters in screen content pictures. IBC can effectively remove the redundancy between those repeating patterns. In HEVC-SCC, IBC can be applied when an inter-coded Coding Unit (CU) selects the current picture as its reference picture. In this case, the Motion Vector (MV) is renamed as Block Vector (BV), and the BV always has integer pixel accuracy. To be compatible with the main profile of HEVC, the current picture is marked as a "long-term" reference picture in the Decoded Picture Buffer (DPB). Similarly, in the multi-view / 3D video coding standard, the inter-view reference picture is also marked as a "long-term" reference picture.
[0179] After the BV finds its reference block, prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Then, transformation and quantization can be applied as in other coding modes.
[0180] However, if the reference block is outside the picture, overlaps with the current block, is outside the reconstructed area, or is outside the valid area restricted by some constraints, some or all of the pixel values are not determined. Basically, there are two solutions to address such problems. One is, for example, in bitstream compliance, not to allow such situations. The other is to apply padding to those undetermined pixel values. In the following subsections, those solutions will be described in detail.
[0181] 2.3.1 IBC in the VVC Test Model (VTM4.0) In the current VVC test model, i.e., the VTM-4.0 design, the entire reference block should be the current coding tree unit (CTU) and should not overlap with the current block. Therefore, there is no need to pad the reference block or the prediction block. The IBC flag is coded as the prediction mode of the current CU. Thus, for each CU, there are three prediction modes: MODE_INTRA, MODE_INTER, and MODE_IBC.
[0182] 2.3.1.1 IBC Merge Mode In the IBC merge mode, an index pointing to an entry in the IBC merge candidate list is syntax-analyzed from the bitstream. The construction of the IBC merge list can be summarized according to the following series of steps: · Step 1: Derivation of spatial candidates; · Step 2: Insertion of HMVP candidates; · Step 3: Insertion of pairwise average candidates; In the derivation of spatial merge candidates, as shown in Figure 2, A 1 , B 1 , B 0 , A 0 , and B 2 Among the candidates at the positions shown, up to four merge candidates are selected. The order of derivation is A 1 , B 1 , B 0 , A0 and B 2 is. Location B 2 is the location A 1 , B 1 , B 0 , A 0 is considered only when any of the PUs of A, B is not available (for example, because it belongs to another slice or tile) or is not encoded in IBC mode. Location A 1 After the candidates of A are added, the insertion of the remaining candidates is subjected to a redundancy check that ensures that candidates with the same motion information are excluded from the list so that the coding efficiency is improved.
[0183] After the insertion of the spatial candidates, if the size of the IBC merge list is still smaller than the maximum IBC merge list size, IBC candidates from the HMVP table can be inserted. A redundancy check is performed when inserting HMVP candidates.
[0184] Finally, the pairwise average candidates are inserted into the IBC merge list.
[0185] If the reference block specified by the merge candidate is outside the picture, overlaps with the current block, is outside the reconstructed area, or is outside the valid area restricted by some constraints, the merge candidate is called an invalid merge candidate.
[0186] Note that an invalid merge candidate may be inserted into the IBC merge list.
[0187] 2.3.1.2 IBC AMVP Mode In the IBC AMVP mode, the AMVP index pointing to the entry in the IBC AMVP list is syntax-analyzed from the bitstream. The construction of the IBC AMVP list can be summarized according to the following series of steps: · Step 1: Derivation of spatial candidates; - Examine A 0 , A 1 until available candidates are found; - Inspect B until possible candidates are found 0 B 1 B 2 for inspection; · Step 2: Insertion of HMVP candidates; · Step 3: Insertion of zero candidates; After the insertion of spatial candidates, if the size of the IBC AMVP list is still smaller than the maximum IBC AMVP list size, IBC candidates from the HMVP table can be inserted.
[0188] Finally, zero candidates are inserted into the IBC AMVP list.
[0189] 2.3.1.3 Chroma IBC mode In the current VVC, motion compensation in the chroma IBC mode is performed at the sub-block level. The chroma block will be divided into several sub-blocks. Each sub-block determines whether the corresponding luma block has a block vector and its validity if it exists. There are encoder constraints in the current VTM, and the chroma IBC mode will be tested when all sub-blocks within the current chroma CU have valid luma block vectors. For example, in a YUV 420 video, the chroma block is N×M, and the collocated luma region is 2N×2M. The sub-block size of the chroma block is 2×2. There are several steps for chroma mv derivation and the block copy process: 1) First, the chroma block is divided into (N>>1)*(M>>1) sub-blocks; 2) Each sub-block with the coordinates of the top-left sample as (x,y) fetches the corresponding luma block covering the same top-left sample with the coordinates of (2x,2y); 3) The encoder inspects the block vector (bv) of the fetched luma block. If any of the following conditions is met, the bv is considered invalid: a. The bv of the corresponding luma block does not exist; b. The predicted block specified by the bv has not been reconstructed yet; c. The predicted block specified by bv overlaps partially or completely with the current block; 4) The chroma motion vector of the sub-block is set to the motion vector of the corresponding luma sub-block.
[0190] When all sub-blocks find valid bv, the IBC mode is permitted in the encoder.
[0191] 2.3.2 Single BV List for IBC (in VTM5.0) JVET-N0843 is adopted in VVC. In JVET-N0843, the BV predictors for the merge mode and the AMVP mode in IBC share a common predictor list composed of the following elements: · Two spatial adjacent positions (A1, B1 as in Figure 2); · Five HMVP entries; · Default zero vector.
[0192] The number of candidates in the list is controlled by a variable derived from the slice header. In the merge mode, the first up to 6 entries of this list are used, and in the AMVP mode, the first 2 entries of this list are used. Also, this list conforms to the shared merge list area requirement (sharing the same list within the SMR).
[0193] In addition to the above BV predictor candidate list, JVET-N0843 also proposed to simplify the pruning process between the HMVP candidates and the existing merge candidates (A1, B1). In this simplification, only the first HMVP candidate is compared with the (one or more) spatial merge candidates, so there are at most 2 pruning processes.
[0194] 3. Problems The current design of the merge mode may have the following problems: 1. The normal merge list construction process depends on the use of TPM for the current block: a. In the TPM symbolization block, complete pruning is applied among the spatial merge candidates; b. In the non-TPM symbolization block, partial pruning is applied among the spatial merge candidates; 2. According to the current design, all merge-related tools (including IBC merge, normal merge, MMVD, sub-block merge, CIIP, TPM) are signaled by one flag named general_merge_flag. However, when this flag is true, it has become possible that all merge-related tools are signaled or derived to be disabled. It is not known how to handle this case. Also, turning off the merge mode is not allowed, that is, the maximum number of merge candidates is not considered to be equal to 0. However, in high-throughput encoders / decoders, it may be necessary to forcibly disable the merge mode; 3. The determination of the initial MV in the ATMVP process depends on the slice type, POC values of all reference pictures, collocated_from_l0_flag, etc., which delays the throughput of the MV; 4. The process of deriving the collocated MV depends on the use of sub-block technologies such as the conventional TMVP process or ATMVP process that require additional logic, for example; 5. For the sub-CU within the ATMVP symbolization block, even if the corresponding collocated block is inter-coded, the motion information of the sub-CU can be filled with other motion information instead of being derived from the corresponding collocated block. Such a design is sub-optimal for both coding efficiency and throughput.
[0195] 6. The HEVC specification currently determines the availability of one adjacent block within a picture or a reference picture based on whether the block is formed or within a different CTU row / slice, etc. However, in VVC, multiple coding methods have been introduced. It may be necessary to define different definitions for the availability of blocks.
[0196] 4. Examples of Technologies and Embodiments The following detailed enumerations should be regarded as examples for explaining general concepts. These embodiments should not be construed narrowly. Furthermore, these technologies can be combined in any way. For example, the embodiments described in this document are applicable to the Geometric Partitioning Mode (GPM) in which a current video block is divided into at least two non-rectangular sub-blocks. The non-rectangular blocks can have any geometric shape other than a rectangle. For example, GPM has dividing a first video block into a plurality of prediction partitions to which motion prediction is applied separately, and at least one partition has a non-rectangular shape. Also, the embodiments herein are described using examples of Alternative Temporal Motion Vector Prediction Coding (ATMVP), but in some embodiments, Sub-block-based Temporal Motion Vector Predictor Coding (SbTMVP) is also applicable.
[0197] Adjacent blocks denoted as A0, A1, B0, B1, B2, etc. are shown in FIG. 2: 1. The conventional merge and normal merge list construction process for TPM-encoded blocks is decoupled from the encoding method of the current block: a. In one example, when TPM is applied to one block, partial pruning is applied to the spatial merge candidates: i. In one example, whether two candidates are compared with each other is determined in the same way as that used for non-TPM merge-encoded blocks; ii. In one example, B1 is compared with A1, B0 is compared with B1, A0 is compared with A1, and B2 is compared with B1 and A1; iii. Alternatively, even if TPM is not used for one block, full pruning is applied to the spatial merge candidates; iv. Alternatively, furthermore, full pruning may be applied to some specific block dimensions: 1. For example, full pruning can be applied to block dimensions for which TPM is allowed; 2. For example, when the block size includes samples less than M*H, such as 16 or 32 or 64 luma samples for example, full pruning is not permitted; 3. For example, when the width of the block > th1 or >= th1, and / or the height of the block > th2 or >= th2, full pruning is not permitted; b. In one example, whether to check B2 is based on the number of available merge candidates before checking B2 when TPM is applied to one block: i. Alternatively, regardless of the number of available merge candidates for non-TPM merged coded blocks, B2 is always checked; 2. The initialization MV used to identify the block to determine whether ATMVP is available can simply rely on the list X information of spatially adjacent blocks (e.g., A1), where X is set to where the collocated picture used for temporal motion vector prediction is derived from (e.g., collocated_from_l0_flag): a. Alternatively, X is determined according to whether all reference pictures within all reference lists have POC values less than or equal to the POC value of the current picture: i. In one example, if it is true, X is set to 1. Otherwise, X is set to 0; b. Alternatively, when the reference picture related to the list X of spatially adjacent blocks (e.g., A1) is available and has the same POC value as the collocated picture, the initialization MV is set to the MV related to the list X of that spatially adjacent block. Otherwise, the default MV (e.g., (0,0)) is used; c. Alternatively, the motion information stored in the HMVP table may be used as the initialization MV in ATMVP: i. For example, the first available motion information stored in the HMVP table may be used; ii. For example, the first available motion information stored in the HMVP table associated with a specific reference picture (e.g., the collocated picture) may be used; d. Alternatively, X is a fixed number such as 0 or 1, for example; 3. The processes for deriving collocated MVs used for sub-block-based encoding tools and non-sub-block-based encoding tools can be made consistent, i.e., the process becomes independent of the use of a particular encoding tool: a. In one example, all or part of the process for deriving collocated MVs for sub-block-based encoding tools is made consistent with that used for TMVP: i. In one example, if it is a uni-prediction from list Y, the motion vectors of list Y are scaled to the target reference picture list X; ii. In one example, if it is a bi-prediction and the target reference picture list is X, the motion vectors of list Y are scaled to the target reference picture list X, and Y can be determined according to the following rules: - If none of the reference pictures have a POC greater than that of the current picture, or if all reference pictures have a POC value less than that of the current picture, Y is set equal to X; - Otherwise, Y is set equal to collocated_from_l0_flag; b. In one example, all or part of the process for deriving collocated MVs for TMVP is made consistent with that used for sub-block-based encoding tools; 4. The motion candidate list construction process (e.g., normal merge list, IBC merge / AMVP list) may depend on the block size and / or the merge sharing condition. Let the width and height of the block be denoted as W and H, respectively. Condition C may depend on W and H and / or the merge sharing condition: a. In one example, when condition C is satisfied, the derivation of spatial merge candidates is skipped; b. In one example, when condition C is satisfied, the derivation of HMVP candidates is skipped; c. In one example, when condition C is satisfied, the derivation of pairwise merge candidates is skipped; d. In one example, when condition C is satisfied, the maximum pruning process count is reduced or set to 0; e. In one example, condition C is satisfied when W*H is less than or equal to a threshold value (e.g., 64 or 32); f. In one example, condition C is satisfied when W and / or H is less than or equal to a threshold value (e.g., 4 or 8); g. In one example, condition C is satisfied when the current block is under a shared node; 5. The maximum number of allowed normal merge candidates / the maximum number of allowed IBC candidates / the maximum number of allowed sub-block merge candidates can be set to 0. Therefore, a specific tool can be disabled and there is no need to signal related syntax elements: a. In one example, when the maximum number of allowed normal merge candidates is equal to 0, the encoding tool relying on the normal merge list can be disabled. The encoding tool can be normal merge, MMVD, CIIP, TPM, DMVR, etc.; b. In one example, when the maximum number of allowed IBC candidates is equal to 0, IBC AMVP and / or IBC merge can be disabled; c. In one example, when the maximum number of allowed sub-block based merge candidates is equal to 0, sub-block based techniques such as ATMVP, affine merge mode, etc. can be disabled; d. When a tool is disabled according to the maximum number of allowed candidates, signaling of related syntax elements is skipped: i. Alternatively, further, signaling of merge-related tools may need to check whether the maximum number of allowed candidates is not equal to 0; ii. Alternatively, further, the call of the process related to the merge-related tool may need to check whether the maximum number of allowed candidates is not equal to 0; 6. The signaling of general_merge_flag and / or cu_skip_flag can depend on the maximum number of normal merge candidates allowed / the maximum number of IBC candidates allowed / the maximum number of sub-block merge candidates allowed / the use of merge-related coding tools: a. In one example, the merge-related coding tools may include IBC merge, normal merge, MMVD, sub-block merge, CIIP, TPM, DMVR, etc.; b. In one example, when the maximum number of normal merge candidates allowed, the maximum number of IBC merge / AMVP candidates allowed, and the maximum number of sub-block merge candidates allowed are equal to 0, general_merge_flag and / or cu_skip_flag are not signaled: i. Alternatively, furthermore, general_merge_flag and / or cu_skip_flag are presumed to be 0; 7. The compliant bitstream shall satisfy that at least one of the merge-related tools including IBC merge, normal merge, MMVD, sub-block merge, CIIP, TPM, DMVR, etc. is enabled when the general_merge_flag or cu_skip_flag of the current block is true; 8. The compliant bitstream shall satisfy that at least one of the merge-related tools including normal merge, MMVD, sub-block merge, CIIP, TPM, DMVR, etc. is enabled when (general_merge_flag or cu_skip_flag) of the current block is true and IBC is disabled for one slice / tile / brick / picture / current block; 9. The compliant bitstream shall satisfy that at least one of the merge-related tools including IBC merge, normal merge, sub-block merge, CIIP, TPM is enabled when (general_merge_flag or cu_skip_flag) of the current block is true and MMVD is disabled for one slice / tile / brick / picture / current block; 10. A compliant bitstream shall satisfy that when the (general_merge_flag or cu_skip_flag) of the current block is true and CIIP is disabled for one slice / tile / block / picture / current block, at least one of the merge-related tools including IBC merge, normal merge, MMVD, sub-block merge, and TPM is enabled; 11. A compliant bitstream shall satisfy that when the (general_merge_flag or cu_skip_flag) of the current block is true and TPM is disabled for one slice / tile / block / picture / current block, at least one of the merge-related tools including IBC merge, normal merge, MMVD, sub-block merge, and CIIP is enabled; 12. A compliant bitstream shall satisfy that when the general_merge_flag or cu_skip_flag of the current block is true, at least one of the enabled merge-related tools including IBC merge, normal merge, MMVD, sub-block merge, CIIP, and TPM is applied. When encoding the first block, the check for the availability of the second block can depend on the encoding mode information of the first block. For example, if different modes are used for the first block and the second block, the second block may be treated as unavailable regardless of the check result of other conditions (e.g., already constructed): a. In one example, when the first block is inter-encoded and the second block is IBC-encoded, the second block is marked as unavailable; b. In one example, when the first block is IBC-encoded and the second block is inter-encoded, the second block is marked as unavailable; c. When the second block is marked as unavailable, the related encoding information (e.g., motion information) is not allowed to be used for encoding the first block.
[0198] 5. Embodiment In addition to the latest VVC working draft (JVET-N1001_v7), the changes proposed are given below. The text to be deleted is marked with a font in bold uppercase letters. The newly added parts are emphasized with a bold italic font with underline.
[0199] 5.1 Embodiment #1 This embodiment aligns the pruning process for non-TPM encoded blocks with the pruning process for TPM encoded blocks, i.e., it performs a complete pruning operation on non-TPM encoded blocks.
[0200] 8.5.2.3 Derivation Process of Spatial Merge Candidates The inputs to this process are as follows: - The luma position (xCb, yCb) of the top-left sample of the current luma encoded block with respect to the top-left luma sample of the current picture; - A variable cbWidth that defines the width of the current encoded block within the luma sample; - A variable cbHeight that defines the height of the current encoded block within the luma sample.
[0201] The output of this process is X, which is 0 or 1, and is as follows: - The availability flag availableFlagA of the adjacent encoded unit 0 , availableFlagA 1 , availableFlagB 0 , availableFlagB 1 , and availableFlagB 2 ; - The reference index refIdxLXA of the adjacent encoded unit 0 , refIdxLXA 1 , refIdxLXB 0 , refIdxLXB 1 , and refIdxLXB 2 ; - The prediction list usage flag predFlagLXA of the adjacent encoded unit 0 , predFlagLXA 1, predFlagLXB 0 , predFlagLXB 1 , and predFlagLXB 2 ; - Motion vector mvLXA at 1 / 16 fractional sample accuracy of the neighboring symbolization unit 0 , mvLXA 1 , mvLXB 0 , mvLXB 1 , and mvLXB 2 ; - Dual prediction weight index gbiIdxA 0 , gbiIdxA 1 , gbiIdxB 0 , gbiIdxB 1 , and gbiIdxB 2 .
[0202] availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 For the derivation of, the following applies: - The luma position (xNbA 1 , yNbA 1 ) inside the neighboring luma coded block is set equal to (xCb - 1, yCb + cbHeight - 1); - The block availability derivation process defined in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the neighboring luma position (xNbA 1 , yNbA 1 ) as input, and its output is assigned to the block availability flag availableA 1 ; - The variables availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 are derived as follows: - availableA 1 is equal to FALSE, X which is 0 or 1, availableFlagA 1 is set equal to 0, mvLXA 1 both components of are set equal to 0, refIdxLXA 1 is set equal to -1, predFlagLXA 1 is set equal to 0, and, gbiIdxA 1 is set equal to 0; - Otherwise, availableFlagA 1 is set equal to 1, and the following assignments are made: mvLXA 1 =MvLX[xNbA 1 [yNbA 1 (8 - 294) refIdxLXA 1 =RefIdxLX[xNbA 1 [yNbA 1 (8 - 295) predFlagLXA 1 =PredFlagLX[xNbA 1 [yNbA 1 (8 - 296) gbiIdxA 1 =GbiIdx[xNbA 1 [yNbA 1 (8 - 297)
[0203] availableFlagB 1 、refIdxLXB 1 、predFlagLXB 1 、and mvLXB 1 For the derivation of, the following applies: - The luma position (xNbB 1 ,yNbB 1 ) inside the adjacent luma coded block is set equal to (xCb + cbWidth - 1,yCb - 1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) and the neighbouring luma position (xNbB 1 , yNbB 1 ) set equal to (xCb, yCb) as inputs, and its output is assigned to the block availability flag availableB 1 ; - The variables availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 are derived as follows: - If one or more of the following conditions are true, then X, which is 0 or 1, has availableFlagB 1 set equal to 0, both components of mvLXB 1 set equal to 0, refIdxLXB 1 set equal to -1, predFlagLXB 1 set equal to 0, and gbiIdxB 1 set equal to 0; - availableB 1 is equal to FALSE; - availableA 1 is equal to TRUE, and the luma positions (xNbA 1 , yNbA 1 ) and (xNbB 1 , yNbB 1 ) have the same motion vector and the same reference index; - Otherwise, availableFlagB 1 is set equal to 1 and the following assignments are made: mvLXB 1 = MvLX[xNbB 1 [yNbB 1 (8 - 298) refIdxLXB 1=RefIdxLX[xNbB 1 ][yNbB 1 ] (8-299) predFlagLXB 1 =PredFlagLX[xNbB 1 ][yNbB 1 ] (8-300) gbiIdxB 1 =GbiIdx[xNbB 1 ][yNbB 1 ] (8-301)
[0204] availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 To derive, the following applies: - The luma position inside the adjacent luma coding block (xNbB 0 ,yNbB 0 ) is set equal to (xCb+cbWidth,yCb-1); - The block availability derivation process specified in section 6.4.X [Ed.(BB):Neighbouring blocks availability checking process tbd] determines whether the current luma position (xCurr, yCurr) and the neighboring luma position (xNbB 0 ,yNbB 0 ) and outputs the block availability flag availableB. 0 assigned to; - Variable availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 is derived as follows: - availableFlagB, with X being 0 or 1 if one or more of the following conditions are true: 0 is set equal to 0, and mvLXB 0 Both components of refIdxLXB are set equal to 0.0 is set equal to -1, and predFlagLXB 0 is set equal to 0, and gbiIdxB 0 is set equal to 0; (Outer 2) TIFF0007683985000013.tif34170 - Otherwise, availableFlagB 0 is set equal to 1, and the following assignments are made: mvLXB 0 =MvLX[xNbB 0 [yNbB 0 (8 - 302) refIdxLXB 0 =RefIdxLX[xNbB 0 [yNbB 0 (8 - 303) predFlagLXB 0 =PredFlagLX[xNbB 0 [yNbB 0 (8 - 304) gbiIdxB 0 =GbiIdx[xNbB 0 [yNbB 0 (8 - 305)
[0205] availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 For the derivation of, the following applies: - The luma position (xNbA 0 ,yNbA 0 ) inside the adjacent luma - coded block is set equal to (xCb - 1,yCb + cbWidth); - The block availability derivation process specified in Section 6.4.X [Ed.(BB):Neighbouring blocks availability checking process tbd] is applied to the current luma position (xCurr,yCurr) set equal to (xCb,yCb) and the adjacent luma position (xNbA 0 ,yNbA0 ) is called with the input, and its output is assigned to the block availability flag availableA 0 ; - The variable availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 are derived as follows: - If one or more of the following conditions are true, X, which is 0 or 1, sets availableFlagA 0 equal to 0, sets both components of mvLXA 0 equal to 0, sets refIdxLXA 0 equal to -1, sets predFlagLXA 0 equal to 0, and sets gbiIdxA 0 equal to 0; (Outer 3) TIFF0007683985000014.tif47170 - Otherwise, availableFlagA 0 is set equal to 1 and the following assignments are made: mvLXA 0 = MvLX[xNbA 0 [yNbA 0 (8 - 306) refIdxLXA 0 = RefIdxLX[xNbA 0 [yNbA 0 (8 - 307) predFlagLXA 0 = PredFlagLX[xNbA 0 [yNbA 0 (8 - 308) gbiIdxA 0 = GbiIdx[xNbA 0 [yNbA 0 (8 - 309)
[0206] availableFlagB 2 , refIdxLXB 2, predFlagLXB 2 , and mvLXB 2 For the derivation of, the following applies: - The luma position (xNbB 2 , yNbB 2 ) inside the adjacent luma-encoded block is set equal to (xCb - 1, yCb - 1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbB 2 , yNbB 2 ) as input, and its output is assigned to the block availability flag availableB 2 ; - The variables availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 are derived as follows: - If one or more of the following conditions are true, X, which is 0 or 1, sets availableFlagB 2 equal to 0, sets both components of mvLXB 2 equal to 0, sets refIdxLXB 2 equal to -1, sets predFlagLXB 2 equal to 0, and sets gbiIdxB 2 equal to 0; (Outer 4) TIFF0007683985000015.tif69170 - Otherwise, availableFlagB 2 is set equal to 1 and the following assignments are made: mvLXB 2 = MvLX[xNbB 2 [yNbB 2 (8 - 310) refIdxLXB 2=RefIdxLX[xNbB 2 [yNbB 2 (8-311) predFlagLXB 2 =PredFlagLX[xNbB 2 [yNbB 2 (8-312) gbiIdxB 2 =GbiIdx[xNbB 2 [yNbB 2 (8-313)
[0207] 5.2 Embodiment #2 This embodiment aligns the pruning process for the TPM-encoded block with the pruning process for the non-TPM-encoded block, that is, a limited pruning operation is performed on the TPM-encoded block.
[0208] 8.5.2.3 Derivation Process of Spatial Merge Candidates The inputs to this process are as follows: - The luma position (xCb, yCb) of the top-left sample of the current luma-encoded block with respect to the top-left luma sample of the current picture; - A variable cbWidth that defines the width of the current encoded block within the luma sample; - A variable cbHeight that defines the height of the current encoded block within the luma sample.
[0209] The output of this process is an X which is 0 or 1 and is as follows: - The availability flags availableFlagA of adjacent encoded units 0 , availableFlagA 1 , availableFlagB 0 , availableFlagB 1 , and availableFlagB 2 ; - The reference indices refIdxLXA of adjacent encoded units 0 , refIdxLXA 1 , refIdxLXB 0 , refIdxLXB1 and refIdxLXB 2 ; - Prediction list usage flag predFlagLXA of adjacent coding units 0 , predFlagLXA 1 , predFlagLXB 0 , predFlagLXB 1 , and predFlagLXB 2 ; - Motion vector mvLXA at 1 / 16 fractional sample precision of adjacent coding units 0 , mvLXA 1 , mvLXB 0 , mvLXB 1 , and mvLXB 2 ; - Dual-prediction weight index gbiIdxA 0 , gbiIdxA 1 , gbiIdxB 0 , gbiIdxB 1 , and gbiIdxB 2 。
[0210] availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 For the derivation of, the following applies: - The luma position (xNbA 1 , yNbA 1 ) inside the adjacent luma-coded block is set equal to (xCb - 1, yCb + cbHeight - 1); - The block availability derivation process defined in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbA 1 , yNbA 1 ) as input, and its output is assigned to the block availability flag availableA 1 ; - variable availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 are derived as follows: - When availableA 1 is equal to FALSE, X which is 0 or 1, availableFlagA 1 is set equal to 0, both components of mvLXA 1 are set equal to 0, refIdxLXA 1 is set equal to -1, predFlagLXA 1 is set equal to 0, and gbiIdxA 1 is set equal to 0; - Otherwise, availableFlagA 1 is set equal to 1 and the following assignments are made: mvLXA 1 = MvLX[xNbA 1 [yNbA 1 (8 - 294) refIdxLXA 1 = RefIdxLX[xNbA 1 [yNbA 1 (8 - 295) predFlagLXA 1 = PredFlagLX[xNbA 1 [yNbA 1 (8 - 296) gbiIdxA 1 = GbiIdx[xNbA 1 [yNbA 1 (8 - 297)
[0211] For the derivation of availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 the following applies: - The luma position (xNbB 1 , yNbB1 ) is set equal to (xCb + cbWidth - 1, yCb - 1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the neighbouring luma position (xNbB 1 , yNbB 1 ) as input, and its output is assigned to the block availability flag availableB 1 ; - The variables availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 are derived as follows: - If one or more of the following conditions are true, then X, which is 0 or 1, sets availableFlagB 1 equal to 0, sets both components of mvLXB 1 equal to 0, sets refIdxLXB 1 equal to -1, sets predFlagLXB 1 equal to 0, and sets gbiIdxB 1 equal to 0; - availableB 1 is equal to FALSE; - availableA 1 is equal to TRUE and the luma positions (xNbA 1 , yNbA 1 ) and (xNbB 1 , yNbB 1 ) have the same motion vector and the same reference index; - Otherwise, availableFlagB 1 is set equal to 1 and the following assignments are made: mvLXB 1 = MvLX[xNbB 1 [yNbB1 ] (8-298) refIdxLXB 1 =RefIdxLX[xNbB 1 ][yNbB 1 ] (8-299) predFlagLXB 1 =PredFlagLX[xNbB 1 ][yNbB 1 ] (8-300) gbiIdxB 1 =GbiIdx[xNbB 1 ][yNbB 1 ] (8-301)
[0212] availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 To derive, the following applies: - The luma position inside the adjacent luma coding block (xNbB 0 ,yNbB 0 ) is set equal to (xCb+cbWidth,yCb-1); - The block availability derivation process specified in section 6.4.X [Ed.(BB):Neighbouring blocks availability checking process tbd] determines whether the current luma position (xCurr, yCurr) and the neighboring luma position (xNbB 0 ,yNbB 0 ) and outputs the block availability flag availableB. 0 assigned to; - Variable availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 is derived as follows: - availableFlagB, with X being 0 or 1 if one or more of the following conditions are true: 0is set equal to 0, mvLXB 0 both components of 0 are set equal to 0, refIdxLXB 0 is set equal to -1, predFlagLXB 0 is set equal to 0, and gbiIdxB 0 is set equal to 0; (Outer 5) TIFF0007683985000016.tif34170 - Otherwise, availableFlagB 0 is set equal to 1 and the following assignments are made: mvLXB 0 = MvLX[xNbB 0 [yNbB 0 (8 - 302) refIdxLXB 0 = RefIdxLX[xNbB 0 [yNbB 0 (8 - 303) predFlagLXB 0 = PredFlagLX[xNbB 0 [yNbB 0 (8 - 304) gbiIdxB 0 = GbiIdx[xNbB 0 [yNbB 0 (8 - 305)
[0213] availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 For the derivation of - The luma position (xNbA 0 , yNbA 0 ) inside the adjacent luma coded block is set equal to (xCb - 1, yCb + cbWidth); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) and the neighbouring luma position (xNbA 0 , yNbA 0 ) set equal to (xCb, yCb), and its output is assigned to the block availability flag availableA 0 ; - The variables availableFlagA 0 , refIdxLXA 0 , predFlagLXA 0 , and mvLXA 0 are derived as follows: - If one or more of the following conditions are true, then X, which is 0 or 1, sets availableFlagA 0 equal to 0, sets both components of mvLXA 0 equal to 0, sets refIdxLXA 0 equal to -1, sets predFlagLXA 0 equal to 0, and sets gbiIdxA 0 equal to 0; (Outside 6) TIFF0007683985000017.tif46170 - Otherwise, availableFlagA 0 is set equal to 1 and the following assignments are made: mvLXA 0 = MvLX[xNbA 0 [yNbA 0 (8 - 306) refIdxLXA 0 = RefIdxLX[xNbA 0 [yNbA 0 (8 - 307) predFlagLXA 0 = PredFlagLX[xNbA 0 [yNbA 0 (8 - 308) gbiIdxA 0 =GbiIdx[xNbA 0 [yNbA 0 (8-309)
[0214] availableFlagB 2 、refIdxLXB 2 、predFlagLXB 2 、and mvLXB 2 for the derivation of, the following applies: - The luma position (xNbB 2 , yNbB 2 ) inside the adjacent luma coded block is set equal to (xCb-1, yCb-1); - The block availability derivation process specified in Section 6.4.X [Ed.(BB):Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbB 2 , yNbB 2 ) as input, and its output is assigned to the block availability flag availableB 2 ; - The variables availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 are derived as follows: - If one or more of the following conditions are true, X which is 0 or 1, availableFlagB 2 is set equal to 0, both components of mvLXB 2 are set equal to 0, refIdxLXB 2 is set equal to -1, predFlagLXB 2 is set equal to 0, and gbiIdxB 2 is set equal to 0; (Outer 7) TIFF0007683985000018.tif69170 - Otherwise, availableFlagB 2 is set equal to 1 and the following assignments are made: mvLXB 2 =MvLX[xNbB 2 [yNbB 2 (8 - 310) refIdxLXB 2 =RefIdxLX[xNbB 2 [yNbB 2 (8 - 311) predFlagLXB 2 =PredFlagLX[xNbB 2 [yNbB 2 (8 - 312) gbiIdxB 2 =GbiIdx[xNbB 2 [yNbB 2 (8 - 313)
[0215] 5.3 Embodiment #3 This embodiment arranges the conditions for calling the inspection of B2.
[0216] 8.5.2.3 Derivation Process of Spatial Merge Candidates The inputs to this process are as follows: - The luma position (xCb, yCb) of the top - left sample of the current luma - coded block with respect to the top - left luma sample of the current picture; - A variable cbWidth that defines the width of the current coded block within the luma sample; - A variable cbHeight that defines the height of the current coded block within the luma sample.
[0217] The output of this process is an X which is 0 or 1 and is as follows: - The availability flag availableFlagA of the adjacent coded unit 0 , availableFlagA 1 , availableFlagB 0 , availableFlagB1 and availableFlagB 2 ; - Reference index refIdxLXA of the adjacent coding unit 0 refIdxLXA 1 refIdxLXB 0 refIdxLXB 1 and refIdxLXB 2 ; - Prediction list usage flag predFlagLXA of the adjacent coding unit 0 predFlagLXA 1 predFlagLXB 0 predFlagLXB 1 and predFlagLXB 2 ; - Motion vector mvLXA at 1 / 16 fractional sample accuracy of the adjacent coding unit 0 mvLXA 1 mvLXB 0 mvLXB 1 and mvLXB 2 ; - Dual prediction weight index gbiIdxA 0 gbiIdxA 1 gbiIdxB 0 gbiIdxB 1 and gbiIdxB 2 .
[0218] availableFlagA 1 refIdxLXA 1 predFlagLXA 1 and mvLXA 1 For the derivation of, the following applies: - The luma position (xNbA 1 , yNbA 1 ) inside the adjacent luma coding block is set equal to (xCb - 1, yCb + cbHeight - 1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the neighbouring luma position (xNbA 1 , yNbA 1 ) as inputs, and its output is assigned to the block availability flag availableA 1 ; - The variables availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 are derived as follows: ...
[0219] For the derivation of availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 the following applies: - The luma position (xNbB 1 , yNbB 1 ) inside the neighbouring luma coded block is set equal to (xCb + cbWidth - 1, yCb - 1); ...
[0220] For the derivation of availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 the following applies: - The luma position (xNbB 0 , yNbB 0 ) inside the neighbouring luma coded block is set equal to (xCb + cbWidth, yCb - 1); ...
[0221] availableFlagA 0 , refIdxLXA0 , predFlagLXA 0 , and mvLXA 0 For the derivation of, the following applies: - The luma position (xNbA 0 , yNbA 0 ) inside the adjacent luma - coded block is set equal to (xCb - 1, yCb+cbWidth); ...
[0222] availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 For the derivation of, the following applies: - The luma position (xNbB 2 , yNbB 2 ) inside the adjacent luma - coded block is set equal to (xCb - 1, yCb - 1); - The block availability derivation process defined in item 6.4.X [Ed.(BB):Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbB 2 , yNbB 2 ) as input, and its output is assigned to the block availability flag availableB 2 ; - The variables availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 are derived as follows: - If one or more of the following conditions are true, then X, which is 0 or 1, is such that availableFlagB 2 is set equal to 0, both components of mvLXB 2 are set equal to 0, refIdxLXB 2 is set equal to - 1, and predFlagLXB 2is set to 0, and gbiIdxB 2 is set to 0; (Outer 8) TIFF0007683985000019.tif68170 - Otherwise, availableFlagB 2 is set to 1, and the following assignments are made: mvLXB 2 =MvLX[xNbB 2 [yNbB 2 (8 - 310) refIdxLXB 2 =RefIdxLX[xNbB 2 [yNbB 2 (8 - 311) predFlagLXB 2 =PredFlagLX[xNbB 2 [yNbB 2 (8 - 312) gbiIdxB 2 =GbiIdx[xNbB 2 [yNbB 2 (8 - 313)
[0223] 5.4 Embodiment #4 This embodiment arranges the conditions for calling the inspection of B2.
[0224] 8.5.2.3 Derivation Process of Spatial Merge Candidates The inputs to this process are as follows: - The luma position (xCb, yCb) of the top - left sample of the current luma - coded block with respect to the top - left luma sample of the current picture; - A variable cbWidth that defines the width of the current coded block within the luma sample; - A variable cbHeight that defines the height of the current coded block within the luma sample.
[0225] The output of this process is X, which is 0 or 1, and is as follows: - The availability flag availableFlagA of the adjacent coded unit 0, availableFlagA 1 , availableFlagB 0 , availableFlagB 1 , and availableFlagB 2 ; - Reference index refIdxLXA of the adjacent coding unit 0 , refIdxLXA 1 , refIdxLXB 0 , refIdxLXB 1 , and refIdxLXB 2 ; - Prediction list usage flag predFlagLXA of the adjacent coding unit 0 , predFlagLXA 1 , predFlagLXB 0 , predFlagLXB 1 , and predFlagLXB 2 ; - Motion vector mvLXA at 1 / 16 fractional sample accuracy of the adjacent coding unit 0 , mvLXA 1 , mvLXB 0 , mvLXB 1 , and mvLXB 2 ; - Dual prediction weight index gbiIdxA 0 , gbiIdxA 1 , gbiIdxB 0 , gbiIdxB 1 , and gbiIdxB 2 .
[0226] availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 For the derivation of, the following applies: - The luma position (xNbA 1 , yNbA 1 ) inside the adjacent luma coding block is set equal to (xCb - 1, yCb + cbHeight - 1); - The block availability derivation process specified in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the neighbouring luma position (xNbA 1 , yNbA 1 ) as inputs, and its output is assigned to the block availability flag availableA 1 ; - The variables availableFlagA 1 , refIdxLXA 1 , predFlagLXA 1 , and mvLXA 1 are derived as follows: ...
[0227] For the derivation of availableFlagB 1 , refIdxLXB 1 , predFlagLXB 1 , and mvLXB 1 the following applies: - The luma position (xNbB 1 , yNbB 1 ) inside the neighbouring luma coded block is set equal to (xCb + cbWidth - 1, yCb - 1); ...
[0228] For the derivation of availableFlagB 0 , refIdxLXB 0 , predFlagLXB 0 , and mvLXB 0 the following applies: - The luma position (xNbB 0 , yNbB 0 ) inside the neighbouring luma coded block is set equal to (xCb + cbWidth, yCb - 1); ...
[0229] availableFlagA 0 , refIdxLXA0 , predFlagLXA 0 , and mvLXA 0 For the derivation of, the following applies: - The luma position (xNbA 0 , yNbA 0 ) inside the adjacent luma encoded block is set equal to (xCb - 1, yCb + cbWidth); ...
[0230] availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 For the derivation of, the following applies: - The luma position (xNbB 2 , yNbB 2 ) inside the adjacent luma encoded block is set equal to (xCb - 1, yCb - 1); - The block availability derivation process defined in item 6.4.X [Ed.(BB): Neighbouring blocks availability checking process tbd] is called with the current luma position (xCurr, yCurr) equal to (xCb, yCb) and the adjacent luma position (xNbB 2 , yNbB 2 ) as input, and its output is assigned to the block availability flag availableB 2 ; - The variables availableFlagB 2 , refIdxLXB 2 , predFlagLXB 2 , and mvLXB 2 are derived as follows: - If one or more of the following conditions are true, X is 0 or 1, availableFlagB 2 is set equal to 0, both components of mvLXB 2 are set equal to 0, refIdxLXB 2 is set equal to -1, and predFlagLXB 2is set equal to 0, and, gbiIdxB 2 is set equal to 0; (Outer 9) TIFF0007683985000020.tif69170 - Otherwise, availableFlagB 2 is set equal to 1 and the following assignments are made: mvLXB 2 =MvLX[xNbB 2 [yNbB 2 (8 - 310) refIdxLXB 2 =RefIdxLX[xNbB 2 [yNbB 2 (8 - 311) predFlagLXB 2 =PredFlagLX[xNbB 2 [yNbB 2 (8 - 312) gbiIdxB 2 =GbiIdx[xNbB 2 [yNbB 2 (8 - 313)
[0231] 5.5 Embodiment #5 This embodiment simplifies the determination of the initialization MV in the ATMVP process.
[0232] 8.5.5.4 Derivation Process of Sub - block - based Time - merge - based Motion Data The inputs to this process are as follows: - The position (xCtb, yCtb) of the top - left sample of the luma - coded tree block containing the current coded block; - The position (xColCtrCb, yColCtrCb) of the top - left sample of the collocated luma - coded block covering the bottom - right central sample; - The availability flag availableFlagA of the adjacent coded unit 1 ; - The reference index refIdxLXA of the adjacent coded unit 1 ; - Prediction list usage flag predFlagLXA of adjacent symbolization unit 1 ; - Motion vector mvLXA of adjacent symbolization unit with 1 / 16 fractional sample precision 1 。
[0233] The output of this process is as follows: - Motion vectors ctrMvL0 and ctrMvL1; - Prediction list usage flags ctrPredFlagL0 and ctrPredFlagL1; - Temporal motion vector tempMv.
[0234] The variable tempMv is set as follows: tempMv[0]=0 (8-529) tempMv[1]=0 (8-530) The variable currPic defines the current picture.
[0235] availableFlagA 1 When it is equal to TRUE, the following applies: (Outer 10) TIFF0007683985000021.tif89170
[0236] The position (xColCb, yColCb) of the collocate block inside ColPic is derived as follows: xColCb=Clip3(xCtb, Min(CurPicWidthInSamplesY-1,xCtb+(1<<CtbLog2SizeY)+3), xColCtrCb+(tempMv[0]>>4)) (8-531) yColCb=Clip3(yCtb, Min(CurPicHeightInSamplesY-1,yCtb+(1<<CtbLog2SizeY)-1), yColCtrCb+(tempMv[1]>>4)) (8-532)
[0237] The array colPredMode is set equal to the prediction mode array CuPredMode of the collocated picture defined by ColPic.
[0238] The motion vectors ctrMvL0 and ctrMvL1, and the prediction list utilization flags ctrPredFlagL0 and ctrPredFlagL1 are derived as follows: ...
[0239] 5.6 Embodiment #6 An example of aligning the collocated MV derivation process in a sub-block-based method and a non-sub-block-based method.
[0240] 8.5.2.12 Collocated Motion Vector Derivation Process The inputs to this process are as follows: - The variable currCb that defines the current coding block; - The variable colCb that defines the collocated coding block inside the collocated picture defined by ColPic; - The luma position (xColCb, yColCb) that defines the top-left sample of the collocated coding block defined by colCb with respect to the top-left luma sample of the collocated picture defined by ColPic; - The reference index refIdxLX at X which is 0 or 1; - The flag sbFlag that indicates the sub-block temporal merge candidate.
[0241] The outputs of this process are as follows: - The motion vector prediction mvLXCol with 1 / 16 fractional sample accuracy; - The availability flag availableFlagLXCol.
[0242] The variable currPic defines the current picture.
[0243] The arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] are respectively set equal to PredFlagL0[x][y], MvDmvrL0[x][y], and RefIdxL0[x][y] of the collocated picture defined by ColPic, and the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] are respectively set equal to PredFlagL1[x][y], MvDmvrL1[x][y], and RefIdxL1[x][y] of the collocated picture defined by ColPic.
[0244] The variables mvLXCol and availableFlagLXCol are derived as follows: - When colCb is encoded in the intra or IBC prediction mode, both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0; - Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows: (Outer 11) TIFF0007683985000022.tif82170(Outer 12) TIFF0007683985000023.tif79170 - When availableFlagLXCol is equal to TRUE, mvLXCol and availableFlagLXCol are derived as follows: - When LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic(ColPic, colCb, refIdxCol, listCol), both components of mvLXCol are set equal to 0, and availableFlagLXCol is set equal to 0; - Otherwise, the variable availableFlagLXCol is set equal to 1, and refPicList[listCol][refIdxCol] is set to be a picture having a reference index refIdxCol within the reference picture list listCol of the slice that includes the coded block colCb within the collocated picture defined by ColPic, and the following applies: ...
[0245] FIG. 22 is a block diagram of a video processing apparatus 2200. The apparatus 2200 may be used to implement one or more of the methods described herein. The apparatus 2200 may be embodied in a smartphone, a tablet, a computer, or a mono Internet of Things (IoT) receiver. The apparatus 2200 may include one or more processors 2202, one or more memories 2204, and video processing hardware 2206. The (one or more) processors 2202 may be configured to execute one or more of the methods described in this document. The (one or more) memories 2204 may be used to store data and code used to execute the methods and techniques described herein. The video processing hardware 2206 may be used to implement some of the techniques described in this document in hardware circuits. The video processing hardware 2206 may be partially or fully included within the (one or more) processors 2202 in the form of dedicated hardware or a graphical processing unit (GPU) or a specialized signal processing block.
[0246] Here, some embodiments of this document are presented in a clause-based format.
[0247] Some example embodiments of the techniques described in item 1 of Section 4 include the following.
[0248] 1. A method of video processing (e.g., method 2300 shown in FIG. 23), the method comprising the step (2302) of applying a pruning process to the construction of a merge list of a current video block that is divided using a triangular partitioning mode (TMP) in which the current video block is divided into at least two non-rectangular sub-blocks, wherein the pruning process is the same as another pruning process of another video block divided using a non-TMP partition; and the step of performing a conversion between the current video block and a bitstream representation of the current video block based on the construction of the merge list (2304).
[0249] 2. The method according to item 1, wherein the pruning process uses partial pruning for spatial merge candidates of the current video block.
[0250] 3. The method according to item 1, wherein the pruning process applies full pruning or partial pruning to the current video block based on a block dimension rule that defines whether to use full pruning or partial pruning based on the dimensions of the current video block.
[0251] 4. The method according to item 1, wherein the pruning process uses adjacent blocks in different orders in the merge list construction process.
[0252] Some embodiments of the technology described in item 2 of Section 4 include the following.
[0253] 1. A method of video processing, the method comprising the step of determining, for the availability of an alternative temporal motion vector predictor coding (ATMVP) mode for the conversion between a current video block and a bitstream representation of the current video block, based on a list X of adjacent blocks of the current video block, where X is an integer and the value of X depends on the coding conditions of the current video block; and the step of performing the conversion based on the availability of the ATMVP mode.
[0254] 2. The method according to claim 1, wherein X indicates the position of a collocated video picture from which the temporal motion vector prediction used in the conversion between the current video block and the bitstream representation is performed.
[0255] 3. The method according to claim 1, wherein X is determined by comparing the picture order count (POC) of all reference pictures in all reference lists for the current video block with the POC of the current video picture of the current video block.
[0256] 4. The method according to claim 3, wherein when the comparison indicates that the POC is less than or equal to the POC of the current video block, X is set to 1, and otherwise X is set to 0.
[0257] 5. The method according to claim 1, wherein motion information stored in a history-based motion vector predictor table is used to initialize motion information in the ATMVP mode.
[0258] Some embodiments of the technology described in item 3 of Section 4 include the following.
[0259] 1. A video processing method, comprising: determining, in a conversion between a current video block and a bitstream representation of the current video block, that a sub-block based coding technique in which the current video block is divided into at least two sub-blocks capable of deriving their own motion information is used in the conversion; and performing the conversion using a merge list construction process for the current video block aligned with a block-based derivation process for a collocated motion vector.
[0260] 2. The method according to claim 1, wherein the merge list construction process and the derivation process perform a single prediction from list Y, and the motion vector of list Y is scaled to a target reference picture list X.
[0261] 3. The merge list construction process and the derivation process have double prediction performed using a target reference picture list X, and the motion vectors of list Y are scaled to the motion vectors of list X, where Y is determined according to rules, the method according to item 1.
[0262] Some embodiments of the technology described in item 4 of Section 4 include the following.
[0263] 1. A method for video processing, comprising determining between a condition being satisfied and the condition not being satisfied based on the size of a current video block of a video picture and / or activation of a merge sharing state in which merge candidates from different coding tools are shared, and performing a conversion between the current video block and a bitstream representation of the current video block based on the condition.
[0264] 2. The method according to item 1, wherein the step of performing the conversion includes skipping deriving a spatial merge candidate when the condition is satisfied.
[0265] 3. The method according to item 1, wherein the step of performing the conversion includes skipping deriving a history-based motion vector candidate when the condition is satisfied.
[0266] 4. The method according to any one of items 1 to 3, wherein the determination that the condition is satisfied is based on the current video block being under a shared node within the video picture.
[0267] Some embodiments of the technology described in item 5 of Section 4 include the following.
[0268] 1. A method for video processing, comprising: determining, in a conversion between a current video block and a bitstream representation of the current video block, whether an encoding tool is to be disabled for the conversion, wherein the bitstream representation is configured to provide an indication that a maximum number of merge candidates for the encoding tool is zero; and performing the conversion between the current video block and the bitstream representation of the current video block using the determination that the encoding tool is to be disabled.
[0269] 2. The method according to claim 1, wherein the encoding tool corresponds to an intra block copy in which pixels of the current video block are encoded from other pixels within a video region of the current video block.
[0270] 3. The method according to claim 1, wherein the encoding tool is a sub-block encoding tool.
[0271] 4. The method according to claim 3, wherein the sub-block encoding tool is an affine encoding tool or an alternative motion vector predictor tool.
[0272] 5. The method according to any one of claims 1 to 4, wherein performing the conversion includes processing the bitstream by skipping syntax elements related to the encoding tool.
[0273] Some embodiments of the technology described in item 6 of Section 4 include the following.
[0274] 1. A method for video processing, comprising: determining, in a conversion between a current video block and a bitstream representation of the current video block, using a rule, wherein the rule defines that a first syntax element in the bitstream representation conditionally exists based on a second syntax element indicating a maximum number of merge candidates used by an encoding tool used in the conversion; and performing the conversion between the current video block and the bitstream representation of the current video block based on the determination.
[0275] 2. The method according to claim 1, wherein the first syntax element corresponds to a merge flag.
[0276] 3. The method according to claim 1, wherein the first syntax element corresponds to a skip flag.
[0277] 4. The method according to any one of claims 1 to 3, wherein the encoding tool is a sub-band encoding tool, and the second syntax element corresponds to a maximum allowable merge candidate for the sub-band encoding tool.
[0278] 34. The method according to any one of claims 1 to 33, wherein the conversion includes generating the bitstream representation from the current video block.
[0279] 35. The method according to any one of claims 1 to 33, wherein the conversion includes generating samples of the current video block from the bitstream representation.
[0280] 36. A video processing apparatus having a processor configured to implement the method according to any one of claims 1 to 35.
[0281] 37. A computer-readable medium storing code, which causes a processor to execute the method according to any one of claims 1 to 35 when executed.
[0282] FIG. 24 is a block diagram showing an example of a video processing system 2400 in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 2400. System 2400 may include an input 2402 for receiving video content. The video content may be received in a raw or uncompressed format, such as, for example, 8-bit or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 2402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet®, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi® or cellular interfaces.
[0283] System 2400 may include an encoding component 2404 that may implement various coding or encoding methods described in this document. Encoding component 2404 may reduce the average bit rate of the video from input 2402 to the output of encoding component 2404 to generate an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression techniques or video transcoding techniques. The output of encoding component 2404 may be stored or connected as represented by component 2406 and transmitted via communication. The stored or communicated bitstream (or encoded) representation of the video received at input 2402 may be used by component 2408 to generate pixel values or a displayable video that is sent to display interface 2410. The process of generating a video that a user can view from the bitstream representation is sometimes referred to as video decompression. Also, although certain video processing operations may be referred to as "encoding" operations or tools, it is understood that encoding tools or operations are used in an encoder and corresponding decoding tools or operations that reverse the encoding result are executed in a decoder.
[0284] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI (registered trademark)), a DisplayPort, and the like. Examples of a storage interface include SATA (serial advanced technology attachment), PCI, an IDE interface, and the like. The technology described in this document may be embodied in various electronic devices such as, for example, a mobile phone, a laptop, a smartphone, or other devices capable of performing digital data processing and / or video display.
[0285] FIG. 25 is a flowchart of an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 1 of Section 4 of this document. In step 2502, the process determines that a first video block of visual media data uses a Geometric Partitioning Mode (GPM) and a second video block of visual media data uses a non-GPM mode. In step 2504, the process constructs a first merge list for the first video block and a second merge list for the second video block based on a unified pruning process, the first merge list and the second merge list having merge candidates, and the unified pruning process including adding a new merge candidate to the merge list based on comparing motion information of the new merge candidate with motion information of at least one merge candidate in the merge list, GPM having dividing the first video block into a plurality of prediction partitions to which motion prediction is separately applied, and at least one partition having a non-rectangular shape.
[0286] FIG. 26 is a flowchart of an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 2 of Section 4 of this document. At step 2602, the process determines, based on a rule, an initial value of motion information applicable to the current video block for conversion between the current video block of the visual media data and the bitstream representation of the visual media data, the rule specifying to examine whether a sub-block based temporal motion vector predictor coding (SbTMVP) mode is available for the current video block based on a reference list (denoted as list X) of adjacent blocks of the current video block, where X is an integer and the value of X depends at least on the coding conditions of the adjacent blocks. At step 2604, the process performs the conversion based on the determination.
[0287] FIG. 27 is a flowchart of an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 3 of Section 4 of this document. At step 2702, the process derives, based on a rule, one or more collocated motion vectors for one or more sub-blocks of the current video block for conversion between the current video block of the visual media data and the bitstream representation of the visual media data, the rule specifying to use a unified derivation process for deriving one or more collocated motion vectors regardless of the coding tools used to encode the current video block into the bitstream representation. At step 2704, the process performs the conversion using a merge list having one or more collocated motion vectors.
[0288] FIG. 28 is a flowchart of an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 4 of Section 4 of this document. At step 2802, the process identifies one or more conditions related to the dimensions of the current video block of the visual media data, and the intra-block copy (IBC) mode is applied to the current video block. At step 2804, the process determines a motion candidate list construction process for the current video block based on whether one or more conditions related to the dimensions of the current video block are satisfied. At step 2806, the process performs a conversion between the current video block and the bitstream representation of the current video block based on the motion candidate list.
[0289] FIG. 29 is a flowchart of an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 5 of Section 4 of this document. At step 2902, the process determines that an encoding technique is disabled for conversion between the current video block of the visual media data and the bitstream representation of the visual media data, and the bitstream representation is configured to include a field indicating that the maximum number of merge candidates for the encoding technique is zero. At step 2904, the process performs the conversion based on the determination that the encoding technique is disabled.
[0290] Figure 30 is a flowchart for an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 6 of Section 4 of this document. In step 3002, the process makes a determination using a rule for conversion between a current video block and a bitstream representation of the current video block, the rule specifying that a first syntax element in the bitstream representation is conditionally included based on a second syntax element in the bitstream representation that indicates the maximum number of merge candidates related to at least one coding technique applied to the current video block. In step 3004, the process performs the conversion between the current video block and the bitstream representation of the current video block based on the determination.
[0291] Figure 31 is a flowchart for an example of a visual media processing method. The steps of this flowchart are described in relation to Embodiment 4 of Section 4 of this document. In step 3102, the process constructs a motion candidate list using a motion list construction process based on one or more conditions related to the dimensions of a current video block of a video for conversion between the current video block of the video and a bitstream representation of the video. In step 3104, the process performs the conversion using the motion candidate list, the motion candidate list including zero or more intra block copy mode candidates and / or zero or more advanced motion vector predictor candidates.
[0292] Here, some embodiments of this document are presented in a clause-based format.
[0293] A1. A visual media processing method, comprising: determining that a first video block of visual media data uses a geometric partitioning mode (GPM) and that a second video block of visual media data uses a non-GPM mode; Based on a unified pruning process, constructing a first merge list for the first video block and a second merge list for the second video block, wherein the first merge list and the second merge list have merge candidates, and the unified pruning process includes adding a new merge candidate to the merge list based on comparing the motion information of the new merge candidate with the motion information of at least one merge candidate in the merge list, the step of, having, The GPM has dividing the first video block into a plurality of prediction partitions to which motion prediction is separately applied, and at least one partition has a non-rectangular shape. Method.
[0294] Therefore, the pruning process is a unified pruning process that is similarly applied to different video blocks regardless of whether the video block is processed using GPM or other conventional splitting modes.
[0295] For example, item A1 may be implemented as a method having a step of constructing a merge list having motion candidates based on construction rules for conversion between the current video block of visual media data and the bitstream representation of the visual media data. The method may further include the step of performing the conversion using the merge list. The construction rules may use a unified construction procedure to construct a merge candidate list. The unified construction procedure may, for example, apply the pruning process in a unified manner such that the merge list for the current video block divided by GPM can be constructed using the same construction procedure as the second video block not divided by the GPM procedure, for example, including using the same pruning process to generate both the first and second merge lists for the first and second video blocks, respectively.
[0296] The method according to item A1, wherein the merge candidates in the merge list are spatial merge candidates, and the unified pruning process is a partial pruning process.
[0297] The method according to any one of items A1 to A2, wherein the first merge list or the second merge list constructed based on the unified pruning process includes a maximum of four spatial merge candidates from adjacent video blocks of the first video block and / or the second video block.
[0298] The method according to any one of items A1 to A3, wherein the adjacent video blocks of the first video block and / or the second video block include one video block (denoted as B2) at the upper left corner, two video blocks (denoted as B1 and B0) at the upper right corner sharing a common side, and two video blocks (denoted as A1 and A0) at the lower left corner sharing a common side.
[0299] Comparing the motion information of the new merge candidate with the motion information of the at least one merge candidate in the merge list is associated with performing pair - by - pair comparison in an ordered sequence including comparison between B1 and A1, comparison between B1 and B0, comparison between A0 and A1, comparison between B2 and A1, and comparison between B2 and B1. Further, comparing the motion information is the same for the first video block and the second video block. The method according to item A4.
[0300] The method according to item A1, wherein the unified pruning process is a full pruning process selectively applied to the spatial merge candidates based on the first video block or the second video block satisfying one or more threshold conditions.
[0301] The method according to item A6, wherein the one or more threshold conditions are related to the dimension of the first video block, and / or the dimension of the second video block, and / or the number of samples in the first video block, and / or the number of samples in the second video block.
[0302] The comparison involving A8. B2 is selectively executed based on the number of available merge candidates in the first merge list or the second merge list, the method according to any one of items A4 to A7.
[0303] A9. The adjacent video blocks are divided using the non-GPM mode, and the comparison involving B2 is always executed based on the number of available merge candidates in the first merge list or the second merge list, the method according to any one of items A4 to A7.
[0304] B1. A method for visual media processing, For the conversion between the current video block of the visual media data and the bitstream representation of the visual media data, determining an initial value of the motion information applicable to the current video block based on a rule, the rule stipulating to examine whether a sub-block based temporal motion vector predictor coding (SbTMVP) mode is available for the current video block based on the reference list (denoted as list X) of the adjacent blocks of the current video block, provided that X is an integer and the value of X depends at least on the coding conditions of the adjacent blocks, the step and, Executing the conversion based on the determination; A method having.
[0305] B2. The method according to item B1, wherein X indicates the position of the collocated video picture where the temporal motion vector prediction used for the conversion between the current video block and the bitstream representation is performed therefrom.
[0306] B3. The method according to item B1, wherein X is determined by comparing the picture order count (POC) of all reference pictures in all reference lists for the current video block with the POC of the current video picture of the current video block.
[0307] B4. The method according to item B3, wherein when the result of the comparison indicates that the POC of all reference pictures in all reference lists for the current video block is less than or equal to the POC of the current video picture of the current video block, X is set to 1, and in other cases, X is set to 0.
[0308] B5. When the reference picture related to list X of the adjacent blocks is available and the POC of the reference picture is the same as the POC of the collocated video picture, the method further sets the initial value of the motion information in the SbTMVP mode to the motion information related to list X of the adjacent blocks of the current video block. The method according to item B2, comprising
[0309] B6. The method according to item B1, wherein motion information stored in a history-based motion vector predictor table is used to set the initial value of the motion information in the SbTMVP mode.
[0310] B7. The method according to item B6, wherein the motion information stored in the history-based motion vector predictor table is the first available motion information in the history-based motion vector predictor table.
[0311] B8. The method according to item B7, wherein the first available motion information is related to a reference picture.
[0312] B9. The method according to item B8, wherein the reference picture is a collocated picture.
[0313] B10. The method according to item B1, wherein X is a predetermined value.
[0314] B11. The method according to item B10, wherein X is 0 or 1.
[0315] C1. A visual media processing method, comprising For converting between the current video block of the visual media data and the bitstream representation of the visual media data, a step of deriving one or more collocated motion vectors for one or more sub-blocks of the current video block based on a rule, the rule using a unified derivation process for deriving the one or more collocated motion vectors regardless of the encoding tool used to encode the current video block into the bitstream representation, the step; Executing the conversion using a merge list having the one or more collocated motion vectors; A method having.
[0316] C2. If the derivation process utilizes uni-prediction from a reference picture list denoted as list Y, the motion vectors of list Y are scaled to a target reference picture list denoted as list X, provided that X and Y are integers and the value of X depends at least on the encoding tool used for the current video block. The method according to item C1.
[0317] C3. If the derivation process utilizes bi-prediction from a target reference picture list denoted as list Y, the rule further provides for scaling the motion vectors of list Y to a target reference picture list denoted as list X, provided that X and Y are integers and the value of X depends at least on the encoding tool used for the current video block. The method according to item C1.
[0318] C4. The rule further provides for determining X by comparing the picture order count (POC) of all reference pictures in all reference lists for the current video block with the POC of the current video picture of the current video block. The method according to any one of items C2 to C3.
[0319] C5. The method according to item C4, wherein the rule further specifies determining Y by comparing the POCs of all reference pictures in all reference lists for the current video block with the POC of the current video picture of the current video block.
[0320] C6. If the result of the comparison indicates that the POCs of all reference pictures in all reference lists for the current video block are less than or equal to the POC of the current video picture of the current video block, the rule specifies setting Y = X; otherwise, the rule specifies setting X to the position of the collocated video picture from which the temporal motion vector prediction used for the conversion between the current video block and the bitstream representation is performed, according to the method of item C5.
[0321] D1. A method for visual media processing, comprising: identifying one or more conditions related to the dimensions of a current video block of visual media data, wherein an intra-block copy (IBC) mode is applied to the current video block; determining a motion candidate list construction process for a motion candidate list for the current video block based on whether the one or more conditions related to the dimensions of the current video block are satisfied; performing a conversion between the current video block and a bitstream representation of the current video block based on the motion candidate list; and having the method.
[0322] D2. A method for visual media processing, comprising: constructing a motion candidate list using a motion list construction process based on one or more conditions related to the dimensions of a current video block for conversion between the current video block of a video and a bitstream representation of the video; A step of performing the conversion using the motion candidate list, the motion candidate list including zero or more intra-block copy mode candidates and / or zero or more advanced motion vector predictor candidates, and the step; A method having
[0323] D3. The method according to any one of items D1 to D2, wherein the motion candidate list construction process skips derivation of spatial merge candidates when the one or more conditions are satisfied.
[0324] D4. The method according to any one of items D1 to D2, wherein the motion candidate list construction process skips derivation of history-based motion vector candidates when the one or more conditions are satisfied.
[0325] D5. The method according to any one of items D1 to D2, wherein the motion candidate list construction process skips derivation of pairwise merge candidates when the one or more conditions are satisfied.
[0326] D6. The method according to any one of items D1 to D2, wherein the motion candidate list construction process reduces the total number of maximum pruning processes when the one or more conditions are satisfied.
[0327] D7. The method according to item D6, wherein the total number of the maximum pruning processes is reduced to zero.
[0328] D8. The method according to any one of items D1 to D7, wherein the one or more conditions are satisfied when the product of the width of the current video block and the height of the current video block is less than or equal to a threshold value.
[0329] D9. The method according to item D8, wherein the threshold value is 64, 32 or 16.
[0330] D10. The method according to any one of items D1 to D9, wherein the one or more conditions are satisfied when the width and / or the height of the current video block is less than a threshold value.
[0331] D11. The method according to item D10, wherein the threshold value is 4 or 8.
[0332] D12. The method according to any one of items D1 to D2, wherein the motion candidate list includes an IBC merge list or an IBC motion vector prediction list.
[0333] E1. A method for visual media processing, comprising: determining that an encoding technique is disabled for conversion between a current video block of visual media data and a bitstream representation of the visual media data, wherein the bitstream representation is configured to include a field indicating that a maximum number of merge candidates for the encoding technique is zero; executing the conversion based on the determination that the encoding technique is disabled; and a method having the above steps.
[0334] E2. The method according to item E1, wherein the bitstream representation is further configured to skip signaling of one or more syntax elements based on the field indicating that the maximum number of merge candidates for the encoding technique is zero.
[0335] E3. The method according to item E1, wherein the encoding technique corresponds to intra block copy in which samples of the current video block are encoded from other samples within the video region of the current video block.
[0336] E4. The method according to item E1, wherein the encoding technique is a sub-block encoding technique.
[0337] E5. The method according to item E4, wherein the sub-block encoding technique is an affine encoding technique or an alternative motion vector predictor technique.
[0338] E6. The method according to any one of items E1 to E2, wherein the encoding technique applied includes one of combined intra-intra prediction (CIIP), geometric partitioning mode (GPM), decoder motion vector refinement (DMVR), sub-block merge, intra-block copy merge, normal merge, or motion vector difference merge (MMVD).
[0339] F1. A method for visual media processing, performing a determination using a rule for conversion between a current video block and a bitstream representation of the current video block, the rule stipulating that a first syntax element in the bitstream representation is conditionally included based on a second syntax element in the bitstream representation indicating the maximum number of merge candidates related to at least one encoding technique applied to the current video block; executing the conversion between the current video block and the bitstream representation of the current video block based on the determination; and having the method.
[0340] F2. The method according to item F1, wherein the first syntax element corresponds to a merge flag.
[0341] F3. The method according to item F1, wherein the first syntax element corresponds to a skip flag.
[0342] F4. The method according to any one of items F1 to F3, wherein the maximum number of merge candidates is zero.
[0343] F5. The second syntax element is skipped from being included in the bitstream representation, and the method further includes estimating that the second syntax element is zero. The method according to item F4, having
[0344] F6. The method according to any one of items F1 to F5, wherein the at least one encoding technique includes combined intra-intra prediction (CIIP), geometric partitioning mode (GPM), decoder motion vector refinement (DMVR), sub-block merge, intra-block copy merge, normal merge, or motion vector difference merge (MMVD).
[0345] F7. The method according to any one of items F1 to F6, wherein the rule further specifies that the first syntax element corresponds to boolean true and the at least one encoding technique is enabled.
[0346] F8. The method according to item F7, wherein the rule further specifies disabling the intra-block copy (IBC) technique for a slice, tile, block, picture, or current video block.
[0347] F9. The method according to item F7, wherein the rule further specifies disabling the motion vector difference merge (MMVD) technique for a slice, tile, block, picture, or current video block.
[0348] F10. The method according to item F7, wherein the rule further specifies disabling the combined intra-intra prediction (CIIP) technique for a slice, tile, block, picture, or current video block.
[0349] F11. The method according to item F7, wherein the rule further specifies disabling the geometric partitioning mode (GPM) technique for a slice, tile, block, picture, or current video block.
[0350] G1. The method according to any one of items A1 to F11, wherein the transformation includes generating the bitstream representation from the current video block.
[0351] G2. The method according to any one of clauses A1 to F11, wherein the conversion includes generating samples of the current video block from the bitstream representation.
[0352] G3. A video processing apparatus having a processor configured to implement the method according to any one of clauses A1 to F11.
[0353] G4. A video encoding apparatus having a processor configured to implement the method according to any one of clauses A1 to F11.
[0354] G5. A video decoding apparatus having a processor configured to implement the method according to any one of clauses A1 to F11.
[0355] G6. A computer-readable medium storing code, the code causing a processor to execute the method according to any one of clauses A1 to F11 when executed.
[0356] In this document, the terms "video processing" or "visual media processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from the pixel representation of a video to the corresponding bitstream representation or vice versa. The bitstream representation of the current video block may correspond to bits that are collocated within the bitstream or spread out at different locations, as defined by the syntax. For example, a macroblock can be encoded with respect to the transformed and encoded error residual values and can also be encoded using bits in the header and other fields within the bitstream. Further, during the conversion, the decoder can parse the bitstream based on a decision, using the knowledge that some fields may or may not exist, as described in the above solutions. Similarly, the encoder can generate an appropriate encoded representation by determining whether specific syntax fields should be included and including or excluding those syntax fields from the encoded representation.
[0357] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, or in combinations of one or more of these, including the structures disclosed in this document and those structurally equivalent thereto. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., as one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that generates a machine-readable propagated signal, or combinations of one or more of these. The term "data processing apparatus" includes, by way of example, any apparatus, device, and machine for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can include, in addition to hardware, code that creates an execution environment for the computer programs in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or combinations of one or more of these. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver device.
[0358] A computer program (also known as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program need not necessarily correspond to a file in a file system. The program may be stored in a part of a file that holds other programs or data (for example, one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (for example, files that hold one or more modules, subprograms, or portions of code). A computer program may be deployed to be executed on one computer or may be deployed to be executed on multiple computers, which may be located in one place or distributed across multiple places and interconnected by a communication network.
[0359] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data to produce output. These processes and logical flows can also be performed by, for example, a special purpose logic circuit such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as such special purpose logic circuits.
[0360] Processors suitable for the execution of a computer program include, by way of example, any one or more processors of both general and special purpose microprocessors, and any one of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory or both. Essential elements of a computer are a processor for executing instructions and one or more memory devices for storing the instructions and data. In general, a computer also includes one or more mass storage devices for storing data, such as, by way of example, magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from or transfer data to a mass storage device. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include, by way of example, semiconductor memory devices, such as, for example, EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and all forms of non-volatile memory, media, and memory devices, including CD ROM and DVD-ROM disks. The processor and the memory may be supplemented or incorporated by dedicated logic circuitry.
[0361] This patent document contains many details, but they should not be construed as limitations on any subject matter or the scope of what may be claimed. Rather, they should be construed as descriptions of mechanisms that may be specific to particular embodiments of a particular technology. The specific plurality of mechanisms described in this patent document in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various mechanisms described in the context of a single embodiment can also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although a plurality of mechanisms may be described as acting in a particular combination and may even initially be claimed as such, in some cases, one or more mechanisms can be excluded from the claimed combination, and the claimed combination can also be led to a sub-combination or a variation of a sub-combination.
[0362] Similarly, although the processes are shown in a particular order in the drawings, this should not be understood as requiring that those operations be performed in or in the order shown, or that all of the processes shown be performed, in order to achieve the desired result. Also, the separation of the various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.
[0363] Only a few implementations and examples are described, and other implementations, extensions, and variations can be made based on what is described and illustrated in this patent document.
Claims
1. A method for processing video data, comprising: in a first conversion between a first block of video and a bitstream of the video, determining that the first block is coded in a geometric partitioning mode that is a prediction mode; constructing a candidate list for the first block, the construction including examining the availability of a spatial candidate in a specific adjacent block B2 based on the number of available candidates in the candidate list, the specific adjacent block B2 being at the upper left corner of the first block; determining first motion information for a first geometric partition of the first block and second motion information for a second geometric partition of the first block based on the candidate list; applying a weighting process to generate a final prediction for samples of the first block based on a weighted sum of prediction samples derived based on the first motion information and the second motion information; executing the first conversion based on the application; having; examining the availability of the spatial candidate in the specific adjacent block B2 further includes when the candidate list includes spatial candidates in a specific adjacent block A1 adjacent to the lower left corner of the first block and a specific adjacent block B1 adjacent to the upper right corner of the first block, comparing motion information between the spatial candidate in the specific adjacent block B2 and the spatial candidates in the specific adjacent blocks A1 and B1, determining that the spatial candidate in the specific adjacent block B2 is not available based on the result of the comparison indicating that the motion information of the spatial candidate in B2 is the same as that in A1 or B1; including; method.
2. The method according to claim 1, wherein when the number of available candidates in the candidate list is 4, the spatial candidate in the specific adjacent block B2 is not available.
3. The method according to claim 1, wherein the candidate list is constructed based on a pruning process, the pruning process including comparing motion information between at least two spatial candidates to avoid the same spatial candidate.
4. Constructing the candidate list for the first block based on the pruning process includes examining the availability of spatial candidates in the specific adjacent block A1, and the examination includes when the candidate list includes spatial candidates in the specific adjacent block B1, comparing motion information between the spatial candidates in the specific adjacent block A1 and the spatial candidates in the specific adjacent block B1, determining that the spatial candidates in the specific adjacent block A1 are not available based on the result of the comparison indicating that the motion information of the spatial candidates in A1 is the same as that in B1. The method according to claim 3, comprising this.
5. Constructing the candidate list for the first block based on the pruning process includes examining the availability of spatial candidates in a specific adjacent block B0 at the upper right corner of the first block, and the examination includes when the candidate list includes spatial candidates in the specific adjacent block B1, comparing motion information between the spatial candidates in the specific adjacent block B0 and the spatial candidates in the specific adjacent block B1, determining that the spatial candidates in the specific adjacent block B0 are not available based on the result of the comparison indicating that the motion information of the spatial candidates in B0 is the same as that in B1. The method according to claim 3, comprising this.
6. Constructing the candidate list for the first block based on the pruning process includes examining the availability of spatial candidates in a specific adjacent block A0 at the lower left corner of the first block, and the examination includes when the candidate list includes spatial candidates in the specific adjacent block A1, comparing motion information between the spatial candidates in the specific adjacent block A0 and the spatial candidates in the specific adjacent block A1, determining that the spatial candidates in the specific adjacent block A0 are not available based on the result of the comparison indicating that the motion information of the spatial candidates in A0 is the same as that in A1. The method according to claim 3, comprising this.
7. In response to the first motion information being from the first reference picture list LX and the second motion information being from the second reference picture list L(1 - X), a step of storing dual-prediction motion information for 4×4 sub-blocks within the weighted region of the first block, where the dual-prediction motion information is formed by combining the first motion information and the second motion information, and X = 0 or X = 1. The method according to claim 1, further comprising.
8. A step of storing single-prediction motion information for 4×4 sub-blocks within the non-weighted region of the first block, where the single-prediction motion information is based on the first motion information or the second motion information. The method according to claim 1, further comprising.
9. Spatial candidates corresponding to adjacent blocks coded in the intra-block copy mode are excluded from the candidate list for the first block. In the intra-block copy mode, the predicted samples of the adjacent blocks are derived from the block of sample values in the decoded video region of the same picture as the adjacent blocks determined by the block vector. The method according to claim 1.
10. The conversion includes encoding the first block into the bitstream. The method according to any one of claims 1 to 9.
11. The conversion includes decoding the first block from the bitstream. The method according to any one of claims 1 to 9.
12. An apparatus for processing video data having a processor and a non-transitory memory having instructions, where when the instructions are executed by the processor, the processor is caused to, In a first conversion between a first block of a video and the bitstream of the video, determine that the first block is coded in a geometric partitioning mode that is a prediction mode. Construct a candidate list for the first block, where the construction includes examining the availability of spatial candidates in a specific adjacent block B2 based on the number of available candidates in the candidate list, and the specific adjacent block B2 is at the upper left corner of the first block. Based on the candidate list, determine first motion information for a first geometric partition of the first block and second motion information for a second geometric partition of the first block. Apply a weighting process to generate a final prediction for the samples of the first block based on the weighted sum of the prediction samples derived based on the first motion information and the second motion information. Execute the first transformation based on the application. Further, examining the availability of the spatial candidate in the specific adjacent block B2 When the candidate list includes spatial candidates in a specific adjacent block A1 adjacent to the lower left corner of the first block and a specific adjacent block B1 adjacent to the upper right corner of the first block, compare the motion information between the spatial candidate in the specific adjacent block B2 and the spatial candidates in the specific adjacent blocks A1 and B1. Based on the result of the comparison indicating that the motion information of the spatial candidate in B2 is the same as that in A1 or B1, determine that the spatial candidate in the specific adjacent block B2 is not available. including device **Claim 13** A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to In a first transformation between a first block of a video and the bitstream of the video, determine that the first block is coded in a geometric partitioning mode that is a prediction mode. Construct a candidate list for the first block, the construction including examining the availability of a spatial candidate in a specific adjacent block B2 based on the number of available candidates in the candidate list, the specific adjacent block B2 being at the upper left corner of the first block. Based on the candidate list, determine first motion information for a first geometric partition of the first block and second motion information for a second geometric partition of the first block. Apply a weighting process to generate a final prediction for the samples of the first block based on the weighted sum of the prediction samples derived based on the first motion information and the second motion information. Execute the first transformation based on the application. Further, examining the availability of the spatial candidate in the specific adjacent block B2 When the candidate list includes spatial candidates in a specific adjacent block A1 adjacent to the lower left corner of the first block and a specific adjacent block B1 adjacent to the upper right corner of the first block, compare motion information between the spatial candidate in the specific adjacent block B2 and the spatial candidates in the specific adjacent blocks A1 and B1. Based on the result of the comparison indicating that the motion information of the spatial candidate in B2 is the same as that in A1 or B1, determine that the spatial candidate in the specific adjacent block B2 is not available. Including that. A computer-readable storage medium.
14. A method for storing a bitstream of a video, comprising: In a first conversion between a first block of a video and the bitstream of the video, determining that the first block is coded in a geometric partitioning mode that is a prediction mode; Constructing a candidate list for the first block, the construction including examining the availability of a spatial candidate in a specific adjacent block B2 based on the number of available candidates in the candidate list, the specific adjacent block B2 being at the upper left corner of the first block; Determining first motion information for a first geometric partition of the first block and second motion information for a second geometric partition of the first block based on the candidate list; Applying a weighting process to generate a final prediction for samples of the first block based on a weighted sum of prediction samples derived based on the first motion information and the second motion information; Generating the bitstream based on the application; Storing the bitstream in a non-transitory computer-readable recording medium; Having Examining the availability of the spatial candidate in the specific adjacent block B2 further includes: When the candidate list includes spatial candidates in a specific adjacent block A1 adjacent to the lower left corner of the first block and a specific adjacent block B1 adjacent to the upper right corner of the first block, compare motion information between the spatial candidate in the specific adjacent block B2 and the spatial candidates in the specific adjacent blocks A1 and B1. Based on the result of the comparison indicating that the movement information of the spatial candidate in B2 is the same as that in A1 or B1, determine that the spatial candidate in the specific adjacent block B2 is not available. Including Method.
Citation Information
Patent Citations
amvp and merge candidate list derivation for integration of intra-bc and inter-prediction
JP2017535180A