Constraints on Merge Candidates of Blocks in Intra Block Copy Coding and Decoding
By using motion vector difference (MVD)-related syntax elements in the intra-block copy (IBC) advanced motion vector prediction (AMVP) mode in video encoding and decoding, the current sample points are predicted from other samples in the video area, which solves the problem of intra-block encoding and decoding efficiency in the prior art, and improves the video decoding quality and efficiency.
Patent Information
- Application Number
- CN202080038735.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-25
- Filing Date
- 2020-05-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-05-25
AI Technical Summary
The existing video codec standards have problems with inefficiency in intra-block codec, especially when using block vector signaling notification and Merge candidates, it is difficult to effectively improve the codec efficiency.
During the video processing, the motion vector difference (MVD)-related syntax elements of the intra-block copy (IBC) advanced motion vector prediction (AMVP) mode are selectively used, and combined with the motion vector prediction of the intra-block copy (IBC) mode, the current sample point is predicted from other samples in the video area, and the encoding and decoding process of intra-block copy is optimized.
It improves the efficiency of intra-block encoding and decoding, enhances the quality and efficiency of video decoding, and is suitable for existing video encoding and decoding standards such as HEVC and future video encoding and decoding standards VVC.
Smart Images

Figure CN113950840B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application is filed to timely claim the priority and benefits of International Patent Application No. PCT / CN2019 / 088454, filed on May 25, 2019, in accordance with the provisions of applicable patent laws and / or the Paris Convention. For all legal purposes, the entire disclosure of the above - mentioned application is incorporated by reference as part of the disclosure of this application. Technical field
[0003] This document relates to video and image encoding, decoding, and transcoding technologies. Background art
[0004] Digital video still occupies the largest bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video usage is expected to continue to grow. Summary of the invention
[0005] The disclosed technology can be implemented by video or image decoder or encoder embodiments, where decoding or encoding of intra - block - coded video is performed when using block - vector signaling notification and / or Merge candidates.
[0006] In one representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bit - stream representation of the video, where the bit - stream representation selectively includes syntax elements related to the motion - vector difference (MVD) of the intra - block copy (IBC) advanced motion - vector prediction (AMVP) mode based on a maximum number of first - type IBC candidates used during the conversion of the video region, and where, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0007] In another representative aspect, a video processing method is disclosed. The method includes: determining an indication to disable the use of the intra - block copy (IBC) mode for a video region of a video and enable the use of the IBC mode at the sequence level of the video for the conversion between the video region of the video and the bit - stream representation of the video, and performing the conversion based on the determination, and where, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0008] In yet another representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes an indication of the use of an intra block copy (IBC) mode and / or one or more IBC-related syntax elements based on a maximum number of first type of IBC candidates used during the conversion of the video region, and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0009] In yet another representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bitstream representation of the video, wherein an indication of a maximum number of first type of intra block copy (IBC) candidates used during the conversion of the video region is signaled in the bitstream representation independently of a maximum number of Merge candidates of an inter mode used during the conversion, and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0010] In yet another representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates used during the conversion of the video region (denoted as maxIBCCandNum) is a function of a maximum number of IBC Merge candidates (denoted as maxIBCMrgNum) and a maximum number of IBC advanced motion vector prediction (AMVP) candidates (denoted as maxIBCAMVPNum), and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0011] In yet another representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates used during the conversion of the video region (denoted as maxIBCCandNum) is based on mode information of encoding and decoding of the video region.
[0012] In yet another representative aspect, a video processing method is disclosed. The method includes: converting between a video region of a video and a bitstream representation of the video, wherein a decoded intra block copy (IBC) advanced motion vector prediction (AMVP) Merge index or a decoded IBC Merge index is less than a maximum number of intra block copy (IBC) motion candidates (denoted as maxIBCCandNum).
[0013] In yet another representative aspect, a video processing method is disclosed. The method includes: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list, and based on the determination, using a default prediction block during the conversion.
[0014] In yet another representative aspect, a video processing method is disclosed. The method includes: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list, and based on the determination, performing the conversion by treating the video region as having invalid block vectors.
[0015] In yet another representative aspect, a video processing method is disclosed. The method includes: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to meet a condition, generating a supplementary block vector (BV) candidate list based on the determination, and using the supplementary BV candidate list for the conversion.
[0016] In yet another representative aspect, a video processing method is disclosed. The method includes: performing a conversion between the video region of a video and the bitstream representation of the video, wherein the maximum number of Intra Block Copy (IBC) Advanced Motion Vector Predictor (AMVP) candidates (denoted as maxIBCAMVPNum) is not equal to two.
[0017] In another example aspect, the above method can be implemented by a video decoder device including a processor.
[0018] In another example aspect, the above method can be implemented by a video encoder device including a processor.
[0019] In yet another example aspect, these methods can be implemented in the form of processor-executable instructions and stored on a computer-readable program medium.
[0020] These and other aspects will be further described in this document. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 An example of the derivation process for Merge candidate list construction is shown.
[0022] Figure 2 Example locations of spatial merge candidates are shown.
[0023] Figure 3 An example of candidate pairs considered for redundancy checking of spatial merge candidates is shown.
[0024] Figure 4A and 4B Examples of the location of the second prediction unit (PU) for Nx2N and 2NxN partitions are shown.
[0025] Figure 5 FIG2 is a diagram illustrating motion vector scaling for temporal Merge candidates.
[0026] Figure 6 Examples of candidate positions of time-domain merge candidates C0 and C1 are shown.
[0027] Figure 7 An example of a combined bi-predictive Merge candidate is shown.
[0028] Figure 8 The process of deriving motion vector prediction candidates is summarized.
[0029] Figure 9 is a diagram of motion vector scaling for spatial motion vector candidates.
[0030] Figure 10A and 10B A 4-parameter affine motion model and a 6-parameter affine motion model are shown respectively.
[0031] Figure 11 is an example of an affine motion vector field (MVF) for each sub-block.
[0032] Figure 12 Examples showing candidate positions for the affine merge mode
[0033] Figure 13 An example of a modified Merge list building process is shown.
[0034] Figure 14 An example of inter-frame prediction based on triangular partitioning is shown.
[0035] Figure 15 An example of the final motion vector expression (UMVE) search process is shown.
[0036] Figure 16 An example of a UMVE search point is shown.
[0037] Figure 17 An example of MVD(0, 1) mirrored between list 0 and list 1 in DMVR is shown.
[0038] Figure 18 Shows an example of MVs that can be checked in one iteration.
[0039] Figure 19 Shows an example of Intra Block Copy (IBC).
[0040] Figure 20A - 20K Is a flowchart of an example of a method for video processing.
[0041] Figure 21 Is a block diagram of an example of a video processing device.
[0042] Figure 22 Is a block diagram of an example video processing system in which the disclosed technology can be implemented. Detailed Description
[0043] This document provides various techniques that can be used by a decoder of an image or video bitstream to improve the quality of decompressed or decoded digital video or images. For simplicity, the term "video" as used herein includes both a sequence of pictures (traditionally called video) and a single image. Additionally, a video encoder can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0044] The use of section headings in this document is for ease of understanding and does not limit the embodiments and techniques to the corresponding sections. Thus, embodiments of one section can be combined with embodiments of other sections.
[0045] 1. Overview
[0046] The present invention relates to video coding and decoding techniques. Specifically, it relates to motion vector coding and decoding. It can be applied to existing video coding and decoding standards such as HEVC or to standards pending finalization (Universal Video Coding). It can also be applicable to future video coding and decoding standards or video codecs.
[0047] 2. Background
[0048] Video coding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced the H.261 and H.263 standards, ISO / IEC produced the MPEG-1 and MPEG-4 Visual standards, and the two organizations jointly produced the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established, working on the VVC standard with the goal of reducing the bitrate by 50% compared to HEVC.
[0049] The latest version of the VVC draft, i.e., Versatile Video Coding (Draft 5) can be found at:
[0050] phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N1001-v2.zip
[0051] The latest reference software for VVC, called VTM, can be found at:
[0052] vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0
[0053] 2.1 Inter-frame prediction in HEVC / H.265
[0054] For a coding unit (CU) for inter-frame coding, it can be coded with one prediction unit (PU) or two PUs according to the split mode. Each PU for inter-frame prediction has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. inter_pred_idc can also be used to signal the use of one of the two reference picture lists. The motion vector can be explicitly coded as an increment relative to the predictor.
[0055] When encoding / decoding a CU using the skip mode, a PU is associated with the CU, and there are no significant residual coefficients, no encoding / decoding motion vector deltas, or reference picture indices. The Merge mode is specified, whereby motion parameters for the current PU are obtained from neighboring PUs - including spatial and temporal candidates. The Merge mode can be applied to any inter-predicted PU, not just those in the skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) relative to the motion vector predictor), the corresponding reference picture indices for each reference picture list, and the reference picture lists are explicitly signaled for each PU. Such a mode is named Advanced Motion Vector Prediction (AMVP) in this disclosure.
[0056] When the signaling indicates that one of the two reference picture lists is to be used, the PU is generated from the samples of a block. This is referred to as "unidirectional prediction". Unidirectional prediction can be used for P-slices and B-slices.
[0057] When the signaling indicates that two reference picture lists are to be used, the PU is generated from the samples of two blocks. This is referred to as "bidirectional prediction". Bidirectional prediction can only be used for B-slices.
[0058] Details of the inter-prediction modes specified in HEVC are provided below. The description will start with the Merge mode.
[0059] 2.1.1 Reference Picture Lists
[0060] In HEVC, the term inter-prediction is used to denote prediction derived from data elements (e.g., sample values or motion vectors) of reference pictures other than the current decoded picture. As in H.264 / AVC, pictures can be predicted from multiple reference pictures. The reference pictures used for inter-prediction are organized in one or more reference picture lists. The reference index identifies which reference pictures in the list should be used to create the prediction signal.
[0061] A single reference picture list - list 0 (List 0) is used for P-slices, and two reference picture lists - list 0 (List 0) and list 1 (List 1) are used for B-slices. It should be noted that, in terms of the capture / display order, the reference pictures contained in list 0 / 1 can be pictures from the past and the future.
[0062] 2.1.2 Merge Mode
[0063] 2.1.2.1 Derivation of Merge Mode Candidates
[0064] When predicting a PU in Merge mode, an index pointing to an entry in the Merge candidates list is parsed from the bitstream and this index is used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized in the following steps in sequence:
[0065] Step 1: Initial candidate derivation
[0066] Step 1.1: Spatial candidate derivation
[0067] Step 1.2: Redundancy check of spatial candidates
[0068] Step 1.3: Temporal candidate derivation
[0069] Step 2: Additional candidate insertion
[0070] Step 2.1: Creation of bi-predictive candidates
[0071] Step 2.2: Insertion of zero-motion candidates
[0072] In Figure 1 these steps are also schematically depicted. For spatial Merge candidate derivation, up to four Merge candidates are selected among candidates located at five different positions. For temporal Merge candidate derivation, up to one Merge candidate is selected among two candidates. Since the number of candidates for each PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates obtained from Step 1 does not reach the maximum number of Merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is constant, truncated unary binary (TU) is used to encode the index of the best Merge candidate. If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of a 2N×2N prediction unit.
[0073] In the following, the operations associated with the foregoing steps are described in detail.
[0074] Figure 1 An example of the derivation process for constructing the Merge candidate list is shown.
[0075] 2.1.2.2 Spatial candidate derivation
[0076] In the derivation of spatial Merge candidates, among those located at Figure 2Select up to four Merge candidates from the candidates at the positions depicted. The order of derivation is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, A0 are not available (e.g., because the PU belongs to another slice or tile) or is intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrows in Figure 3 are considered, and the candidate is added to the list only if the corresponding candidate used for the redundancy check has different motion information. Another source of duplicate motion information is the "second PU" associated with a partition different from 2Nx2N. As an example, Fig. 4 depicts the second PU for the cases of N×2N and 2N×N, respectively. When the current PU is partitioned into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units with the same motion information, which is redundant for a coding unit having only one PU. Similarly, when the current PU is partitioned into 2N×N, position B1 is not considered.
[0077] Figure 2 Fig. 6 shows an example position of a spatial-domain Merge candidate.
[0078] Figure 3 Fig. 10 shows an example of candidate pairs considered for the redundancy check of spatial-domain Merge candidates.
[0079] Fig. 4 shows an example of the positions of the second PU for Nx2N and 2NxN partitions.
[0080] 2.1.2.3 Temporal Candidate Derivation
[0081] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal Merge candidate, a scaled motion vector is derived based on a co-located PU that belongs to the picture with the minimum POC difference relative to the current picture within a given reference picture list. The reference picture list used for the derivation of the co-located PU is signaled explicitly in the slice header. As Figure 5As shown by the dashed line, a scaled motion vector for a temporal Merge candidate is obtained. The scaled motion vector for the temporal Merge candidate is scaled from the motion vector of a collocated PU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. The actual implementation of the scaling process is described in the HEVC specification. For B slices, two motion vectors are obtained and combined to generate a bi-predictive Merge candidate. One of the two motion vectors is for reference picture list 0 (list 0) and the other is for reference picture list 1 (list 1).
[0082] Figure 5 Fig. shows an illustration of the motion vector scaling for a temporal Merge candidate.
[0083] As Figure 6 shown, in a collocated PU (Y) belonging to a reference frame, a position for a temporal candidate is selected between candidates C0 and C1. If the PU at position C0 is unavailable, intra-coded or outside the current coding tree unit (CTU, i.e., LCU, largest coding unit), then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate.
[0084] Figure 6 Fig. shows an example of the candidate positions of temporal Merge candidates C0 and C1.
[0085] 2.1.2.4 Additional candidate insertion
[0086] In addition to the spatial and temporal Merge candidates, there are two additional types of Merge candidates: combined bi-predictive Merge candidates and zero Merge candidates. The combined bi-predictive Merge candidates are generated by leveraging the spatial and temporal Merge candidates. The combined bi-predictive Merge candidates are only used for B slices. The combined bi-predictive candidates are generated by combining the first reference picture list motion parameters of an initial candidate with the second reference picture list motion parameters of another candidate. If the two tuples provide different motion hypotheses, they will form a new bi-predictive candidate. As an example, Figure 7 depicts the following situation where two candidates with mvL0 and refIdxL0 or mvL1 and refIdxL1 in the original list (on the left) are used to create a combined bi-predictive Merge candidate that is added to the final list (on the right). There are numerous rules regarding the combinations that are considered to generate these additional Merge candidates.
[0087] Figure 7 Shows an example of a combined bidirectional prediction Merge candidate.
[0088] Zero motion candidates are inserted to fill the remaining entries in the Merge candidate list and thus reach the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference picture index that starts from zero and increases whenever a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0089] 2.1.3 AMVP
[0090] AMVP utilizes the spatio-temporal correlation of motion vectors with adjacent PUs, and this spatio-temporal correlation is used for the explicit transmission of motion parameters. For each reference picture list, a motion vector candidate list is constructed by the following operations: First, check the availability of the left and upper temporally adjacent PU positions, remove redundant candidates, and add zero vectors to make the candidate list a constant length. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to the Merge index signaling, a truncated unary is used to code the index of the best motion vector candidate. The maximum value to be coded in this case is 2 (see Figure 8 ). In the following sections, details of the derivation process of motion vector prediction candidates are provided.
[0091] 2.1.3.1 Derivation of AMVP Candidates
[0092] Figure 8 Summarizes the derivation process for motion vector prediction candidates.
[0093] Figure 8 Shows an example of the derivation process of motion vector prediction candidates.
[0094] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. As Figure 2 shown, for the derivation of spatial motion vector candidates, two motion vector candidates are finally derived based on the motion vectors of each PU located at five different positions.
[0095] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived from two different co-located positions. After creating the first list of spatial-temporal candidates, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than 2, motion vector candidates whose reference picture index in the associated reference picture list is greater than 1 are removed from the list. If the number of spatial-temporal motion vector candidates is less than 2, additional zero motion vector candidates are added to the list.
[0096] 2.1.3.2 Spatial motion vector candidates
[0097] In the derivation of spatial motion vector candidates, at most two candidates are considered among five potential candidates, which are from PUs located at the positions as Figure 2 shown, and these positions are the same as those for motion Merge. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Thus, for each side, there are four cases available as motion vector candidates, where two cases do not require the use of spatial scaling, and two cases use spatial scaling. The four different cases are summarized as follows.
[0098] No spatial scaling
[0099] -(1) The same reference picture list and the same reference picture index (the same POC)
[0100] -(2) Different reference picture lists, but the same reference picture (the same POC)
[0101] Spatial scaling
[0102] -(3) The same reference picture list, but different reference pictures (different POCs)
[0103] -(4) Different reference picture lists and different reference pictures (different POCs)
[0104] First, the no-spatial-scaling cases are checked, and then the spatial-scaling cases are checked. Regardless of the reference picture list, when the POC is different between the reference pictures of adjacent PUs and the reference picture of the current PU, spatial scaling is considered. If all PUs of the left candidate are unavailable or are intra-coded, spatial scaling of the upper motion vector is allowed to assist in the parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling of the upper motion vector is not allowed.
[0105] Figure 9 Figure is a diagram of the motion vector scaling for spatial motion vector candidates.
[0106] During the spatial domain scaling process, the motion vectors of adjacent PUs are scaled in a similar manner to the temporal domain scaling, as shown. The main difference is that the reference picture list and index of the current PU are given as input; the actual scaling process is the same as the temporal domain scaling process.
[0107] 2.1.3.3 Temporal Motion Vector Candidates
[0108] All the processes for the derivation of the temporal Merge candidates are the same as those for the derivation of the spatial motion vector candidates (see ) except for the reference picture index derivation. The reference picture index is signaled to the decoder.
[0109] 2.2 Inter - frame Prediction Methods in VVC
[0110] There are several new coding - decoding tools for inter - frame prediction improvement, such as Adaptive Motion Vector Difference Resolution (AMVR) for signaling the MVD, Merge with Motion Vector Difference (MMVD), Triangle Prediction Mode (TPM), Combined Intra - Inter Prediction (CIIP), Advanced TMVP (ATMVP, also known as SbTMVP), Affine Prediction Mode, Generalized Bi - directional Prediction (GBI), Decoder - side Motion Vector Refinement (DMVR), and Bi - directional Optical Flow (BIO, also known as BDOF).
[0111] Three different Merge list construction processes are supported in VVC:
[0112] 1) Sub - block Merge candidate list: It includes ATMVP and Affine Merge candidates. Both the Affine mode and the ATMVP mode share a Merge list construction process. Here, ATMVP and Affine Merge candidates can be added in sequence. The size of the sub - block Merge list is signaled in the slice header, with a maximum value of 5.
[0113] 2) Regular Merge list: For the remaining coding - decoding blocks, a Merge list construction process is shared. Here, spatial / temporal / HMVP, pairwise - combined bi - directional prediction Merge candidates, and zero - motion candidates can be inserted in sequence. The size of the regular Merge list is signaled in the slice header, and the maximum value is 6. MMVD, TPM, CIIP rely on the regular Merge list.
[0114] 3) IBC Merge list: It is completed in a similar way to the regular Merge list.
[0115] Similarly, VVC supports three AMVP lists:
[0116] 1) Affine AMVP candidate list
[0117] 2) Conventional AMVP candidate list
[0118] 3) IBC AMVP candidate list: Since JVET-N0843 is adopted, the construction process is the same as that of the IBC Merge list
[0119] 2.2.1 Coding and decoding block structure in VVC
[0120] In VVC, the picture is divided into square or rectangular blocks using a quadtree / binary tree / trinary tree (QT / BT / TT) structure
[0121] In addition to QT / BT / TT, a separate tree (also called a dual coding tree) is also adopted in VVC for I-frames. With the separate tree, the coding and decoding block structures are signaled for the luminance and chrominance components respectively
[0122] In addition, the CU is set to be equal to the PU and TU, except for blocks coded with several specific coding and decoding methods (such as intra sub-division prediction, where the PU is equal to the TU but smaller than the CU, and sub-block transform for inter-coded blocks, where the PU is equal to the CU but the TU is smaller than the PU)
[0123] 2.2.2 Affine prediction mode
[0124] In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). However, in the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, a simplified affine transform motion compensation prediction is applied to 4-parameter and 6-parameter affine models. As shown, the affine motion field of the block is described by two control point motion vectors (CPMVs) of the 4-parameter affine model and 3 CPMVs of the 6-parameter affine model
[0125] and 10B show: 10A: Simplified affine motion model parameter affine, 10B: 6-parameter affine mode
[0126] The motion vector field (MVF) of the block is described by the following equations for the 4-parameter affine model in Equation (1) (where the 4 parameters are defined as variables a, b, e, and f) and the 6-parameter affine model in Equation (2) (where the 4 parameters are defined as variables a, b, c, d, e, and f) respectively
[0127]
[0128]
[0129] where (mv h 0, mv h 0) is the motion vector of the top - left control point, and (mv h 1, mv h 1) is the motion vector of the top - right control point, and (mv h 2, mv h 2) is the motion vector of the bottom - left control point. All three motion vectors are referred to as control - point motion vectors (CPMVs). (x, y) represents the coordinates of the representative point relative to the top - left sample point within the current block, and (mv h (x, y), mv v (x, y)) is the motion vector derived for the sample located at (x, y). The CP motion vectors can be signaled (e.g., in the affine AMVP mode) or derived on - the - fly (e.g., in the affine Merge mode). w and h are the width and height of the current block. In practice, this division is implemented by a right - shift with rounding. In VTM, the representative point is defined as the center position of the sub - block. For example, when the top - left corner of the sub - block has coordinates (xs, ys) relative to the top - left sample point within the current block, the coordinates of the representative point are defined as (xs + 2, ys + 2). For each sub - block (i.e., 4X4 in VTM), the representative point is used to derive the motion vector of the entire sub - block.
[0130] To further simplify motion - compensation prediction, sub - block - based affine - transform prediction is applied. To derive the motion vector for each M×N (in the current VVC, both M and N are set to 4) sub - block, as shown, the motion vector of the center sample of each sub - block is calculated according to equations (1) and (2) and rounded to 1 / 16 - fraction accuracy. Then, a 1 / 16 - pixel motion - compensation interpolation filter can be applied to generate the prediction for each sub - block using the derived motion vector. The 1 / 16 - pixel interpolation filter is introduced through the affine mode.
[0131] is an example of the affine MVF for each sub - block.
[0132] After MCP, the high - accuracy motion vectors of each sub - block are rounded and saved with the same accuracy as the normal motion vectors.
[0133] 2.2.3 MERGE for the whole block
[0134] 2.2.3.1 Construction of the Merge list for the regular Merge mode of translation
[0135] 2.2.3.1.1 History - based motion - vector prediction (HMVP)
[0136] Different from the Merge list design, in VVC, a history-based motion vector prediction (HMVP) method is adopted.
[0137] In HMVP, the previously encoded / decoded motion information is stored. The motion information of the previously encoded / decoded block is defined as an HMVP candidate. Multiple HMVP candidates are stored in a table called the HMVP table, and this table is maintained during the real-time encoding / decoding process. When starting to encode / decoded a new slice / LCU row / strip, the HMVP table is cleared. Whenever there is an inter-coded block and it is not a sub-block and not in the TPM mode, the associated motion information is added as a new HMVP candidate to the last entry of the table. The entire encoding / decoding process is as shown.
[0138] 2.2.3.1.2 Conventional Merge List Construction Process
[0139] The construction of the conventional Merge list (for translational motion) can be summarized according to the following step sequence:
[0140] · Step 1: Derive spatial candidates
[0141] · Step 2: Insert HMVP candidates
[0142] · Step 3: Insert paired average candidates
[0143] · Step 4: Default motion candidates
[0144] HMVP candidates can be used in the construction processes of both the AMVP and the Merge candidate list. Depicts the modified Merge candidate list construction process (highlighted in blue). When, after inserting the TMVP candidate, the Merge candidate list is not full, the HMVP candidates stored in the HMVP table can be used to fill the Merge candidate list. Considering that a block usually has a higher correlation with the nearest neighboring block in terms of motion information, the HMVP candidates in the table are inserted in descending order of the index. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy removal is applied to the HMVP candidates. Once the total number of available Merge candidates reaches the maximum number of Merge candidates allowed to be signaled, the Merge candidate list construction process terminates.
[0145] Note that all spatial / temporal / HMVP candidates should be encoded / decoded in non-IBC mode. Otherwise, it is not allowed to add them to the conventional Merge candidate list.
[0146] The HMVP table contains up to 5 conventional motion candidates, and each candidate is unique.
[0147] 2.2.3.2 Triangular Prediction Mode (TPM)
[0148] In VTM4, triangular partitioning mode is supported for inter prediction. The triangular partitioning mode is only applicable to CUs that are 8x8 or larger and are encoded and decoded in Merge mode rather than in MMVD or CIIP mode. For CUs that meet these conditions, a CU-level flag is signaled to indicate whether the triangular partitioning mode is applied.
[0149] When this mode is used, as shown, the CU is evenly divided into two triangular partitions using a diagonal partition or an anti-diagonal partition. Each triangular partition in the CU is inter predicted using its own motion; each partition only allows unidirectional prediction, that is, each partition has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure that, like conventional bi-directional prediction, each CU only requires two motion-compensated predictions.
[0150] An example of inter prediction based on triangular partitioning is shown.
[0151] If the CU-level flag indicates that the current CU is encoded and decoded using the triangular partitioning mode, then a flag indicating the direction of the triangular partition (diagonal or anti-diagonal) and two Merge indices (one for each partition) are further signaled. After predicting each triangular partition, a hybrid process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edge. This is the prediction signal for the entire CU, and like in other prediction modes, the transform and quantization processes are applied to the entire CU. Finally, the motion field of the CU predicted using the triangular partitioning mode is stored in 4x4 units.
[0152] The regular Merge candidate list is reused for triangular partition Merge prediction without additional motion vector pruning. For each Merge candidate in the regular Merge candidate list, only one of its L0 or L1 motion vectors is used for triangular prediction. Additionally, the order of selecting the L0 and L1 motion vectors is based on the parity of their Merge indices. With this scheme, the regular Merge list can be directly used.
[0153] 2.2.3.3 MMVD
[0154] In JVET-L0054, the Ultimate Motion Vector Representation (UMVE, also known as MMVD) was proposed. UMVE can be used for skip or Merge mode through the proposed motion vector representation method.
[0155] UMVE reuses the same merge candidates as those included in the regular merge candidate list in VVC. Among the merge candidates, a basic candidate can be selected and further extended by the proposed motion vector representation method.
[0156] UMVE provides a new method for representing the motion vector difference (MVD), where the starting point, motion amplitude, and motion direction are used to represent the MVD.
[0157] An example of the UMVE search process is shown.
[0158] An example of the UMVE search point is shown.
[0159] The proposed technique uses the merge candidate list as it is. However, only the candidates of the default merge type (MRG_TYPE_DEFAULT_N) will consider the extension of UMVE.
[0160] The basic candidate index defines the starting point. The basic candidate index indicates the best candidate among the candidates in the list as follows.
[0161] Table 1: Basic candidate IDX
[0162] 0 1 2 3
[0163] If the number of basic candidates is equal to 1, the basic candidate IDX is not signaled.
[0164] The distance index is the motion amplitude information. The distance index indicates the predefined distance from the starting point information. The predefined distances are as follows:
[0165] Table 2. Distance IDX
[0166]
[0167] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions as follows.
[0168] Table 3. Direction IDX
[0169] 00 01 10 11 + – + –
[0170] Signal the UMVE flag immediately after sending the skip flag or the Merge flag. If the skip or Merge flag is true, parse the UMVE flag. If the UMVE flag equals 1, parse the UMVE syntax. However, if it is not 1, parse the affine flag. If the affine flag equals 1, it is the affine mode; however, if it is not equal to 1, parse the skip / Merge index for the skip / Merge mode of VTM.
[0171] Due to UMVE candidates, no additional line buffer is required. Because the software skip / Merge candidates are directly used as the base candidates. Using the input UMVE index, the supplement of the MV can be determined before motion compensation. There is no need to reserve a long line buffer for this.
[0172] Under the current general test conditions, the first or second Merge candidate in the Merge candidate list can be selected as the base candidate.
[0173] UMVE is also known as Merge with MV Difference (MMVD).
[0174] 2.2.3.4 Combined Intra-Inter Prediction (CIIP)
[0175] In JVET-L0100, multi-hypothesis prediction was proposed, where combined intra prediction and inter prediction is a method to generate multiple hypotheses.
[0176] When applying multi-hypothesis prediction to improve the intra mode, multi-hypothesis prediction combines an intra prediction and a Merge index prediction. In a Merge CU, when the flag is true, signal a flag for the Merge mode to select the intra mode from the intra candidate list. For the luminance component, the intra candidate list is derived only from one intra prediction mode, i.e., the planar mode. The weights applied to the prediction block from intra and inter predictions are determined by the coding / decoding modes (intra or non-intra) of two adjacent blocks (A1 and B1).
[0177] 2.2.4 Sub-Block Based Merge Techniques
[0178] It is recommended to put all motion candidates related to sub-blocks not only into the regular Merge list of non-sub-block Merge candidates but also into a separate Merge list.
[0179] Put the motion candidates related to sub-blocks in a separate Merge list, which is named "sub-block Merge candidate list".
[0180] In one example, the sub-block Merge candidate list includes ATMVP candidates and affine Merge candidates.
[0181] The sub-block Merge candidate list is populated with candidates in the following order:
[0182] a. ATMVP candidates (possibly available or not);
[0183] b. Affine Merge list (including inherited affine candidates; and constructed affine candidates)
[0184] c. Zero-filled MV 4-parameter affine model
[0185] 2.2.4.1.1 ATMVP (also known as Sub-block Temporal Motion Vector Predictor, SbTMVP)
[0186] The basic idea of ATMVP is to derive multiple sets of temporal motion vector predictors for a block. Each sub-block is assigned a set of motion information. When generating ATMVP Merge candidates, motion compensation is performed at the 8x8 level instead of the entire block level.
[0187] 2.2.5 Conventional Inter Prediction Mode (AMVP)
[0188] 2.2.5.1 AMVP Motion Candidate List
[0189] Similar to the AMVP design in HEVC, up to 2 AMVP candidates can be derived. However, HMVP candidates can also be added after TMVP candidates. The HMVP candidates in the HMVP table are traversed in ascending order of index (i.e., starting from the index equal to 0, the oldest index). Up to 4 HMVP candidates can be checked to find if their reference pictures are the same as the target reference picture (i.e., the same POC value).
[0190] 2.2.5.2 AMVR
[0191] In HEVC, when the use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector and the predicted motion vector of the PU) is signaled in units of quarter luminance samples. In VVC, Local Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, the MVD can be coded and decoded in units of quarter luminance samples, integer luminance samples, or four luminance samples (i.e., 1 / 4 pixel, 1 pixel, 4 pixels). The MVD resolution is controlled at the Coding Unit (CU) level, and the MVD resolution flag is conditionally signaled for each CU with at least one non-zero MVD component.
[0192] For a CU with at least one non-zero MVD component, signal a first flag to indicate whether quarter-luma sample MV precision is used in the CU. When the first flag (equal to 1) indicates that quarter-luma sample MV precision is not used, signal another flag to indicate whether integer-luma sample MV precision or four-luma sample MV precision is used.
[0193] When the first MVD resolution flag of a CU is zero or the CU is not decoded (meaning all MVDs in the CU are zero), quarter-luma sample MV resolution is used for the CU. When the CU uses integer-luma sample MV precision or four-luma sample MV precision, the MVP in the CU's AMVP candidate list is rounded to the corresponding precision.
[0194] 2.2.5.3 Symmetric Motion Vector Difference in JVET-N1001-v2
[0195] In JVET-N1001-v2, symmetric motion vector difference (SMVD) is applied to the coding of motion information in bi-prediction.
[0196] First, at the slice level, the variables RefIdxSymL0 and RefIdxSymL1 are derived separately in the following steps specified in N1001-v2 to indicate the reference picture indices of list 0 / 1 used in the SMVD mode. When at least one of the two variables is equal to -1, the SMVD mode shall be disabled.
[0197] 2.2.6 Refinement of Motion Information
[0198] 2.2.6.1 Decoder-Side Motion Vector Refinement (DMVR)
[0199] In bi-prediction operation, for the prediction of a block region, two prediction blocks formed by the motion vectors (MVs) of list 0 and list 1 are respectively combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of bi-prediction are further refined.
[0200] For DMVR in VVC, assume the MVD mirroring between list 0 and list 1 as As shown, bilateral matching is performed to refine the MV, that is, to find the best MVD among several MVD candidates. The MVs of two reference picture lists are represented by MVL0 (L0X, L0Y) and MVL1 (L1X, L1Y). The MVD that minimizes the cost function (e.g., SAD) represented by (MvdX, MvdY) of list 0 is defined as the best MVD. For the SAD function, it is defined as the SAD between the reference block of list 0 derived from the motion vector (L0X + MvdX, L0Y + MvdY) in the reference picture of list 0 and the reference block of list 1 derived from the motion vector (L1X - MvdX, L1Y - MvdY) in the reference picture of list 1.
[0201] The motion vector refinement process can be iterated twice. In each iteration, up to 6 MVDs (with integer pixel precision) can be checked in two steps, as shown. In the first step, MVDs (0, 0), (-1, 0), (1, 0), (0, -1), (0, 1) are checked. In the second step, one of MVDs (-1, -1), (-1, 1), (1, -1) or (1, 1) can be selected and further checked. Assume the function Sad(x, y) returns the SAD value of MVD (x, y). The MVD represented by (MvdX, MvdY) checked in the second step is determined as follows:
[0202] MvdX = -1;
[0203] MvdY = -1;
[0204] If (Sad(1, 0) < Sad(-1, 0))
[0205] MvdX = 1;
[0206] If (Sad(0, 1) < Sad(0, -1))
[0207] MvdY = 1;
[0208] In the first iteration, the starting point is the signaled MV, and in the second iteration, the starting point is the signaled MV plus the best MVD selected in the first iteration. DMVR is applied only when one reference picture is the previous picture and the other reference picture is the next picture and the picture order count distances of the two reference pictures from the current picture are the same.
[0209] An example of the MVD (0, 1) mirrored between list 0 and list 1 in DMVR is shown.
[0210] An example of the MVs that can be checked in one iteration is shown.
[0211] To further simplify the DMVR process, JVET-M0147 proposed some improvements to the design in JEM. More specifically, the DMVR design adopted by VTM-4.0 (to be released soon) has the following main functions:
[0212] · Terminate early when the SAD at the (0, 0) position between list 0 and list 1 is less than the threshold.
[0213] · Terminate early when the SAD between list 0 and list 1 is zero for some positions.
[0214] · Block size of DMVR: W*H >= 64 && H >= 8, where W and H are the width and height of the block.
[0215] · For DMVR with CU size > 16*16, the CU is divided into multiple 16x16 sub-blocks. If only the width or height of the CU is greater than 16, it is divided only in the vertical or horizontal direction.
[0216] · Reference block size (W+7)*(H+7) (for luminance).
[0217] · 25-point SAD-based integer pixel search (i.e., (+-)2 refined search range, single stage)
[0218] · DMVR based on bilinear interpolation.
[0219] · Sub-pixel refinement based on the "parameter error surface equation". This process is performed only when the minimum SAD cost is not equal to zero and the best MVD was (0, 0) in the previous MV refinement iteration.
[0220] · Luminance / chroma MC w / reference block filling (if required).
[0221] · Refined MVs only for MC and TMVP.
[0222] 2.2.6.1.1 Usage of DMVR
[0223] DMVR can be enabled when all of the following conditions are true:
[0224] – The DMVR enable flag in the SPS (i.e., sps_dmvr_enabled_flag) is equal to 1
[0225] – The TPM flag, inter-frame affine flag, and sub-block Merge flag (ATMVP or affine Merge), and the MMVD flag are all equal to 0
[0226] – The Merge flag is equal to 1
[0227] – The current block is bi - directionally predicted and the POC distance between the current picture in list 1 and the reference picture is equal to the POC distance between the reference picture in list 0 and the current picture
[0228] – The height of the current CU is greater than or equal to 8
[0229] – The number of luma samples (CU width * height) is greater than or equal to 64
[0230] 2.2.6.1.2 Sub - pixel refinement based on the "parameter error surface equation"
[0231] The method is summarized as follows:
[0232] 1. Calculate the parameter error surface fitting only when the center position is the best cost position in a given iteration.
[0233] 2. The center position cost and the costs at positions (-1, 0), (0, -1), (1, 0), and (0, 1) from the center are used to fit a two - dimensional parabolic error surface equation of the following form
[0234] E(x,y) = A(x - x0) 2 + B(y - y0) 2 + C
[0235] where (x0, y0) corresponds to the position with the minimum cost and C corresponds to the minimum cost value. By solving 5 equations for 5 unknowns, (x0, y0) can be calculated as:
[0236] x0 = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0)))
[0237] y0 = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0)))
[0238] The sub - pixel accuracy of (x0, y0) can be calculated to any desired level by adjusting the precision of the division (i.e., how many bits of the quotient are calculated). For 1 / 16 pixel accuracy, only 4 bits of the absolute value of the quotient need to be calculated, which helps in the implementation of fast - shift - based subtraction for the 2 divisions required per CU.
[0239] 3. Add the calculated (x0, y0) to the integer - distance refined MV to obtain a sub - pixel accurate refined increment MV.
[0240] 2.2.6.2 Bidirectional optical flow (BDOF)
[0241] 2.3 Intra - block copy
[0242] The High Efficiency Video Coding (HEVC) Screen Content Coding extension (HEVC-SCC) and the current Versatile Video Coding (VVC) Test Model (VTM-4.0) have adopted Intra Block Copy (IBC), also known as current picture reference. IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. As shown, when IBC is applied, the current block is predicted from a reference block in the same picture. The samples in the reference block must have been reconstructed before the current block is encoded or decoded. Although IBC is not efficient for most sequences captured by cameras, it shows significant coding gain for screen content. The reason is that there are many repetitive patterns in screen content pictures, such as icons and text characters. IBC can effectively remove the redundancy between these repetitive patterns. In HEVC-SCC, if the current picture is selected as its reference picture, the coding unit (CU) for inter-frame coding can apply IBC. In this case, the motion vector (MV) is renamed as the block vector (BV), and the BV always has integer pixel precision. To be compatible with the main profile of HEVC, the current picture is marked as a "long-term" reference picture in the Decoded Picture Buffer (DPB). It should be noted that, similarly, in the multi-view / 3D video coding standard, the inter-view reference pictures are also marked as "long-term" reference pictures.
[0243] After the BV finds its reference block, the prediction can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixels from the original signal. Then, transformation and quantization can be applied as in other coding modes.
[0244] is an example of intra block copy.
[0245] However, when the reference block is outside the picture, or overlaps with the current block, or is outside the reconstructed area, or outside the valid area restricted by certain constraints, some or all of the pixel values are undefined. Basically, there are two solutions to this problem. One is to prohibit such cases, for example, in bitstream conformance. The other is to apply padding to those undefined pixel values. The following subsections describe the solutions in detail.
[0246] 2.3.1 IBC in the VVC Test Model (VTM4.0)
[0247] In the current VVC Test Model (e.g., the VTM-4.0 design), the entire reference block should be consistent with the current coding tree unit (CTU) and not overlap with the current block. Therefore, no padding is required for the reference or prediction blocks. The IBC flag is encoded as the prediction mode of the current CU. Therefore, there are a total of three prediction modes for each CU: MODE_INTRA, MODE_INTER, and MODE_IBC.
[0248] 2.3.1.1 IBC Merge Mode
[0249] In the IBC Merge mode, an index pointing to an entry in the IBC Merge candidate list is parsed from the bitstream. The construction of the IBC Merge list can be summarized according to the following sequence of steps:
[0250] · Step 1: Derive spatial candidates
[0251] · Step 2: Insert HMVP candidates
[0252] · Step 3: Insert paired average candidates
[0253] In the derivation of spatial Merge candidates, up to four Merge candidates are selected among the candidates located at the positions shown as A1, B1, B0, A0, and B2 as depicted in the appendix. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, A0 are unavailable (e.g., because it belongs to another strip or slice) or not coded / decoded in the IBC mode. After adding the candidate at position A1, a redundancy check is performed on the insertion of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving the coding / decoding efficiency.
[0254] After inserting the spatial candidates, if the IBC Merge list size is still smaller than the maximum IBC Merge list size, IBC candidates from the HMVP table can be inserted. A redundancy check is performed when inserting HMVP candidates.
[0255] Finally, the paired average candidates are inserted into the IBC Merge list.
[0256] A Merge candidate is called an invalid Merge candidate when the reference block identified by the Merge candidate is outside the picture, or overlaps with the current block, or is outside the reconstruction area, or is outside the valid area subject to certain constraints.
[0257] Note that invalid Merge candidates can be inserted into the IBC Merge list.
[0258] 2.3.1.2 IBC AMVP Mode
[0259] In the IBC AMVP mode, the AMVP index pointing to an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized according to the following sequence of steps:
[0260] Step 1: Derive airspace candidates
[0261] o Check A0, A1 until a usable candidate is found.
[0262] o Check B0, B1, B2 until a usable candidate is found.
[0263] Step 2: Insert HMVP candidates
[0264] Step 3: Insert zero candidates
[0265] After inserting the spatial candidates, if the IBC AMVP list size is still smaller than the maximum IBC AMVP list size, the IBC candidates from the HMVP table may be inserted.
[0266] Finally, the zero candidate is inserted into the IBC AMVP list.
[0267] 2.3.1.3 Chroma IBC Mode
[0268] In the current VVC, motion compensation in chroma IBC mode is performed at the sub-block level. The chroma block will be split into several sub-blocks. Each sub-block determines whether the corresponding luminance block has a block vector and the validity (if any). There is an encoder constraint in the current VTM, where the chroma IBC mode will be tested if all sub-blocks in the current chroma CU have valid luminance block vectors. For example, on YUV 420 video, the chroma block is NxM, and then the co-located luminance area is 2Nx2M. The sub-block size of the chroma block is 2x2. There are several steps for chroma mv derivation, followed by a block copy process.
[0269] 1) The chroma block will first be split into (N>>1)*(M>>1) sub-blocks.
[0270] 2) Each sub-block with an upper left sample point with coordinates (x, y) retrieves the corresponding luma block that covers the same upper left sample point with coordinates (2x, 2y).
[0271] 3) The encoder checks the block vector (bv) of the retrieved luminance block. If one of the following conditions is met, bv is considered invalid.
[0272] a. The bv corresponding to the luminance block does not exist.
[0273] b. The prediction block identified by bv has not been reconstructed.
[0274] c. The prediction block identified by bv partially or completely overlaps with the current block.
[0275] 4) Set the chroma motion vector of the sub-block to the motion vector of the corresponding luminance sub-block.
[0276] When all sub - blocks find valid bv, the IBC mode is allowed at the encoder.
[0277] The decoding process of the IBC block is listed below. The parts related to the chrominance motion vector derivation in the IBC mode are enclosed in double thick brackets, i.e., {{a}} indicates that "a" is related to the chrominance motion vector derivation in the IBC mode.
[0278] 8.6.1 General decoding process of the coded - decoded unit coded - decoded in IBC prediction
[0279] The inputs to this process are:
[0280] – The luma position (xCb, yCb), which specifies the top - left sample of the current coded - decoded block relative to the top - left luma sample of the current picture,
[0281] – The variable cbWidth, which specifies the width of the current coded - decoded block in luma samples,
[0282] – The variable cbHeight, which specifies the height of the current coded - decoded block in luma samples,
[0283] – The variable treeType, which specifies whether to use a single tree or a dual tree, and if a dual tree is used, it specifies whether the current tree corresponds to the luma component or the chrominance component.
[0284] The output of this process is the modified reconstructed picture before loop filtering.
[0285] The process for deriving the quantization parameters specified in Section 8.7.1 is called with the luma position (xCb, yCb), the width cbWidth of the current coded - decoded block in luma samples, the height cbHeight of the current coded - decoded block in luma samples, and the variable treeType as inputs.
[0286] The decoding process of the coded - decoded unit coded - decoded in the ibc prediction mode includes the following ordered steps:
[0287] 1. The motion vector components of the current coded - decoded unit are derived as follows:
[0288] 1. If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the following applies:
[0289] – The process for deriving the motion vector components specified in Section 8.6.2.1 is called with the luma coded - decoded block position (xCb, yCb), the luma coded - decoded block width cbWidth, and the luma coded - decoded block height cbHeight as inputs and the luma motion vector mvL[0][0] as the output.
[0290] – When treeType is equal to SINGLE_TREE, use the luminance motion vector mvL[0][0] as the input and the chrominance motion vector mvC[0][0] as the output, and call the chrominance motion vector derivation process in Section 8.6.2.9.
[0291] – The number of luminance coding / decoding sub-blocks numSbX in the horizontal direction and the number of luminance coding / decoding sub-blocks numSbY in the vertical direction are both set to be equal to 1.
[0292] 1. Otherwise, if treeType is equal to DUAL_TREE_CHROMA, the following applies:
[0293] – The number of luminance coding / decoding sub-blocks numSbX in the horizontal direction and the number of luminance coding / decoding sub-blocks numSbY in the vertical direction are derived as follows:
[0294] numSbX = ((cbWidth >> 2) (8 - 886)
[0295] numSbY = ((cbHeight >> 2)) (8 - 887)
[0296] – {{For xSbIdx = 0..numSbX - 1, ySbIdx = 0..numSbY - 1, the chrominance motion vector mvC[xSbIdx][ySbIdx] is derived as follows:
[0297] – The luminance motion vector mvL[xSbIdx][ySbIdx] is derived as follows:
[0298] – The position (xCuY, yCuY) of the co-located luminance coding / decoding unit is derived as follows:
[0299] xCuY = xCb + xSbIdx * 4 (0 - 1)
[0300] yCuY = yCb + ySbIdx * 4 (0 - 2)
[0301] – If CuPredMode[xCuY][yCuY] is equal to MODE_INTRA, the following applies:
[0302] mvL[xSbIdx][ySbIdx][0] = 0 (0 - 3)
[0303] mvL[xSbIdx][ySbIdx][1] = 0 (0 - 4)
[0304] predFlagL0[xSbIdx][ySbIdx] = 0 (0 - 5)
[0305] predFlagL1[xSbIdx][ySbIdx] = 0 (0 - 6)
[0306] – Otherwise (CuPredMode[xCuY][yCuY] is equal to MODE_IBC), the following applies:
[0307] mvL[xSbIdx][ySbIdx][0] = MvL0[xCuY][yCuY][0] (0 - 7)
[0308] mvL[xSbIdx][ySbIdx][1] = MvL0[xCuY][yCuY][1] (0 - 8)
[0309] predFlagL0[xSbIdx][ySbIdx] = 1 (0 - 9)
[0310] predFlagL1[xSbIdx][ySbIdx] = 0 (0 - 10)}}
[0311] – Call the derivation process of the chrominance motion vector in Section 8.6.2.9 with mvL[xSbIdx][ySbIdx] as the input and mvC[xSbIdx][ySbIdx] as the output.
[0312] – The requirement for bitstream consistency is that the chrominance motion vector mvC[xSbIdx][ySbIdx] shall comply with the following constraints:
[0313] – When calling the derivation process of block availability specified in Section 6.4.X [Edit (BB): Adjacent block availability check process tbd] with the current chrominance position (xCurr, yCurr) set to be equal to (xCb / SubWidthC, yCb / SubHeightC) and the adjacent chrominance position (xCb / SubWidthC + (mvC[xSbIdx][ySbIdx][0] >> 5), yCb / SubHeightC + (mvC[xSbIdx][ySbIdx][1] >> 5)) as the input, the output shall be equal to true.
[0314] – When the current chrominance position (xCurr, yCurr) is set to be equal to (xCb / SubWidthC, yCb / SubHeightC) and the adjacent chrominance position (xCb / SubWidthC + (mvC[xSbIdx][ySbIdx][0] >> 5) + cbWidth / SubWidthC – 1, yCb / SubHeightC + (mvC[xSbIdx][ySbIdx][1] >> 5) + cbHeight / SubHeightC - 1) is used as the input to call the export process of block availability specified in Section 6.4.X [Edit (BB): Adjacent block availability check process tbd], the output shall be equal to true.
[0315] – One or both of the following conditions shall be true:
[0316] – (mvC[xSbIdx][ySbIdx][0] >> 5>) + xSbIdx * 2 + 2 is less than or equal to 0.
[0317] – (mvC[xSbIdx][ySbIdx][1] >> 5) + ySbIdx * 2 + 2 is less than or equal to 0.
[0318] 2. The predicted samples of the current coding unit are exported as follows:
[0319] – If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the predicted samples of the current coding unit are exported as follows:
[0320] – Using the luma coding block position (xCb, yCb), luma coding block width cbWidth and luma coding block height cbHeight, the number of luma coding sub-blocks numSbX in the horizontal direction and the number of luma coding sub-blocks numSbY in the vertical direction, the luma motion vector mvL[xSbIdx][ySbIdx] (where xSbIdx = 0..numSbX - 1 and ySbIdx = 0..numSbY - 1), the variable cIdx is set to be equal to 0 as the input, and the ibc predicted samples (predSamples) of the (cbWidth) x (cbHeight) matrix predSamples L as the output, call the decoding process of the ibc block specified in Section 8.6.3.1.
[0321] – Otherwise, if treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the predicted samples of the current coding unit are exported as follows:
[0322] – With the luminance coding / decoding block position (xCb, yCb), luminance coding / decoding block width cbWidth and luminance coding / decoding block height cbHeight, the number of luminance coding / decoding sub-blocks numSbX in the horizontal direction and the number of luminance coding / decoding sub-blocks numSbY in the vertical direction, the chrominance motion vectors mvC[xSbIdx][ySbIdx] (where xSbIdx = 0..numSbX - 1 and ySbIdx = 0..numSbY - 1), the variable cIdx set to be equal to 1 as input, and the (cbWidth / 2) x (cbHeight / 2) matrix predSamples of predicted chrominance samples for the chrominance component Cb Cb The ibc predicted samples (predSamples) are used as output to call the decoding process of the ibc block specified in Section 8.6.3.1.
[0323] – With the luminance coding / decoding block position (xCb, yCb), luminance coding / decoding block width cbWidth and luminance coding / decoding block height cbHeight, the number of luminance coding / decoding sub-blocks numSbX in the horizontal direction and the number of luminance coding / decoding sub-blocks numSbY in the vertical direction, the chrominance motion vectors mvC[xSbIdx][ySbIdx] (where xSbIdx = 0..numSbX - 1 and ySbIdx = 0..numSbY – 1), the variable cIdx set to be equal to 2 as input, and the (cbWidth / 2) x (cbHeight / 2) matrix predSamples of predicted chrominance samples for the chrominance component Cr Cr The ibc predicted samples (predSamples) are used as output to call the decoding process of the ibc block specified in Section 8.6.3.1.
[0324] 3. The variables NumSbX[xCb][yCb] and NumSbY[xCb][yCb] are respectively set to be equal to numSbX and numSbY.
[0325] 4. The residual samples of the current coding / decoding unit are derived as follows:
[0326] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_LUMA, with the position (xTb0, yTb0) set to be equal to the luminance position (xCb, yCb), the width nTbW set to be equal to the luminance coding / decoding block width cbWidth, the height nTbH set to be equal to the luminance coding / decoding block height cbHeight, and the variable cIdxset set to be equal to 0 as input, and the matrix resSamples LAs output, call the decoding process of the residual signal of the coded block coded in the inter prediction mode as specified in Section 8.5.8.
[0327] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, set the position (xTb0, yTb0) to be equal to the chroma position (xCb / 2, yCb / 2), the width nTbW to be equal to the chroma coded block width cbWidth / 2, the height nTbH to be equal to the chroma coded block height cbHeight / 2, and the variable cIdxset to be equal to 1 as input, and the matrix resSamples Cb As output, call the decoding process of the residual signal of the coded block coded in the inter prediction mode as specified in Section 8.5.8.
[0328] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, set the position (xTb0, yTb0) to be equal to the chroma position (xCb / 2, yCb / 2), the width nTbW to be equal to the chroma coded block width cbWidth / 2, the height nTbH to be equal to the chroma coded block height cbHeight / 2, and the variable cIdxset to be equal to 2 as input, and the matrix resSamples Cr As output, call the decoding process of the residual signal of the coded block coded in the inter prediction mode as specified in Section 8.5.8.
[0329] 5. The reconstructed samples of the current coded unit are derived as follows:
[0330] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_LUMA, set the block position (xB, yB) to be equal to (xCb, yCb), the block width bWidth to be equal to cbWidth, the block height bHeight to be equal to cbHeight, the variable cIdx to be equal to 0, set the (cbWidth)x(cbHeight) matrix predSamples to be equal to predSamples L and set the (cbWidth)x(cbHeight) matrix resSamples to be equal to resSamples L As input, and the output is the modified reconstructed picture before loop filtering, call the picture reconstruction process of the color component as specified in Section 8.7.5.
[0331] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, set the block position (xB, yB) to be equal to (xCb / 2, yCb / 2), set the block width bWidth to cbWidth / 2, set the block height bHeight to cbHeight / 2, set the variable cIdx to 1, and set the (cbWidth / 2) x (cbHeight / 2) matrix predSamples to be equal to predSamples Cb and set the (cbWidth / 2) x (cbHeight / 2) matrix resSamples to be equal to resSamples Cb As input, and the output is the modified reconstructed picture before loop filtering, to call the picture reconstruction process of the color component specified in Section 8.7.5
[0332] – When treeType is equal to SINGLE_TREE or treeType is equal to DUAL_TREE_CHROMA, set the block position (xB, yB) to be equal to (xCb / 2, yCb / 2), set the block width bWidth to cbWidth / 2, set the block height bHeight to cbHeight / 2, set the variable cIdx to 2, and set the (cbWidth / 2) x (cbHeight / 2) matrix predSamples to be equal to predSamples Cr and set the (cbWidth / 2) x (cbHeight / 2) matrix resSamples to be equal to resSamples Cr As input, and the output is the modified reconstructed picture before loop filtering, to call the picture reconstruction process of the color component specified in Section 8.7.5
[0333] 2.3.2 Latest progress of IBC (in VTM5.0)
[0334] 2.3.2.1 Single BV list
[0335] VVC adopted JVET-N0843. In JVET-N0843, the BV predictors for the Merge mode and the AMVP mode in IBC will share a common predictor list, which contains the following elements:
[0336] · 2 spatial neighboring positions (such as A1, B1 in)
[0337] · 5 HMVP entries
[0338] · Default zero vector
[0339] The number of candidates in the list is controlled by a variable derived from the stripe head. For the Merge mode, the first 6 entries of this list will be used at most; for the AMVP mode, the first 2 entries of this list will be used. And this list meets the requirements of the shared Merge list area (sharing the same list in the SMR).
[0340] In addition to the above BV predictor candidate list, JVET-N0843 also proposed to simplify the pruning operation between the HMVP candidate and the existing Merge candidates (A1, B1). In the simplification, at most 2 pruning operations will be performed because it only compares the first HMVP candidate with the (one or more) spatial domain Merge candidates.
[0341] 2.3.2.1.1 Decoding process
[0342] 8.6.2.2 Derivation process of IBC luminance motion vector prediction
[0343] This process is called only when CuPredMode[xCb][yCb] is equal to MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luminance coding block relative to the top-left luminance sample of the current picture.
[0344] The inputs to this process are:
[0345] – The luminance position (xCb, yCb) of the top-left sample of the current luminance coding block relative to the top-left luminance sample of the current picture,
[0346] – The variable cbWidth, which specifies the width of the current coding block in luminance samples,
[0347] – The variable cbHeight, which specifies the height of the current coding block in luminance samples.
[0348] The outputs of this process are:
[0349] – The luminance motion vector in mvL with 1 / 16-fraction sample accuracy.
[0350] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows:
[0351] xSmr = IsInSmr[xCb][yCb]? SmrX[xCb][yCb] : xCb (0 - 11)
[0352] ySmr = IsInSmr[xCb][yCb]? SmrY[xCb][yCb] : yCb (0 - 12)
[0353] smrWidth = IsInSmr[xCb][yCb]? SmrW[xCb][yCb] : cbWidth (0 - 13)
[0354] smrHeight = IsInSmr[xCb][yCb]? SmrH[xCb][yCb] : cbHeight(0 - 14)
[0355] smrNumHmvpIbcCand = IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand : NumHmvpIbcCand(0 - 15)
[0356] The luminance motion vector mvL is derived through the following ordered steps:
[0357] 1. Set the luminance coding / decoding block position (xCb, yCb) to be equal to (xSmr, ySmr), set the luminance coding / decoding block width cbWidth and the height cbHeight of the luminance coding / decoding block to be equal to smrWidth and smrHeight as inputs, and the outputs are the availability flags availableFlagA1, availableFlagB1 and the motion vectors mvA1 and mvB1, and call the process of deriving spatial motion vector candidates from adjacent coding units as specified in Article 8.6.2.3.
[0358] 2. The construction of the motion vector candidate list mvCandList is as follows:
[0359] i = 0
[0360] if(availableFlagA1)
[0361] mvCandList[i++] = mvA1 (8 - 915)
[0362] if(availableFlagB1)
[0363] mvCandList[i++] = mvB1
[0364] 3. Set the variable numCurrCand to be equal to the number of Merge candidates in mvCandList.
[0365] 4. When numCurrCand is less than MaxNumMergeCand and smrNumHmvpIbcCand is greater than 0, the export process of IBC history-based motion vector candidates specified in 8.6.2.4 is called with mvCandList, isInSmr set equal to IsInSmr[xCb][yCb], and numCurrCand as the input, and the modified mvCandList and numCurrCand as the output.
[0366] 5. When numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand is equal to MaxNumMergeCand:
[0367] 1. Set mvCandList[numCurrCand][0] equal to 0.
[0368] 2. Set mvCandList[numCurrCand][1] equal to 0.
[0369] 3. Increment numCurrCand by 1.
[0370] 6. The variable mvIdx is derived as follows:
[0371] mvIdx = general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb] : mvp_l0_flag[xCb][yCb](0 - 17)
[0372] 7. The following assignments are made:
[0373] mvL[0] = mergeCandList[mvIdx][0] (0 - 18)
[0374] mvL[1] = mergeCandList[mvIdx][1] (0 - 19)
[0375] 2.3.2.2 Size Limitations of IBC
[0376] In the latest VVC and VTM5, it is recommended to explicitly use syntax constraints to disable the 128x128 IBC mode on top of the current bitstream constraints in previous VTM and VVC versions, which makes the presence of the IBC flag dependent on the CU size < 128x128.
[0377] 2.3.2.3 Shared Merge List of IBC
[0378] To reduce the decoder complexity and support parallel encoding, JVET-M0147 proposed that all leaf codec units (CUs) of an ancestor node in the CU partition tree share the same Merge candidate list to enable parallel processing of small skip / Merge CUs. The ancestor node is called the Merge sharing node. Assuming the Merge sharing node is a leaf CU, a shared Merge candidate list is generated at the Merge sharing node.
[0379] More specifically, the following can be applied:
[0380] - If a block has luminance samples no greater than 32 and is partitioned into 2 4x4 child blocks, a shared Merge list between very small blocks (e.g., two adjacent 4x4 blocks) is used.
[0381] - If a block has luminance samples greater than 32, but after partitioning, at least one sub-block is smaller than the threshold (32), all child blocks of this partition share the same Merge list (e.g., 16x4 or 4x16 ternary partition or 8x8 quaternary partition).
[0382] Such a restriction is only applied to the IBC Merge mode.
[0383] 2.4 Syntax Tables and Semantics of Codec Units and Merge Modes
[0384] 7.3.5.1 General Slice Header Syntax
[0385]
[0386]
[0387] 7.3.7.5 Codec Unit Syntax
[0388]
[0389]
[0390]
[0391] 7.3.7.7 Merge Data Syntax
[0392]
[0393]
[0394] 7.4.6.1 General Slice Header Semantics
[0395] six_minus_max_num_merge_cand specifies the maximum number of Merge motion vector prediction (MVP) candidates supported in the slice subtracted from 6. The maximum number of Merge MVP candidates, MaxNumMergeCand, is derived as follows:
[0396] MaxNumMergeCand = 6 - six_minus_max_num_merge_cand (7-57)
[0397] The value of MaxNumMergeCand shall be in the range of 1 to 6, inclusive of the end values.
[0398] five_minus_max_num_subblock_merge_cand specifies the maximum number of sub-block based Merge motion vector prediction (MVP) candidates supported in the slice subtracted from 5. When five_minus_max_num_subblock_merge_cand does not exist, it is inferred to be equal to 5 - sps_sbtmvp_enabled_flag. The maximum number of sub-block based Merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows:
[0399] MaxNumSubblockMergeCand = 5 - five_minus_max_num_subblock_merge_cand (7-58)
[0400] The value of MaxNumSubblockMergeCand shall be in the range of 0 to 5, inclusive of the end values.
[0401] 7.4.8.5 Coding Unit Semantics
[0402] pred_mode_flag being equal to 0 specifies that the current coding unit is coded in the inter prediction mode. pred_mode_flag being equal to 1 specifies that the current coding unit is coded in the intra prediction mode.
[0403] When pred_mode_flag does not exist, it can be inferred as follows:
[0404] – If cbWidth is equal to 4 and cbHeight is equal to 4, then pred_mode_flag is inferred to be equal to 1.
[0405] – Otherwise, when decoding an I slice, pred_mode_flag is inferred to be equal to 1, and when decoding a P or B slice respectively, pred_mode_flag is inferred to be equal to 0.
[0406] For x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1, the variable CuPredMode[x][y] is derived as follows:
[0407] – If pred_mode_flag is equal to 0, then set CuPredMode[[x][y] to be equal to MODE_INTER.
[0408] – Otherwise (pred_mode_flag is equal to 1), set CuPredMode[x][y] to be equal to MODE_INTRA.
[0409] pred_mode_ibc_flag being equal to 1 indicates that the current coding / decoding unit is coded / decoded in the IBC prediction mode. pred_mode_ibc_flag being equal to 0 indicates that the current coding / decoding unit is not coded / decoded in the IBC prediction mode.
[0410] When there is no pred_mode_ibc_flag, it can be inferred as follows:
[0411] – If cu_skip_flag[x0][y0] is equal to 1, and cbWidth is equal to 4, and cbHeight is equal to 4, then pred_mode_ibc_flag is inferred to be equal to 1.
[0412] – Otherwise, if both cbWidth and cbHeight are equal to 128, then pred_mode_ibc_flag is inferred to be equal to 0.
[0413] – Otherwise, when decoding an I slice, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag; and when decoding a P or B slice respectively, the value of pred_mode_ibc_flag is inferred to be equal to 0.
[0414] When pred_mode_ibc_flag is equal to 1, for x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1, set the variable CuPredMode[x][y] to be equal to MODE_IBC.
[0415] general_merge_flag[x0][y0] specifies whether the inter prediction parameters of the current coding / decoding unit are inferred from adjacent inter prediction partitions. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding / decoding block under consideration relative to the top-left luma sample of the picture.
[0416] When general_merge_flag[x0][y0] does not exist, the following inference can be made:
[0417] – If cu_skip_flag[x0][y0] is equal to 1, then infer that general_merge_flag[x0][y0] is equal to 1.
[0418] – Otherwise, infer that general_merge_flag[x0][y0] is equal to 0.
[0419] mvp_l0_flag[x0][y0] specifies the motion vector predictor index for list 0, where x0 and y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma picture sample of the picture.
[0420] When mvp_l0_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0421] mvp_l1_flag[x0][y0] has the same semantics as mvp_l0_flag, where l0 and list 0 are replaced by l1 and list 1 respectively.
[0422] inter_pred_idc[x0][y0] specifies whether list 0, list 1, or bi-prediction is used for the current coding unit according to Table 7-10. The matrix indices x0 and y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma picture sample of the picture.
[0423] Table 7-10 – Names associated with inter prediction modes
[0424]
[0425] When inter_pred_idc[x0][y0] does not exist, it is inferred to be equal to PRED_L0.
[0426] 7.4.8.7 Merge data semantics
[0427] regular_merge_flag[x0][y0] being equal to 1 specifies that the inter prediction parameters for the current coding unit are generated using the regular Merge mode. The matrix indices x0 and y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma picture sample of the picture.
[0428] When regular_merge_flag[x0][y0] does not exist, the following inferences can be made:
[0429] – If all of the following conditions are true, then regular_merge_flag[x0][y0] is inferred to be equal to 1:
[0430] – sps_mmvd_enabled_flag is equal to 0.
[0431] – general_merge_flag[x0][y0] is equal to 1.
[0432] – cbWidth * cbHeight is equal to 32.
[0433] – Otherwise, regular_merge_flag[x0][y0] is inferred to be equal to 0.
[0434] mmvd_merge_flag[x0][y0] being equal to 1 specifies that the inter-prediction parameters of the current coding unit are generated using the Merge mode with motion vector difference. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0435] When mmvd_merge_flag[x0][y0] does not exist, the following inferences can be made:
[0436] – If all of the following conditions are true, then mmvd_merge_flag[x0][y0] is inferred to be equal to 1:
[0437] – sps_mmvd_enabled_flag is equal to 1.
[0438] – general_merge_flag[x0][y0] is equal to 1.
[0439] – cbWidth * cbHeight is equal to 32.
[0440] – regular_merge_flag[x0][y0] is equal to 0.
[0441] – Otherwise, mmvd_merge_flag[x0][y0] is inferred to be equal to 0.
[0442] mmvd_cand_flag[x0][y0] specifies whether the first (0) or second (1) candidate in the Merge candidate list is used together with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0443] When mmvd_cand_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0444] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0], as shown in Table 7-12. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0445] Table 7-12 – Specification of MmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0]
[0446]
[0447]
[0448] mmvd_direction_idx[x0][y0] specifies the index used to derive MmvdSign[x0][y0], as specified in Table 7-13. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0449] Table 7-13 – Specification of MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]
[0450]
[0451] The two components MmvdOffset[x0][y0] of the Merge plus MVD offset are derived as follows:
[0452] MmvdOffset[x0][y0][0] = (MmvdDistance[x0][y0] << 2) * MmvdSign[x0][y0][0] (0 - 21)
[0453] MmvdOffset[x0][y0][1] = (MmvdDistance[x0][y0] << 2) * MmvdSign[x0][y0][1](0 - 22)
[0454] merge_subblock_flag[x0][y0] specifies whether to infer sub - block - based inter - prediction parameters of the current coding unit from adjacent blocks. Matrix indices x0, y0 specify the position (x0, y0) of the top - left luma sample of the coding block under consideration relative to the top - left luma sample of the picture. When merge_subblock_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0455] merge_subblock_idx[x0][y0] specifies the Merge candidate index of the sub - block - based Merge candidate list, where x0, y0 specify the position (x0, y0) of the top - left luma sample of the coding block under consideration relative to the top - left luma sample of the picture.
[0456] When merge_subblock_idx[x0][y0] does not exist, it is inferred to be equal to 0.
[0457] ciip_flag[x0][y0] specifies whether to apply combined inter - picture Merge and intra - picture prediction to the current coding unit. Matrix indices x0, y0 specify the position (x0, y0) of the top - left luma sample of the coding block under consideration relative to the top - left luma sample of the picture.
[0458] When ciip_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0459] When ciip_flag[x0][y0] is equal to 1, the variable IntraPredModeY[x][y] with x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1 is set to be equal to INTRA_PLANAR.
[0460] The variable MergeTriangleFlag[x0][y0] that specifies whether to use triangle - based motion compensation to generate prediction samples of the current coding unit when decoding B - slices is derived as follows:
[0461] – If all of the following conditions are true, then MergeTriangleFlag[x0][y0] is set to be equal to 1:
[0462] – sps_triangle_enabled_flag is equal to 1.
[0463] – The slice_type is equal to B.
[0464] – The general_merge_flag[x0][y0] is equal to 1.
[0465] – The MaxNumTriangleMergeCand is greater than or equal to 2.
[0466] – The cbWidth * cbHeight is greater than or equal to 64.
[0467] – The regular_merge_flag[x0][y0] is equal to 0.
[0468] – The mmvd_merge_flag[x0][y0] is equal to 0.
[0469] – The merge_subblock_flag[x0][y0] is equal to 0.
[0470] – The ciip_flag[x0][y0] is equal to 0.
[0471] – Otherwise, set the MergeTriangleFlag[x0][y0] to be equal to 0.
[0472] The merge_triangle_split_dir[x0][y0] specifies the splitting direction of the Merge triangle mode. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding / decoding block relative to the top-left luma sample of the picture.
[0473] When there is no merge_triangle_split_dir[x0][y0], it is inferred to be equal to 0.
[0474] The merge_triangle_idx0[x0][y0] specifies the first Merge candidate index of the triangle-based motion compensation candidate list, where x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding / decoding block relative to the top-left luma sample of the picture.
[0475] When there is no merge_triangle_idx0[x0][y0], it is inferred to be equal to 0.
[0476] The merge_triangle_idx1[x0][y0] specifies the second Merge candidate index of the triangle-based motion compensation candidate list, where x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered coding / decoding block relative to the top-left luma sample of the picture.
[0477] When merge_triangle_idx1[x0][y0] does not exist, it is inferred to be equal to 0.
[0478] merge_idx[x0][y0] specifies the Merge candidate index in the Merge candidate list, where x0 and y0 specify the position (x0, y0) of the upper-left luma sample of the coding block under consideration relative to the upper-left luma sample of the picture.
[0479] When merge_idx[x0][y0] does not exist, it can be inferred as follows:
[0480] – If mmvd_merge_flag[x0][y0] is equal to 1, then merge_idx[x0][y0] is inferred to be equal to mmvd_cand_flag[x0][y0].
[0481] – Otherwise (mmvd_merge_flag[x0][y0] is equal to 0), then merge_idx[x0][y0] is inferred to be equal to 0.
[0482] 3 Examples of technical problems solved by the embodiments
[0483] Current IBC may have the following problems:
[0484] 1. For P / B slices, the number of IBC motion (BV) candidates that can be added to the IBC motion list is set to be the same as the conventional Merge list size. Therefore, if the conventional Merge mode is disabled, the IBC Merge mode will also be disabled. However, considering its huge benefits, it is desired to enable the IBC Merge mode even if the conventional Merge mode is disabled.
[0485] 2. IBC AMVP and IBC Merge mode share the same process of constructing the motion (BV) candidate list. The list size indicated by the maximum number of Merge candidates (e.g., MaxNumMergeCand) is signaled in the slice header. For the case of IBC AMVP, the BV predictor can only be selected from one of the two IBC motion candidates.
[0486] a. When the BV candidate list size is 1, in this case, signaling of the BV predictor index is not required.
[0487] b. When the BV candidate list size is 0, both IBC Merge and AMVP modes should be disabled. However, in the current design, the BV predictor index is still signaled.
[0488] 3. IBC can be enabled at the sequence level. However, due to the size of the BV list signaled, IBC can be disabled, but the IBC mode and indications regarding the syntax are still signaled, which wastes bits.
[0489] 4 Example Techniques and Embodiments
[0490] The following detailed inventions should be considered as examples to explain the general concepts. These inventions should not be interpreted narrowly. In addition, these inventions can be combined in any way.
[0491] In the present invention, decoder-side motion vector derivation (DMVD) includes DMVR and FRUC that perform motion estimation to derive or refine block / sub-block motion information, and methods such as BIO that perform motion refinement in terms of samples.
[0492] The maximum number of BV candidates (i.e., the size of the BV candidate list) can be represented by maxIBCCandNum based on the BV derived or predicted in the BV candidate list.
[0493] The maximum number of IBC Merge candidates is represented as maxIBCMrgNum, and the maximum number of IBC AMVP candidates is represented as maxIBCAMVPNum. Note that the skip mode can be regarded as a special Merge mode where all coefficients are equal to 0.
[0494] 1. Even if IBC is enabled for the sequence, IBC (e.g., IBC AMVP and / or IBC Merge) mode may be disabled for pictures / strips / slices / slice groups / tile groups or other video units.
[0495] a. In one example, an indication of whether IBC AMVP and / or IBC Merge mode is enabled can be signaled at the picture / strip / slice / slice group / tile group or other video unit level (such as in PPS, APS, strip headers, picture headers, etc.).
[0496] i. In one example, the indication can be a flag.
[0497] ii. In one example, the indication can be a plurality of allowed BV predictors. When the number of allowed BV predictors is equal to 0, it indicates that the IBC AMVP and / or Merge mode is disabled.
[0498] 1) Alternatively, the number of allowed BV predictors can be signaled in a predictive manner.
[0499] a. For example, it can be signaled that X minus the number of allowed BV predictors, where X is a fixed number such as 2.
[0500] iii. Alternatively, in addition, the indication may be signaled conditionally, e.g., based on whether the video content is screen content or the sps_ibc_enabled_flag.
[0501] b. In one example, when the video content is not screen content (e.g., camera-captured or mixed content), the IBC mode may be disabled.
[0502] c. In one example, when the video content is camera-captured content, the IBC mode may be disabled.
[0503] 2. Whether to signal the use of IBC (e.g., pred_mode_ibc_flag) and IBC-related syntax elements may depend on maxIBCCandNum (e.g., maxIBCCandNum is set to be equal to MaxNumMergeCand).
[0504] a. In one example, when maxIBCCandNum is equal to 0, signaling for using the IBC skip mode (e.g., cu_skip_flag) for I slices may be skipped, and it is inferred that IBC is disabled.
[0505] b. In one example, when maxIBCMrgNum is equal to 0, signaling for using the IBC skip mode (e.g., cu_skip_flag) for I slices may be skipped, and it is inferred that the IBC skip mode is disabled.
[0506] c. In one example, when maxIBCCandNum is equal to 0, signaling for using the IBC Merge / IBC AMVP mode (e.g., pred_mode_ibc_flag) may be skipped, and it is inferred that IBC is disabled.
[0507] d. In one example, when maxIBCCandNum is equal to 0 and the current block is coded in the IBC mode, signaling for the Merge mode (e.g., general_merge_flag) may be skipped, and it is inferred that the IBC Merge mode is disabled.
[0508] i. Alternatively, it is inferred that the current block is coded in the IBC AMVP mode.
[0509] e. In one example, when maxIBCMrgNum is equal to 0 and the current block is coded in the IBC mode, signaling for the Merge mode (e.g., general_merge_flag) may be skipped, and it is inferred that the IBC Merge mode is disabled.
[0510] i. Alternatively, it is inferred that the current block is coded in IBC AMVP mode.
[0511] 3. Whether the syntax elements related to the motion vector difference in IBC AMVP mode are signaled can depend on maxIBCCandNum (e.g., maxIBCCandNum is set to be equal to MaxNumMergeCand).
[0512] a. In one example, when maxIBCCandNum is equal to 0, signaling of the motion vector predictor index (e.g., mvp_10_flag) in IBC AMVP mode can be skipped, and it is inferred that the IBC AMVP mode is disabled.
[0513] b. In one example, when maxIBCCandNum is equal to 0, signaling of the motion vector difference (e.g., mvd_coding) can be skipped, and it is inferred that the IBC AMVP mode is disabled.
[0514] c. In one example, under the condition that maxIBCCandNum is greater than K (e.g., K = 0 or 1), the motion vector predictor index and / or the precision of the motion vector predictor and / or the precision of the motion vector difference in IBC AMVP mode can be coded.
[0515] i. In one example, if maxIBCCandNum is equal to 1, the motion vector predictor index (e.g., mvp_10_flag) in IBC AMVP mode may not be signaled.
[0516] 1) In one example, in this case, the motion vector predictor index (e.g., mvp_10_flag) in IBC AMVP mode can be inferred to be a value such as 0.
[0517] ii. In one example, if maxIBCCandNum is equal to 0, signaling of the precision of the motion vector predictor and / or the precision of the motion vector difference (e.g., amvr_precision_flag) for IBC AMVP mode can be skipped.
[0518] iii. In one example, the precision of the motion vector predictor and / or the precision of the motion vector difference (e.g., amvr_precision_flag) in IBC AMVP mode can be signaled under the condition that maxIBCCandNum is greater than 0.
[0519] 4. It is recommended that maxIBCCandNum can be decoupled from the maximum number of regular Merge candidates.
[0520] a. In one example, maxIBCCandNum can be signaled directly.
[0521] i. Alternatively, when IBC is enabled for a strip, the consistency bitstream shall satisfy that maxIBCCandNum is greater than 0.
[0522] b. In one example, maxIBCCandNum and the predictive coding of other syntax elements / fixed values can be signaled.
[0523] i. In one example, the difference between the regular Merge list size and maxIBCCandNum can be coded.
[0524] ii. In one example, (K minus maxIBCCandNum) can be coded, for example, K = 5 or 6.
[0525] iii. In one example, (maxIBCCandNum minus K) can be coded, for example, K = 0 or 2.
[0526] c. In one example, the indication of maxIBCMrgNum and / or maxIBCAMVPNum can be signaled according to the above method.
[0527] 5. maxIBCCandNum can be set to be equal to Func(maxIBCMrgNum, maxIBCAMVPNum)
[0528] a. In one example, maxIBCAMVPNum is fixed to 2, and maxIBCCandNum is set to Func(maxIBCMrgNum, 2).
[0529] b. In one example, Func(a, b) returns the larger value between the two variables a and b.
[0530] 6. maxIBCCandNum can be determined according to the coding mode information of a block. That is, how many BV candidates can be added to the BV candidate list can depend on the mode information of the block.
[0531] a. In one example, if a block is coded in IBC Merge mode, maxIBCCandNum can be set to be equal to maxIBCMrgNum.
[0532] b. In one example, if a block is coded in IBC AMVP mode, maxIBCCandNum can be set to be equal to maxIBCAMVPNum.
[0533] 7. The compliant bitstream shall satisfy that the decoded IBC AMVP or IBC Merge index is less than maxIBCCandNum.
[0534] a. In one example, the compliant bitstream shall satisfy that the decoded IBC AMVP index is less than maxIBCAMVPNum.
[0535] b. In one example, the compliant bitstream shall satisfy that the decoded IBC Merge index is less than maxIBCMrgNum (e.g., 2).
[0536] 8. When the IBC AMVP or IBC Merge candidate index cannot identify a BV candidate in the BV candidate list (e.g., the IBC AMVP or Merge candidate index is not less than maxIBCCandNum, or the decoded IBC AMVP index is not less than maxIBCAMVPNum, or the decoded IBC Merge index is not less than maxIBCMrgNum), a default prediction block may be used.
[0537] a. In one example, all samples in the default prediction block are set to (1 << (bit depth - 1)).
[0538] b. In one example, the default BV may be assigned to the block.
[0539] 9. When the IBC AMVP or IBC Merge candidate index cannot identify a BV candidate in the BV candidate list (e.g., the IBC AMVP or Merge candidate index is not less than maxIBCCandNum, or the decoded IBC AMVP index is not less than maxIBCAMVPNum, or the decoded IBC Merge index is not less than maxIBCMrgNum), the block may be regarded as an IBC block with an invalid BV.
[0540] a. In one example, the process applied to the block with an invalid BV may be applied to this block.
[0541] 10. A supplementary BV candidate list may be constructed under certain conditions, such as when the decoded IBC AMVP or IBC Merge index is not less than maxIBCCandNum.
[0542] a. In one example, the supplementary BV candidate list is constructed by one or more of the following steps (sequentially or in an interleaved manner):
[0543] i. Add HMVP candidates
[0544] 1) Starting from the K-th entry in the HMVP table in ascending order of the HMVP candidate index (e.g., K = 0)
[0545] 2) Starting from the K-th entry from the bottom in the HMVP table in descending order of the HMVP candidate index (e.g., K = 0 or K = maxIBCCandNum – IBC AMVP / IBC Merge index, or K = maxIBCCandNum – 1 – IBC AMVP / IBC Merge index)
[0546] ii. Derive virtual BV candidates from the available candidates using a BV candidate list with maxIBCCandNum candidates.
[0547] 1) In one example, an offset can be added to the horizontal and / or vertical components of the BV candidate to obtain a virtual BV candidate.
[0548] 2) In one example, an offset can be added to the horizontal and / or vertical components of the HMVP candidate in the HMVP table to obtain a virtual BV candidate.
[0549] iii. Add default candidates
[0550] 1) In one example, (0, 0) can be added as a default candidate.
[0551] 2) In one example, the default candidate can be derived based on the dimensions of the current block.
[0552] 11. maxIBCAMVPNum may not be equal to 2.
[0553] a. Alternatively, in addition, an index that can be greater than 1 can be signaled instead of signaling a flag (e.g., mvp_10_flag) that indicates a motion vector predictor index.
[0554] i. In one example, the index can be binarized using unary / truncated unary / fixed length / exponential Golomb / other binarization methods.
[0555] ii. In one example, the binary digits (bin) of the binary string of the index can be context decoded or bypass decoded.
[0556] 1) In one example, the first K (e.g., K = 1) bits of the binary digit string of the index can be context decoded, and the remaining bits can be bypass decoded.
[0557] b. In one example, maxIBCAMVPNum can be greater than maxIBCMrgNum.
[0558] i. Alternatively, in addition, the first maxIBCMrgNum BV candidates in the BV candidate list can be used for IBC Merge coded / decoded blocks.
[0559] 12. Even if maxIBCCandNum is set to 0 (e.g., MaxNumMergeCand = 0), the IBC BV candidate list can still be constructed.
[0560] a. In one example, the Merge list can be constructed when the IBC AMVP mode is enabled.
[0561] b. In one example, when the IBC Merge mode is disabled, the Merge list can be constructed to include up to maxIBCAMVPNum BV candidates.
[0562] 13. For the methods disclosed above, the term "maxIBCCandNum" can be replaced by "maxIBCMrgNum" or "maxIBCAMVPNum".
[0563] 14. For the methods disclosed above, the term "maxIBCCandNum" and / or "maxIBCMrgNum" can be replaced by MaxNumMergeCand, which can represent the maximum number of Merge candidates in a regular Merge list.
[0564] 5 Embodiments
[0565] The newly added part in JVET-N1001-v5 is enclosed in double bold curly braces, i.e., {{a}} indicates that "a" has been added, and the deleted part is enclosed in double bold brackets, i.e., [[a]] means that "a" has been deleted.
[0566] 5.1 Embodiment #1
[0567] The indication of the maximum number of IBC motion candidates (e.g., for IBC AMVP and / or IBC Merge) can be signaled in the slice header / PPS / APS / DPS.
[0568] 7.3.5.1 General slice header syntax
[0569]
[0570]
[0571] {{max_num_merge_cand_minus_max_num_IBC_cand specifies the maximum number of IBC Merge mode candidates supported in the slices to be subtracted from MaxNumMergeCand. The maximum number of IBC Merge mode candidates, MaxNumIBCMergeCand, is derived as follows:
[0572] MaxNumIBCMergeCand = MaxNumMergeCand - max_num_merge_cand_minus_max_num_IBC_cand
[0573] When max_num_merge_cand_minus_max_num_IBC_cand exists, the value of MaxNumIBCMergeCand shall be in the range of 2 (or 0) to MaxNumMergeCand, inclusive. When max_num_merge_cand_minus_max_num_IBC_cand does not exist, MaxNumIBCMergeCand is set to be equal to 0. When MaxNumIBCMergeCand is equal to 0, the current slice is not allowed to use the IBC Merge mode and the IBC AMVP mode.}}
[0574] Alternatively, the signaling notification of the indication of MaxNumIBCMergeCand can be replaced with:
[0575]
[0576] MaxNumIBCMergeCand = 5 - five_minus_max_num_IBC_cand}}
[0577] 5.2 Example #2
[0578] It is proposed to change the maximum number of the IBC BV list from MaxNumMergeCand that controls both IBC and the regular inter prediction mode to a separate variable, MaxNumIBCMergeCand.
[0579] 8.6.2.2 Derivation process of IBC luminance motion vector prediction
[0580] This process is called only when CuPredMode[xCb][yCb] is equal to MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture.
[0581] The inputs to this process are:
[0582] – The luminance position (xCb, yCb) of the top-left sample of the current luminance coding block relative to the top-left luminance sample of the current picture.
[0583] – The variable cbWidth, which specifies the width of the current coding block in the luminance samples.
[0584] – The variable cbHeight, which specifies the height of the current coding block in the luminance samples.
[0585] The output of this process is:
[0586] – The luminance motion vector of the 1 / 16 fractional sample accuracy mvL.
[0587] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows:
[0588] xSmr = IsInSmr[xCb][yCb]? SmrX[xCb][yCb] : xCb (8 - 910)
[0589] ySmr = IsInSmr[xCb][yCb]? SmrY[xCb][yCb] : yCb (8 - 911)
[0590] smrWidth = IsInSmr[xCb][yCb]? SmrW[xCb][yCb] : cbWidth
[0591] (8 - 912)
[0592] smrHeight = IsInSmr[xCb][yCb]? SmrH[xCb][yCb] : cbHeight
[0593] (8 - 913)
[0594] smrNumHmvpIbcCand = IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand : NumHmvpIbcCand
[0595] (8 - 914)
[0596] The luminance motion vector mvL is derived through the following ordered steps:
[0597] 1. Call the process of deriving spatial motion vector candidates from adjacent coding units as specified in Clause 8.6.2.3, with the luminance coding block position (xCb, yCb) set to be equal to (xSmr, ySmr), the luminance coding block width cbWidth and the luminance coding block height cbHeight set to be equal to smrWidth and smrHeight as inputs, and the outputs being the availability flags availableFlagA1, availableFlagB1 and the motion vectors mvA1 and mvB1.
[0598] 2. The construction of the motion vector candidate list mvCandList is as follows:
[0599] i = 0
[0600] if (availableFlagA1)
[0601] mvCandList[i++] = mvA1 (8 - 915)
[0602] if (availableFlagB1)
[0603] mvCandList[i++] = mvB1
[0604] 3. The variable numCurrCand is set to be equal to the number of Merge candidates in mvCandList.
[0605] 4. When numCurrCand is less than [[MaxNumMergeCand]]{{MaxNumIBCMergeCand}} and smrNumHmvpIbcCand is greater than 0, call the process of deriving motion vector candidates based on IBC history as specified in 8.6.2.4, with mvCandList, isInSmr set to be equal to IsInSmr[xCb][yCb] and numCurrCand as inputs, and the modified mvCandList and numCurrCand as outputs.
[0606] 5. When numCurrCand is less than [[MaxNumMergeCand]]{{MaxNumIBCMergeCand}}, the following applies until numCurrCand is equal to MaxNumMergeCand:
[0607] 1. mvCandList[numCurrCand][0] is set to be equal to 0.
[0608] 2. Set mvCandList[numCurrCand][1] to be equal to 0.
[0609] 3. Increment numCurrCand by 1.
[0610] 6. The variable mvIdx is derived as follows:
[0611] mvIdx = general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb] : mvp_l0_flag[xCb][yCb](0 - 29)
[0612] 7. Perform the following assignments:
[0613] mvL[0] = mergeCandList[mvIdx][0](8 - 917)
[0614] mvL[1] = mergeCandList[mvIdx][1](8 - 918)
[0615] In one example, if the current block is in IBC Merge mode, set MaxNumIBCMergeCand to MaxNumMergeCand; and if the current block is in IBC AMVP mode, set MaxNumIBCMergeCand to 2.
[0616] In one example, such as using Embodiment #1, derive MaxNumIBCMergeCand from the signaled information.
[0617] 5.3 Embodiment #3
[0618] Conditionally signal IBC-related syntax elements according to the maximum number of allowed IBC candidates. In one example, MaxNumIBCMergeCand is set to be equal to MaxNumMergeCand.
[0619] 7.3.7.5 Coding / Decoding Unit Syntax
[0620]
[0621]
[0622] cu_skip_flag[x0][y0] being equal to 1 specifies that for the current coding unit, when decoding a P or B slice, after cu_skip_flag[x0][y0], no other syntax elements are parsed except for one or more of the following syntax elements: the IBC mode flag pred_mode_ibc_flag[x0][y0] {{if MaxNumIBCMergeCand is greater than 0}}, and the merge_data() syntax structure; when decoding an I slice {{and MaxNumIBCMergeCand is greater than 0}}, no other syntax elements are parsed after cu_skip_flag[x0][y0] except for merge_idx[x0][y0]. cu_skip_flag[x0][y0] being equal to 0 indicates that the coding unit is not skipped. The matrix indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0623] When cu_skip_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0624] pred_mode_ibc_flag being equal to 1 specifies that the current coding unit is coded in the IBC prediction mode. pred_mode_ibc_flag being equal to 0 specifies that the current coding unit is not coded in the IBC prediction mode.
[0625] When pred_mode_ibc_flag does not exist, it can be inferred as follows:
[0626] – If cu_skip_flag[x0][y0] is equal to 1, and cbWidth is equal to 4, and cbHeight is equal to 4, then pred_mode_ibc_flag is inferred to be equal to 1.
[0627] – Otherwise, if both cbWidth and cbHeight are equal to 128, then pred_mode_ibc_flag is inferred to be equal to 0.
[0628] – Otherwise, when decoding an I slice {{and MaxNumIBCMergeCand is greater than 0}}, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag, and when decoding a P or B slice respectively, pred_mode_ibc_flag is inferred to be equal to 0.
[0629] 5.4 Example #3
[0630] Conditionally signal IBC-related syntax elements according to the maximum number of allowed IBC candidates.
[0631]
[0632]
[0633] In one example, MaxNumIBCAMVPCand can be set to MaxNumIBCMergeCand or MaxNumMergeCand.
[0634] An example method 2000 for video processing is shown. Method 2000 includes: at operation 2002, convert between a video region of a video and a bitstream representation of the video, the bitstream representation selectively including syntax elements related to motion vector differences (MVDs) of intra block copy (IBC) advanced motion vector prediction (AMVP) modes based on a maximum number of a first type of IBC candidates used during the conversion of the video region. In some embodiments, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0635] An example method 2005 for video processing is shown. Method 2005 includes: at operation 2007, determine an indication to disable the use of the intra block copy (IBC) mode for a video region of a video and enable the use of the IBC mode at the sequence level of the video for the conversion between the video region of the video and the bitstream representation of the video.
[0636] Method 2005 includes, at operation 2009, perform the conversion based on the determination. In some embodiments, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0637] An example method 2010 for video processing is shown. Method 2010 includes: at operation 2012, convert between a video region of a video and a bitstream representation of the video, the bitstream representation selectively including an indication regarding the use of the intra block copy (IBC) mode and / or one or more IBC-related syntax elements based on a maximum number of a first type of IBC candidates used during the conversion of the video region. In some embodiments, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0638] An example method 2020 for video processing is shown. Method 2020 includes: at operation 2022, converting between a video region of a video and a bitstream representation of the video, and signaling in the bitstream representation an indication of a maximum number of intra block copy (IBC) candidates of a first type used during the conversion of the video region, independent of a maximum number of Merge candidates of an inter prediction mode used during the conversion. In some embodiments, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0639] An example method 2030 for video processing is shown. Method 2030 includes: at operation 2032, converting between a video region of a video and a bitstream representation of the video, and a maximum number of intra block copy (IBC) motion candidates used during the conversion of the video region is a function of a maximum number of IBC Merge candidates and a maximum number of IBC advanced motion vector prediction (AMVP) candidates. In some embodiments, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0640] An example method 2040 for video processing is shown. Method 2040 includes: at operation 2042, converting between a video region of a video and a bitstream representation of the video, and a maximum number of intra block copy (IBC) motion candidates used during the conversion of the video region is based on coding / decoding mode information of the video region.
[0641] An example method 2050 for video processing is shown. Method 2050 includes: at operation 2052, converting between a video region of a video and a bitstream representation of the video, and a decoded intra block copy (IBC) advanced motion vector prediction (AMVP) Merge index or a decoded IBC Merge index is less than a maximum number of intra block copy (IBC) motion candidates.
[0642] An example method 2060 for video processing is shown. Method 2060 includes: during the conversion between a video region of a video and a bitstream representation of the video at operation 2062, determining that an intra block copy (IBC) alternative motion vector predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list.
[0643] Method 2060 includes: at operation 2064, using a default prediction block during the conversion based on the determination.
[0644] An example method 2070 for video processing is shown. Method 2070 includes: at operation 2072, during the conversion between a video region of a video and a bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list.
[0645] Method 2070 includes: at operation 2074, performing the conversion by treating the video region as having an invalid block vector based on the determination.
[0646] An example method 2080 for video processing is shown. Method 2080 includes: at operation 2082, during the conversion between a video region of a video and a bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to meet a condition.
[0647] Method 2080 includes: at operation 2084, generating a Supplementary Block Vector (BV) candidate list based on the determination.
[0648] Method 2080 includes: at operation 2086, performing the conversion using the supplementary BV candidate list.
[0649] An example method 2090 for video processing is shown. Method 2090 includes: at operation 2092, performing a conversion between a video region of a video and a bitstream representation of the video, where the maximum number of Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) candidates is not equal to 2.
[0650] is a block diagram of a video processing device 2100. Device 2100 can be used to implement one or more methods described herein. Device 2100 can be implemented in a smart phone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 2100 can include one or more processors 2102, one or more memories 2104, and video processing hardware 2106. The (one or more) processors 2102 can be configured to implement one or more methods described herein. The memory (multiple memories) 2104 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2106 can be used to implement some of the techniques described herein in hardware circuits. The video processing hardware 2106 can be included, in part or in whole, in the (one or more) processors 2102 in the form of dedicated hardware or an Image Processing Unit (GPU) or a specialized signal processing block.
[0651] In some embodiments, a video encoding / decoding method may be implemented using a device implemented on a hardware platform as described with respect to the disclosed technology.
[0652] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode when processing video blocks, but not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the conversion from video blocks to the bitstream representation of the video will use the video processing tool or mode when it is determined or decided to enable the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream when it is already known that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on the determination or decision.
[0653] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video blocks to the bitstream representation of the video. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream when it is already known that the bitstream has not been modified using the video processing tool or mode enabled based on the determination or decision.
[0654] FIG. 2200 is a block diagram illustrating an example video processing system 2200 in which various technologies disclosed herein may be implemented. Various implementations may include some or all components of system 2200. System 2200 may include an input 2202 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 2202 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical network (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0655] System 2200 may include a codec component 2204 that may implement various codec or encoding methods described herein. The codec component 2204 may reduce the average bit rate of the video from the input 2202 to the output of the codec component 2204 to produce a coded representation of the video. Thus, codec techniques are sometimes referred to as video compression or video transcoding techniques. As represented by component 2206, the output of the codec component 2204 may be stored or transmitted via the connected communication. The stored or transmitted bitstream (or coded) representation of the video received at the input 2202 may be used by component 2208 to generate pixel values or viewable video that is sent to the display interface 2210. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Additionally, although some video processing operations are referred to as "codec" operations or tools, it should be understood that codec tools or operations are used at the encoder and corresponding decoding tools or operations that reverse the coded result will be performed by the codec.
[0656] Examples of a peripheral bus interface or a display interface may include a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI), or a Displayport, etc. Examples of a storage interface include SATA (Serial Advanced Technology Attachment), PCI, IDE interface, etc. The techniques described herein may be implemented in various electronic devices, such as mobile phones, laptop computers, smart phones, or other devices capable of performing digital data processing and / or video display.
[0657] In some embodiments, the following technical solutions may be implemented:
[0658] A1. A video processing method, comprising: converting between a video region of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes syntax elements related to a motion vector difference (MVD) of an intra block copy (IBC) advanced motion vector prediction (AMVP) mode based on a maximum number of a first type of IBC candidate used during the conversion of the video region, and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0659] A2. The method according to solution A1, wherein the syntax elements related to the MVD include at least one of the following: a coded motion vector predictor index, a coded precision of the motion vector predictor, a motion vector difference, and a coded precision of the motion vector difference of the IBC AMVP mode.
[0660] A3. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being less than or equal to K, the bitstream represents excluding signaling of syntax elements related to MDV for the IBC AMVP mode, and infers the IBC AMVP mode as disabled, and wherein, K is an integer.
[0661] A4. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being less than or equal to K, the bitstream represents excluding signaling of the motion vector difference, and infers the IBC AMVP mode as disabled, and wherein, K is an integer.
[0662] A5. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being greater than K, the bitstream represents selectively including at least one of the following: the encoded / decoded motion vector predictor index, the encoded / decoded precision of the motion vector predictor, and the encoded / decoded precision of the motion vector difference of the IBC AMVP mode, and wherein, K is an integer.
[0663] A6. The method as described in any one of Schemes A3 to A5, wherein, K = 0 or K = 1.
[0664] A7. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being equal to 1, the bitstream represents excluding the motion vector predictor index of the IBC AMVP mode.
[0665] A8. The method as described in Scheme A19, wherein, the motion vector predictor index of the IBC AMVP mode is inferred as zero.
[0666] A9. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being equal to zero, the bitstream represents excluding the precision of the motion vector predictor and / or the precision of the motion vector difference of the IBC AMVP mode.
[0667] A10. The method as described in Scheme A1, wherein, based on the maximum number of IBC candidates of the first type being greater than zero, the bitstream represents excluding the precision of the motion vector predictor and / or the precision of the motion vector difference of the IBC AMVP mode.
[0668] A11. The method as described in any one of Schemes A1 to A10, wherein, the maximum number of IBC candidates of the first type is the maximum number of IBC motion candidates (denoted as maxIBCCandNum).
[0669] A12. The method according to any one of Aspects A1 to A10, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC Merge candidates (denoted as maxIBCMrgNum).
[0670] A13. The method according to any one of Aspects A1 to A10, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC Advanced Motion Vector Prediction (AMVP) candidates (denoted as maxIBCAMVPNum).
[0671] A14. The method according to any one of Aspects A1 to A10, wherein the maximum number of IBC candidates of the first type is signaled in the bitstream representation.
[0672] A15. A video processing method, comprising: determining an indication to disable the use of the Intra Block Copy (IBC) mode for a video region of a video and enable the use of the IBC mode at the sequence level of the video for a conversion between the video region of the video and the bitstream representation of the video; and performing the conversion based on the determination, wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0673] A16. The method according to Aspect A15, wherein the video region corresponds to a picture, slice, tile, tile group, or chunk of the video picture.
[0674] A17. The method according to Aspect A15 or A16, wherein the bitstream representation includes an indication associated with the determination.
[0675] A18. The method according to Aspect A17, wherein the indication is signaled in a picture, slice, tile, chunk, or Adaptive Parameter Set (APS).
[0676] A19. The method according to Aspect A18, wherein the indication is signaled in a Picture Parameter Set (PPS), slice header, picture header, tile header, tile group header, or chunk header.
[0677] A20. The method according to Aspect A15, wherein the IBC mode is disabled based on the video content of the video being different from the screen content.
[0678] A21. The method according to Aspect A15, wherein the IBC mode is disabled based on the video content of the video being content captured by a camera.
[0679] A22. A video processing method, comprising: converting between a video region of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes an indication of the use of an intra block copy (IBC) mode and / or one or more IBC-related syntax elements based on a maximum number of first-type IBC candidates used during the conversion of the video region, and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0680] A23. The method according to scheme A22, wherein the maximum number of the first-type IBC candidates is set to be equal to the maximum number of Merge candidates of inter-coded blocks.
[0681] A24. The method according to scheme A22, wherein based on the maximum number of the first-type IBC candidates being equal to zero, the bitstream representation excludes signaling of the IBC skip mode, and the IBC skip mode is inferred to be disabled.
[0682] A25. The method according to scheme A22, wherein based on the maximum number of the first-type IBC candidates being equal to zero, the bitstream representation excludes signaling of the IBC Merge mode or the IBC advanced motion vector prediction (AMVP) mode, and the IBC mode is inferred to be disabled.
[0683] A26. The method according to scheme A22, wherein based on the maximum number of the first-type IBC candidates being equal to zero, the bitstream representation excludes signaling of the Merge mode, and the IBC Merge mode is inferred to be disabled.
[0684] A27. The method according to any one of schemes A22 to A26, wherein the maximum number of the first-type IBC candidates is the maximum number of IBC motion candidates (denoted as maxIBCCandNum), the maximum number of IBC Merge candidates (denoted as maxIBCMrgNum), or the maximum number of IBC advanced motion vector prediction (AMVP) candidates (denoted as maxIBCAMVPNum).
[0685] A28. The method according to any one of schemes A22 to A26, wherein the maximum number of the first-type IBC candidates is signaled in the bitstream representation.
[0686] A29. The method according to any one of schemes A1 to A28, wherein the conversion includes: generating pixel values of the video region from the bitstream representation.
[0687] A30. The method according to any one of A1 to A28, wherein the conversion includes: generating the bitstream representation from the pixel values of the video region.
[0688] A31. A video processing device, including a processor configured to implement the method according to any one or more of A1 to A30.
[0689] A32. A computer-readable medium storing program code, which when executed, causes a processor to implement the method according to any one or more of A1 to A30.
[0690] In some embodiments, the following technical solutions may be implemented:
[0691] B1. A video processing method, including: performing conversion between a video region of a video and the bitstream representation of the video, wherein an indication of a maximum number of intra-block copy (IBC) candidates of a first type used during the conversion of the video region is signaled in the bitstream representation independently of a maximum number of Merge candidates of an inter prediction mode used during the conversion, and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0692] B2. The method according to B1, wherein the maximum number of the first type of IBC candidates is directly signaled in the bitstream representation.
[0693] B3. The method according to B1, wherein since the IBC mode is enabled for the conversion of the video region, the maximum number of the first type of IBC candidates is greater than zero.
[0694] B4. The method according to any one of B1 to B3, further including: predictively encoding and decoding the maximum number of the first type of IBC candidates (denoted as maxNumIBC) using another value.
[0695] B5. The method according to B4, wherein the indication of the maximum number of the first type of IBC candidates is set to a difference after the predictive encoding and decoding.
[0696] B6. The method according to B5, wherein the difference is between the another value and the maximum number of the first type of IBC candidates.
[0697] B7. The method according to B6, wherein the another value is a size (S) of a regular Merge list, wherein (S - maxNumIBC) is signaled in the bitstream, and wherein S is an integer.
[0698] B8. The method as described in embodiment B6, wherein the other value is a fixed integer (K), and (K - maxNumIBC) is signaled in the bitstream.
[0699] B9. The method as described in embodiment B8, wherein K = 5 or K = 6.
[0700] B10. The method as described in any one of embodiments B6 to B9, wherein the maximum number of the first type of IBC candidates is derived as (K - the indication of the maximum number signaled in the bitstream).
[0701] B11. The method as described in embodiment B5, wherein the difference is between the maximum number of the first type of IBC candidates and the other value.
[0702] B12. The method as described in embodiment B11, wherein the other value is a fixed integer (K), and (maxNumIBC - K) is signaled in the bitstream.
[0703] B13. The method as described in embodiment B12, wherein K = 0 or K = 2.
[0704] B14. The method as described in any one of embodiments B11 to B13, wherein the maximum number of the first type of IBC candidates is derived as (the indication of the maximum number signaled in the bitstream + K).
[0705] B15. The method as described in any one of embodiments B1 to B14, wherein the maximum number of the first type of IBC candidates is the maximum number of IBC motion candidates (denoted as maxIBCCandNum).
[0706] B16. The method as described in any one of embodiments B1 to B14, wherein the maximum number of the first type of IBC candidates is the maximum number of IBC Merge candidates (denoted as maxIBCMrgNum).
[0707] B17. The method as described in any one of embodiments B1 to B14, wherein the maximum number of the first type of IBC candidates is the maximum number of IBC advanced motion vector prediction (AMVP) candidates (denoted as maxIBCAMVPNum).
[0708] B18. A video processing method, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates (denoted as maxIBCCandNum) used during the conversion of the video region is a function of a maximum number of IBC Merge candidates (denoted as maxIBCMrgNum) and a maximum number of IBC advanced motion vector prediction (AMVP) candidates (denoted as maxIBCAMVPNum), and wherein when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region.
[0709] B19. The method according to scheme B18, wherein maxIBCAMVPNum is equal to 2.
[0710] B20. The method according to scheme B18 or B19, wherein the function returns its larger object.
[0711] B21. A video processing method, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates (denoted as maxIBCCandNum) used during the conversion of the video region is based on mode information of encoding and decoding of the video region.
[0712] B22. The method according to scheme B21, wherein based on the video region being encoded and decoded in the IBC Merge mode, maxIBCCandNum is set to the maximum number of IBC Merge candidates (denoted as maxIBCMrgNum).
[0713] B23. The method according to scheme B21, wherein based on the video region being encoded and decoded in the IBC AMVP mode, maxIBCCandNum is set to the maximum number of IBC advanced motion vector prediction (AMVP) candidates (denoted as maxIBCAMVPNum).
[0714] B24. A video processing method, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a decoded IBC advanced motion vector prediction (AMVP) Merge index or a decoded IBC Merge index is less than a maximum number of IBC motion candidates (denoted as maxIBCCandNum).
[0715] B25. The method according to scheme B24, wherein the decoded IBC AMVP Merge index is less than the maximum number of IBC AMVP candidates (denoted as maxIBCAMVPNum).
[0716] B26. The method as described in embodiment B24, wherein the decoded IBC Merge index is less than the maximum number of IBC Merge candidates (denoted as maxIBCMrgNum).
[0717] B27. The method as described in embodiment B26, wherein maxIBCMrgNum = 2.
[0718] B28. A video processing method, comprising: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list; and based on the determination, using a default prediction block during the conversion.
[0719] B29. The method as described in embodiment B28, wherein each sample of the default prediction block is set to (1 << (bit depth - 1)), where the bit depth is a positive integer.
[0720] B30. The method as described in embodiment B28, wherein a default block vector is assigned to the default prediction block.
[0721] B31. A video processing method, comprising: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to identify a block vector candidate in a block vector candidate list; and based on the determination, performing the conversion by treating the video region as having invalid block vectors.
[0722] B32. A video processing method, comprising: during the conversion between the video region of a video and the bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or an IBC Merge candidate index fails to meet a condition; based on the determination, generating a supplementary block vector (BV) candidate list; and using the supplementary BV candidate list for the conversion.
[0723] B33. The method as described in embodiment B32, wherein the condition includes that the IBC AMVP candidate index or the IBC Merge candidate index is not less than the maximum number of IBC motion candidates for the video region.
[0724] Method according to embodiment B32 or B33, wherein said supplementary BV candidate vector list is generated using the following steps: adding one or more history-based motion vector prediction (HMVP) candidates, generating one or more virtual BV candidates from other BV candidates; and adding one or more default candidates.
[0725] B35. Method according to embodiment B34, wherein said steps are carried out in sequence.
[0726] B36. Method according to embodiment B34, wherein said steps are carried out in an interleaved manner.
[0727] B37. A video processing method, comprising: converting between a video region of a video and a bitstream representation of said video, wherein the maximum number (denoted as maxIBCAMVPNum) of intra block copy (IBC) advanced motion vector prediction (AMVP) candidates is not equal to two.
[0728] B38. Method according to embodiment B37, wherein said bitstream representation excludes a flag indicating a motion vector predictor index and includes an index having a value greater than one.
[0729] B39. Method according to embodiment B38, wherein said index is binary encoded or decoded using unary, truncated unary, fixed length or exponential Golomb representation.
[0730] B40. Method according to embodiment B38, wherein the binary number of the binary digit string of said index is context encoded or bypass encoded.
[0731] B41. Method according to embodiment B37, wherein maxIBCAMVPNum is greater than the maximum number (denoted as maxIBCMrgNum) of IBC Merge candidates.
[0732] B42. A method of video processing, comprising, during conversion between a video region and a bitstream representation of said video region, determining that the maximum number of intra block copy (IBC) motion candidates is zero, and performing said conversion by generating an IBC block vector candidate list during said conversion.
[0733] B43. Method according to embodiment B42, wherein said conversion further includes generating a Merge list based on the fact that the conversion for said video region is enabled in intra block copy (IBC) advanced motion vector prediction (AMVP) mode.
[0734] B44. The method according to scheme B42, wherein the conversion further includes generating a Merge list having a length up to the maximum number of IBC AMVP candidates based on the fact that the conversion for the video region is disabled in the IBC Merge mode.
[0735] B45. The method according to any one of schemes B1 to B44, wherein the conversion includes: generating pixel values of the video region from the bitstream representation.
[0736] B46. The method according to any one of schemes B1 to B44, wherein the conversion includes: generating the bitstream representation from the pixel values of the video region.
[0737] B47. A video processing device, including a processor configured to implement the method according to any one or more of schemes B1 to B46.
[0738] B48. A computer-readable medium storing program code, which when executed, causes the processor to implement the method according to any one or more of schemes B1 to B46.
[0739] In some embodiments, the following technical solutions can be implemented:
[0740] C1. A video processing method, including: during the conversion between a video region of a video and the bitstream representation of the video region, determining to disable the intra block copy mode for the video region and enable the intra block copy mode for other video regions of the video; and performing the conversion based on the determination; wherein, in the intra block copy mode, pixels of the video region are predicted from other pixels in a video picture corresponding to the video region.
[0741] C2. The method according to scheme C1, wherein the video region corresponds to a video picture or a slice or a tile or a group of slices or a chunk of a video picture.
[0742] C3. The method according to any one of schemes C1 - C2, wherein the bitstream representation includes an indication of the determination.
[0743] C4. The method according to scheme C3, wherein the indication is included at the picture or slice or tile or group of slices or chunk level.
[0744] C5. A method for video processing, comprising: converting between a video region of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes an indication regarding determining that a maximum number (denoted as maxIBCCandNum) of IBC motion candidates used during the conversion of the video region is equal to a maximum number (denoted as MaxNumMergeCand) of Merge candidates used during the conversion.
[0745] C6. The method according to scenario C5, wherein maxIBCCandNum is equal to zero, and since maxIBCCCardNum is equal to zero, the bitstream representation excludes signaling of certain IBC information.
[0746] C7. A method for video processing, comprising: converting between a video region of a video and a bitstream representation of the video, wherein the bitstream representation depends on a value of a maximum number (denoted as maxIBCCandNum) of intra block copy motion candidates used during the conversion of the video region to satisfy conditions including syntax elements related to a motion vector difference of an intra block copy alternative motion vector predictor; wherein, in the intra block copy mode, pixels of the video region are predicted from other pixels in a video picture corresponding to the video region.
[0747] C8. The method according to scenario C7, wherein the condition includes that maxIBCCandNum is equal to a maximum number (denoted as MaxNumMergeCand) of Merge candidates used during the conversion.
[0748] C9. The method according to scenario C8, wherein maxIBCCandNum is equal to 0, and since maxIBCCandNum is equal to 0, signaling of a motion vector predictor index of an intra block copy alternative motion vector predictor is omitted from the bitstream representation.
[0749] C10. The method according to scenario C8, wherein maxIBCCandNum is equal to 0, and since maxIBCCandNum is equal to 0, signaling of a motion vector difference of an intra block copy alternative motion vector predictor is omitted from the bitstream representation.
[0750] C11. A method for video processing, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy motion candidates (denoted as maxIBCCandNum) used during the conversion of the video region is signaled in the bitstream representation independently of a maximum number of Merge candidates (denoted as MaxNumMergeCand) used in the conversion; wherein, intra block copy corresponds to a mode in which pixels of the video region are predicted from other pixels in a video picture corresponding to the video region.
[0751] C12. The method according to scheme C11, wherein, since intra block copy is enabled for the conversion of the video region, MaxIBCCandNum is greater than zero.
[0752] C13. The method according to any one of schemes C11 - C12, wherein maxIBCCandNum is signaled in the bitstream representation by using another value for predictive coding.
[0753] C14. The method according to scheme C13, wherein the another value is MaxNumMergeCand.
[0754] C15. The method according to scheme C13, wherein the another value is a constant K.
[0755] C16. A method for video processing, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates (denoted as maxIBCCandNum) used during the conversion of the video region depends on codec mode information of the video region; wherein, IBC corresponds to a mode in which pixels of the video region are predicted from other pixels in a video picture corresponding to the video region.
[0756] C17. The method according to scheme C16, when a block is coded in IBC Merge mode, setting maxIBCCandNum to a maximum number of IBC Merge candidates, denoted as maxIBCMrgNum.
[0757] C18. The method according to scheme C16, when a block is coded in IBC alternative motion vector prediction candidate (AMVP) mode, setting maxIBCCandNum to a maximum number of IBC AMVP numbers, denoted as maxIBCAMVPNum.
[0758] C19. The method according to scheme C17, wherein maxIBCMrgNum is signaled in a slice header.
[0759] C20. The method as described in Scheme C17, where maxIBCMrgNum is set to be the same as the maximum number of allowed non-IBC translational Merge candidates.
[0760] C21. The method as described in Scheme C18, where maxIBCAMVPNum is set to 2.
[0761] C22. A method for video processing, comprising: converting between a video region of a video and a bitstream representation of the video, wherein a maximum number of intra block copy (IBC) motion candidates (denoted as maxIBCCandNum) used during the conversion of the video region is a first function of a maximum number of IBC Merge candidates (denoted as maxIBCMrgNum) and a maximum number of IBC alternative motion vector prediction candidates (denoted as maxIBCAMVPNum); wherein IBC corresponds to a mode of predicting pixels of the video region from other pixels in a video picture corresponding to the video region.
[0762] C23. The method as described in Scheme C22, wherein maxIBCAMVPNum is equal to 2.
[0763] C24. The method as described in Schemes C22 - C23, wherein during the conversion, a decoded IBC alternative motion vector predictor index is less than maxIBCAMVPNum.
[0764] C25. A method for video processing, comprising: during a conversion between a video region of a video and a bitstream representation of the video, determining that an intra block copy alternative motion vector predictor index or an intra block copy Merge candidate index fails to identify a block vector candidate in a block vector candidate list; and based on the determination, using a default prediction block during the conversion.
[0765] C26. A method for video processing, comprising: during a conversion between a video region of a video and a bitstream representation of the video, determining that an intra block copy alternative motion vector predictor index or an intra block copy Merge candidate index fails to identify a block vector candidate in a block vector candidate list; and based on the determination, performing the conversion by treating the video region as having an invalid block vector.
[0766] C27. A method for video processing, comprising: during a conversion between a video region of a video and a bitstream representation of the video, determining that an intra block copy alternative motion vector predictor index or an intra block copy Merge candidate index fails to meet a condition; and based on the determination, generating a supplementary block vector (BV) candidate list; and using the supplementary block vector candidate list for the conversion.
[0767] C28. The method as described in embodiment C27, wherein the condition includes that the intra block copy alternative motion vector predictor index or the intra block copy Merge candidate index is less than the maximum number of intra block copy candidates in the video region.
[0768] C29. The method as described in any one of embodiments C27 - C28, wherein the supplementary BV candidate vector list is generated using the following steps: generating history - based motion vector predictor candidates, generating virtual BV candidates from other BV candidates; and adding default candidates.
[0769] C30. The method as described in embodiment C29, wherein the steps are performed in sequence.
[0770] C31. The method as described in embodiment C29, wherein the steps are performed in an interleaved manner.
[0771] C32. A method for video processing, comprising: converting between a video region of a video and a bit - stream representation of the video, wherein the maximum number of IBC alternative motion vector prediction candidates (denoted as maxIBCAMVPNum) is not equal to 2; wherein IBC corresponds to a mode of predicting pixels in the video region from other pixels in a video picture corresponding to the video region.
[0772] C33. The method as described in embodiment C32, wherein the bit - stream representation excludes a first flag indicating a motion vector predictor index and includes an index having a value greater than one.
[0773] C34. The method as described in any one of embodiments C32 - C33, wherein the index is encoded and decoded using binarization represented by unary / truncated unary / fixed - length / exponential Golomb / other binarization.
[0774] C35. A method for video processing, comprising: during the conversion between a video region and a bit - stream representation of the video region, determining that the maximum number of intra block copy (IBC) candidates is zero, and performing the conversion by generating an IBC block vector candidate list during the conversion; wherein IBC corresponds to a mode of predicting pixels in the video region from other pixels in a video picture corresponding to the video region.
[0775] C36. The method as described in embodiment C35, wherein the conversion further comprises: generating a Merge list by assuming that the IBC alternative motion vector predictor mode is enabled for the conversion of the video region.
[0776] C37. The method according to scheme C34, wherein the conversion further includes: when the IBC Merge mode is disabled for the conversion of the video region, generating a Merge list, the length of which reaches the maximum number of candidates for the IBC advanced motion vector predictor.
[0777] C38. The method according to any one of schemes C1 - C37, wherein the conversion includes: generating pixel values of the video region from the bitstream representation.
[0778] C39. The method according to any one of schemes C1 - C37, wherein the conversion includes: generating the bitstream representation from the pixel values of the video region.
[0779] C40. A video processing device, including a processor configured to implement the method according to any one or more of schemes C1 to C39.
[0780] C41. A computer - readable medium storing program code, which when executed causes the processor to implement the method according to any one or more of schemes C1 to C39.
[0781] The disclosures and other schemes, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or in a combination of one or more of them. The disclosures and other embodiments can be implemented as one or more computer program products, that is, one or more computer program instruction modules encoded on a computer - readable medium for execution by, or for controlling the operation of, a data - processing device. The computer - readable medium can be a machine - readable storage device, a machine - readable storage substrate, a memory device, a substance composition affecting a machine - readable propagated signal, or a combination of one or more of them. The term "data - processing device" encompasses all devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the device may also include code for creating an execution environment for the computer programs under discussion, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or a combination of one or more of them. A propagated signal is an artificially generated signal, such as a machine - generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to a suitable receiver device.
[0782] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a portion of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program being discussed, or in multiple coordinated files (such as files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0783] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0784] By way of example, processors suitable for the execution of a computer program include any one or more of general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include or be operatively coupled to one or more mass storage devices (such as magnetic disks, magneto-optical disks, or optical disks) for storing data, to receive data from or transfer data to the one or more mass storage devices, or both to receive and transfer data. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry.
[0785] Although this patent document contains many details, these details should not be construed as limiting any subject matter or the scope of what can be claimed, but rather as descriptions of features specific to particular embodiments of a particular technology. In this patent document, certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the above features may be described as acting in certain combinations and even initially claimed as such, in some cases, one or more features from a claimed combination may be excluded from that combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.
[0786] Similarly, although operations are depicted in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the operations shown be performed, to achieve a desired result. Moreover, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[0787] Only a few implementations and examples have been described, and other implementations, enhancements, and variations can be made based on what is described and shown in this patent document.
Claims
1. A video processing method, comprising: Converting between a video region of a video and a bitstream of the video, wherein an indication of a maximum number of intra block copy (IBC) candidates of a first type used during the conversion of the video region is signaled in the bitstream independently of a maximum number of Merge candidates of an inter mode used during the conversion. Wherein, when the IBC mode is applied, samples of the video region are predicted from other samples in a video picture corresponding to the video region. Wherein the indication of the maximum number of the first type of IBC candidates is set to a difference after subtracting a fixed integer K1 from the maximum number of the first type of IBC candidates, and wherein K1 = 0 or K1 = 2. Wherein the maximum number of the first type of IBC candidates includes a maximum number of IBC motion candidates, and when the maximum number of IBC motion candidates is equal to 1, a motion vector predictor index of an IBC advanced motion vector prediction (AMVP) mode is not signaled, and a value of the motion vector predictor index is inferred to be zero.
2. The method according to claim 1, wherein The maximum number of the first type of IBC candidates is directly signaled in the bitstream.
3. The method according to claim 1, wherein, Since the IBC mode is enabled for the conversion of the video region, the maximum number of the first type of IBC candidates is greater than zero.
4. The method according to claim 1, further comprising: Predictively encoding the maximum number of the first type of IBC candidates using another value, wherein the maximum number of the first type of IBC candidates is represented as maxNumIBC.
5. The method according to claim 4, wherein, The indication of the maximum number of the first type of IBC candidates is set to a difference after the predictive encoding.
6. The method according to claim 5, wherein, The difference is between the another value and the maximum number of the first type of IBC candidates.
7. The method according to claim 6, wherein The another value is a size S of a regular Merge list, wherein (S - maxNumIBC) is signaled in the bitstream, and wherein S is an integer.
8. The method according to claim 6, wherein The another value is a fixed integer K2, and wherein, (K2 - maxNumIBC) is signaled in the bitstream.
9. The method according to claim 8, wherein, K2=5。 10. The method according to claim 9, wherein, The maximum number of the first type of IBC candidates is derived as (K2 - the indication of the maximum number signaled in the bitstream).
11. The method according to claim 1, wherein, The maximum number of the first type of IBC candidates is derived as (the indication of the maximum number signaled in the bitstream + K1).
12. The method according to any one of claims 1 to 11, wherein, The maximum number of the IBC motion candidates is represented as maxIBCCandNum.
13. The method according to any one of claims 1 to 11, wherein The maximum number of the first type of IBC candidates includes a maximum number of IBC Merge candidates, and the maximum number of the IBC Merge candidates is represented as maxIBCMrgNum.
14. The method according to any one of claims 1 to 11, wherein, The maximum number of the first type of IBC candidates includes a maximum number of IBC advanced motion vector prediction (AMVP) candidates, and the maximum number of the IBC AMVP candidates is represented as maxIBCAMVPNum.
15. The video processing method according to claim 1, wherein, The maximum number of intra block copy (IBC) motion candidates used during the conversion of the video region is a function of the maximum number of IBC Merge candidates and the maximum number of IBC AMVP candidates. The maximum number of IBC motion candidates is denoted as maxIBCCandNum, the maximum number of IBC Merge candidates is denoted as maxIBCMrgNum, and the maximum number of IBC AMVP candidates is denoted as maxIBCAMVPNum.
16. The method according to claim 15, wherein, maxIBCAMVPNum is equal to 2.
17. The method according to claim 15 or 16, wherein, The function returns the larger of the two.
18. The video processing method according to claim 1, wherein, The maximum number of IBC motion candidates used during the conversion of the video region is set based on the encoded mode information of the video region. The maximum number of IBC motion candidates is denoted as maxIBCCandNum.
19. The method according to claim 18, wherein, Based on the video region being encoded in IBC Merge mode, maxIBCCandNum is set to the maximum number of IBC Merge candidates, which is denoted as maxIBCMrgNum.
20. The method according to claim 18, wherein Based on the video region being encoded in IBC AMVP mode, maxIBCCandNum is set to the maximum number of IBC AMVP candidates, which is denoted as maxIBCAMVPNum.
21. The video processing method according to claim 1, wherein, The decoded IBC AMVP Merge index or the decoded IBC Merge index is less than the maximum number of IBC motion candidates, which is denoted as maxIBCCandNum.
22. The method according to claim 21, wherein, The decoded IBC AMVP Merge index is less than the maximum number of IBC AMVP candidates, which is denoted as maxIBCAMVPNum.
23. The method according to claim 21, wherein, The decoded IBC Merge index is less than the maximum number of IBC Merge candidates, which is denoted as maxIBCMrgNum.
24. The method according to claim 23, wherein, maxIBCMrgNum = 2.
25. The video processing method according to claim 1, further comprising: During the conversion between the video region of the video and the bitstream of the video, determining that the IBC alternative motion vector predictor (AMVP) candidate index or the IBC Merge candidate index cannot identify a block vector candidate in the block vector candidate list; and Based on the determination, using a default prediction block during the conversion.
26. The method according to claim 25, wherein, Each sample point of the default prediction block is set to (1 << (BitDepth - 1)), where BitDepth is a positive integer.
27. The method according to claim 26, wherein, A default block vector is assigned to the default prediction block.
28. The video processing method according to claim 1, further comprising: During conversion between the video region of the video and the bitstream of the video, it is determined that the Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or the IBC Merge candidate index cannot identify a block vector candidate in the block vector candidate list; and Based on the determination, perform the conversion by treating the video region as having invalid block vectors.
29. The video processing method according to claim 1, further comprising: During conversion between the video region of the video and the bitstream of the video, determine that the Intra Block Copy (IBC) Alternate Motion Vector Predictor (AMVP) candidate index or the IBC Merge candidate index does not meet the condition; Based on the determination, generate a supplementary block vector (BV) candidate list; and Use the supplementary BV candidate list to perform the conversion.
30. The method according to claim 29, wherein, The condition includes that the IBC AMVP candidate index or the IBC Merge candidate index is not less than the maximum number of IBC motion candidates for the video region.
31. The method according to claim 29 or 30, wherein The supplementary BV candidate vector list is generated using the following steps: Add one or more History-based Motion Vector Predictor (HMVP) candidates, Generate one or more virtual BV candidates from other BV candidates; and Add one or more default candidates.
32. The method according to claim 31, wherein, The steps are performed in sequence.
33. The method according to claim 31, wherein, The steps are performed in an interleaved manner.
34. The video processing method according to claim 1, wherein, The maximum number of Intra Block Copy (IBC) AMVP candidates is not equal to two, and the maximum number of Intra Block Copy (IBC) AMVP candidates is represented as maxIBCAMVPNum.
35. The method according to claim 34, wherein, The bitstream excludes a flag indicating a motion vector predictor index and includes an index having a value greater than one.
36. The method according to claim 35, wherein, The index is binary-coded using unary, truncated unary, fixed length, or exponential Golomb representation.
37. The method according to claim 35, wherein, The binary number of the binary digit string of the binary of the index is context-coded or bypass-coded.
38. The method according to claim 34, wherein, maxIBCAMVPNum is greater than the maximum number of IBC Merge candidates, and the IBC The maximum number of Merge candidates is represented as maxIBCMrgNum.
39. The method of video processing according to claim 1, further comprising: During conversion between the video region and the bitstream of the video region, determine that the maximum number of Intra Block Copy (IBC) motion candidates is zero, and Based on the determination, perform the conversion by generating an IBC block vector candidate list during the conversion.
40. The method according to claim 39, wherein, The conversion further includes generating a Merge list based on the IBC AMVP mode being enabled for the conversion of the video region.
41. The method according to claim 39, wherein, The conversion further includes generating a Merge list having a length up to the maximum number of IBC AMVP candidates based on the IBC Merge mode being disabled for the conversion of the video region.
42. The method according to any one of claims 1-11, 15, 16, 18-30, 32-41, wherein, The conversion includes generating the pixel values of the video region from the bitstream.
43. The method according to any one of claims 1-11, 15, 16, 18-30, 32-41, wherein, The conversion includes generating the bitstream from the pixel values of the video region.
44. A video processing device, comprising: A processor configured to implement the method according to any one of claims 1 to 43.
45. A computer-readable medium storing program code which, when executed, causes a processor to implement the method according to any one of claims 1 to 43.