Coding of block vectors in intra-block copy coded blocks

IBC and AMVP techniques enhance video coding by optimizing motion vector prediction and encoding, addressing inefficiencies in existing standards to reduce bandwidth and improve compression efficiency.

JP7729858B2Active Publication Date: 2025-08-26DOUYIN VISION CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023130143
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-05-25
Filing Date
2023-08-09
Publication Date
2025-08-26
Estimated Expiration
2040-05-25

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in efficiently encoding and decoding video data, particularly in managing motion vectors and prediction modes, leading to suboptimal compression efficiency and increased bandwidth requirements.

Method used

The implementation of Intra Block Copy (IBC) mode and Advanced Motion Vector Prediction (AMVP) techniques for video processing, which involve converting between a video domain and a bitstream representation, optimizing motion vector prediction and encoding through enhanced candidate selection and signaling methods.

Benefits of technology

Improves the quality of decompressed video by reducing bandwidth demands and enhancing compression efficiency, aligning with the goals of future video coding standards like Versatile Video Coding (VVC).

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007729858000019
    Figure 0007729858000019
  • Figure 0007729858000020
    Figure 0007729858000020
  • Figure 0007729858000021
    Figure 0007729858000021
Patent Text Reader

Abstract

To provide a method, a device, and a system for decoding or encoding a video on the basis of intra-block coding while using a block vector signal notification and / or a merge candidate.SOLUTION: A video processing method includes performing conversion between a video area of the video and a bitstream representation of the video. The bitstream representation selectively includes an MVD (Motion Vector Difference) related syntax element for an IBC (Intra Block Copy) AMVP (Advanced Motion Vector Prediction) mode on the basis of the maximum number of IBC candidates of a first type to be used during the conversion of the video area. When the IBC mode is applied, a sample in the video area is predicted from other samples in the video picture that corresponds to the video area.SELECTED DRAWING: Figure 20A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS Book This application claims priority to and the benefit of International Patent Application No. PCT / CN2019 / 088454, filed May 25, 2019. main stretch This is a divisional application of Japanese Patent Application No. 2021-568847 based on International Patent Application No. PCT / CN2020 / 092032 filed on May 25, 2020. The entire disclosure of the above application is three By lighting Here It is cited.

[0002] This specification relates to video and image encoding and decoding techniques. [Background technology]

[0003] Digital video accounts for the largest bandwidth usage on the Internet and other digital communications networks, and the bandwidth demands for digital video use are expected to continue to grow as the number of connected user devices capable of receiving and displaying video increases. Summary of the Invention

[0004] The disclosed techniques may be used by embodiments of a video or image decoder or encoder that performs intra-block coding-based decoding or encoding of video using block vector signaling and / or merging candidates.

[0005] In one exemplary aspect, a method for video processing is disclosed. The method includes performing a conversion between a video domain of a video and a bitstream representation of the video, the bitstream representation selectively including MVD (Motion Vector Difference)-related syntax elements for an IBC Advanced Motion Vector Prediction (AMVP) mode based on a maximum number of first-type IBC (Intra Block Copy) candidates used during the video domain conversion, wherein, when the IBC mode is applied, samples of the video domain are predicted from other samples in a video picture corresponding to the video domain.

[0006] In another representative aspect, a method for video processing is disclosed that includes determining that an indication of use of an Intra Block Copy (IBC) mode is disabled for the video domain and that use of the IBC mode is enabled at a sequence level of the video for conversion between a video domain of the video and a bitstream representation of the video, and performing the conversion based on the determination, wherein, when the IBC mode is applied, samples of the video domain are predicted from other samples in a video picture that correspond to the video domain.

[0007] In yet another representative aspect, a method for video processing is disclosed that includes performing a conversion between a video domain of a video and a bitstream representation of the video, the bitstream representation selectively including an indication regarding use of an Intra Block Copy (IBC) mode and / or one or more IBC-related syntax elements based on a maximum number of first-type IBC candidates used during the video domain conversion, wherein, when the IBC mode is applied, samples of the video domain are predicted from other samples in a video picture that correspond to the video domain.

[0008] In yet another representative aspect, a method for video processing is disclosed, the method including performing a conversion between a video domain of a video and a bitstream representation of the video, wherein an indication of a maximum number of first-type Intra Block Copy (IBC) candidates used during the video domain conversion is signaled in the bitstream representation independently of a maximum number of inter-mode merge candidates used during the conversion, and wherein, when the IBC mode is applied, samples of the video domain are predicted from other samples in a video picture corresponding to the video domain.

[0009] In yet another representative aspect, a method for video processing is disclosed, the method including performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of Intra Block Copy (IBC) motion candidates (denoted by maxIBCCandNum) used during the video domain conversion is a function of a maximum number of IBC merge candidates (denoted by maxIBCMrgNum) and a maximum number of IBC Advanced Motion Vector Prediction (AMVP) candidates (denoted by maxIBCAMVPNum), and wherein, when an IBC mode is applied, samples of the video domain are predicted from other samples in a video picture that correspond to the video domain.

[0010] In yet another representative aspect, a method for video processing is disclosed, the method including performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of Intra Block Copy (IBC) motion candidates (denoted maxIBCCandNum) used during the video domain conversion is based on coded mode information of the video domain.

[0011] In yet another representative aspect, a method for video processing is disclosed, the method including performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a decoded Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) merge index or a decoded IBC merge index is less than a maximum number of IBC motion candidates (denoted by maxIBCCandNum).

[0012] In yet another representative aspect, a method for video processing is disclosed that includes determining, during conversion between a video domain of a video and a bitstream representation of the video, that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index fails to identify a block vector candidate in a block vector candidate list, and using a default prediction block during the conversion based on the determination.

[0013] In yet another representative aspect, a method for video processing is disclosed that includes, during conversion between a video region of a video and a bitstream representation of the video, determining that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index fails to identify a block vector candidate in a block vector candidate list, and, based on the determination, performing the conversion by treating the video region as having an invalid block vector.

[0014] In yet another representative aspect, a method for video processing is disclosed that includes determining that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index does not satisfy a condition during conversion between a video domain of a video and a bitstream representation of the video, generating a supplemental Block Vector (BV) candidate list based on the determination, and performing the conversion using the supplemental BV candidate list.

[0015] In yet another representative aspect, a method for video processing is disclosed, the method including performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) candidates (denoted maxIBCAMVPNum) is not equal to two.

[0016] In another exemplary aspect, the above-described methods may be implemented by a video decoder including a processor.

[0017] In another exemplary aspect, the above-described methods may be implemented by a video encoder including a processor.

[0018] In yet another exemplary aspect, the methods may be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.

[0019] These and other aspects are further described herein. [Brief explanation of the drawings]

[0020] [Figure 1] 10 illustrates an example of a derivation process for building a merge candidate list. [Figure 2] 10 shows examples of spatial merge candidate locations. [Figure 3]10 shows examples of candidate pairs that are considered for redundancy check of spatial merge candidates. [Figure 4A] 10 shows examples of the location of the second Prediction Unit (PU) for N×2N and 2N×N partitions. [Figure 4B] 10 shows examples of the location of the second Prediction Unit (PU) for N×2N and 2N×N partitions. [Figure 5] FIG. 10 is an illustration of motion vector scaling for temporal merge candidates. [Figure 6] 1 shows example candidate positions for temporal merge candidates, C0 and C1. [Figure 7] 10 illustrates an example of a combined bi-predictive merge candidate. [Figure 8] The process of deriving motion vector prediction candidates will be summarized below. [Figure 9] 10 illustrates a description of motion vector scaling for spatial motion vector candidates. [Figure 10A] A four-parameter affine motion model and a six-parameter affine motion model are shown. [Figure 10B] A four-parameter affine motion model and a six-parameter affine motion model are shown. [Figure 11] This is an example of an affine MVF (Motion Vector Field) for each sub-block. [Figure 12] 10 shows examples of candidate positions for affine merge mode. [Figure 13] 10 shows an example of a modified merge list construction process. [Figure 14] 1 illustrates an example of triangulation-based inter prediction. [Figure 15] An example of UMVE (Ultimate Motion Vector Expression) search processing will be shown. [Figure 16] An example of UMVE search points is shown below. [Figure 17] An example of MVD(0,1) mirrored between list 0 and list 1 in DMVR is shown. [Figure 18]Here is an example of an MV that may be checked in one iteration: [Figure 19] An example of IBC (Intra Block Copy) is shown below. [Figure 20A] 1 is a flowchart illustrating an example of a method for video processing. [Figure 20B] 1 is a flowchart illustrating an example of a method for video processing. [Figure 20C] 1 is a flow chart illustrating an example of a method for video processing. [Figure 20D] 1 is a flowchart illustrating an example of a method for video processing. [Figure 20E] 1 is a flow chart illustrating an example of a method for video processing. [Figure 20F] 1 is a flowchart illustrating an example of a method for video processing. [Figure 20G] 1 is a flow chart illustrating an example of a method for video processing. [Figure 20H] 1 is a flowchart illustrating an example of a method for video processing. [Figure 20I] 1 is a flow chart illustrating an example of a method for video processing. [Figure 20J] 1 is a flow chart illustrating an example of a method for video processing. [Figure 20K] 1 is a flowchart illustrating an example of a method for video processing. [Figure 21] FIG. 1 is a block diagram illustrating an example of a video processing device. [Figure 22] FIG. 1 is a block diagram illustrating an exemplary video processing system in which the disclosed techniques can be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0021] This specification provides various techniques that can be used by a decoder of an image or video bitstream to improve the quality of the decompressed or decoded digital video or image. For simplicity, the term "video" is used herein to include both a series of pictures (conventionally called a video) and individual images. Furthermore, a video encoder may implement these techniques during the encoding process to reconstruct decoded frames for use in further encoding.

[0022] Section headings are used herein for ease of understanding and are not intended to limit embodiments disclosed in one section to only that section, and thus embodiments in one section may be combined with embodiments in other sections.

[0023] 1. Summary of the invention

[0024] The present invention relates to video coding technology. Specifically, it relates to motion vector coding. It may be applied to existing video coding standards such as HEVC, or may be applied to establish a standard (Versatile Video Coding). The present invention is also applicable to future video coding standards or video codecs.

[0025] 2 Background technology

[0026] Video coding standards have evolved primarily through the development of well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, while ISO / IEC created MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 AVC (Advanced Video Coding), and H.265 / HEVC. Since H.262, video coding standards have been based on hybrid video coding architectures that utilize temporal prediction and transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, many new methods have been adopted by JVET and incorporated into reference software called Joint Exploration Mode (JEM). In April 2018, the Joint Video Expert Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG), and is working on formulating the VVC standard with the goal of reducing the bitrate by 50% compared to HEVC.

[0027] The latest version of the VVC draft, namely Versatile Video Coding (Draft 5), can be found at:

[0028] phenix.it-sudparis.eu / jvet / doc_end_user / documents / 14_Geneva / wg11 / JVET-N1001-v2.zip

[0029] The latest reference software for VVC, called VTM, can be found at:

[0030] vcgit.hhi.fraunhofer.de / jvet / VVCSoftware_VTM / tags / VTM-5.0

[0031] 2.1 Inter Prediction in HEVC / H.265

[0032] For an inter-coded Coding Unit (CU), it may be coded with one Prediction Unit (PU) or two PUs depending on the partition mode. Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists may be signaled using inter_pred_idc. The motion vector may be explicitly coded as a differential relative to the predictor.

[0033] When a CU is coded in skip mode, one PU is associated with the CU, there are no significant residual coefficients, and there are no coded motion vector differentials or reference picture indices. A merge mode is specified, whereby motion parameters for the current PU are obtained from neighboring PUs, including spatial and temporal candidates. The merge mode can be applied to any inter-predicted PU, not just for skip mode. As an alternative to the merge mode, there is explicit transmission of motion parameters, whereby the motion vectors (more precisely, the Motion Vector Difference (MVD) compared to the motion vector predictor), the corresponding reference picture index of each reference picture list, and the usage of the reference picture list are explicitly signaled to each PU. This disclosure refers to such a mode as Advanced Motion Vector Prediction (AMVP).

[0034] If the signaling indicates using one of two reference picture lists, the PU is generated from one block of samples. This is called "uni-prediction." Uni-prediction is available for both P slices and B slices.

[0035] If the signaling indicates that both reference picture lists are to be used, the PU is generated from two blocks of samples. This is called "bi-prediction." Bi-prediction is available only for B slices.

[0036] The inter prediction modes defined in HEVC will be described in detail below, with merge mode being first described.

[0037] 2.1.1 Reference Picture Buffer

[0038] In HEVC, the term inter-prediction is used to indicate a prediction derived from data elements (e.g., sample values ​​or motion vectors) of reference pictures other than the currently decoded picture. Similar to H.264 / AVC, an image can be predicted from multiple reference pictures. The reference pictures used for inter-prediction are organized into one or more reference picture lists. A reference index identifies which reference picture in the list to use to generate the prediction signal.

[0039] One reference picture list, List0, is used for P slices, and two reference picture lists, List0 and List1, are used for B slices. Note that the reference pictures included in List0 / 1 may be from past and future pictures in terms of shooting / display order.

[0040] 2.1.2 Merge Mode 2.1.2.1 Deriving Merge Mode Candidates

[0041] When predicting a PU using merge mode, an index pointing to an entry in the merge candidate list is parsed from the bitstream and used to look up the motion information. The construction of this list is specified in the HEVC standard and can be summarized based on the following sequence of steps: ●Step 1: Derive initial candidates Step 1.1: Derive spatial candidates Step 1.2: Check spatial candidates for redundancy Step 1.3: Derive temporal candidates ● Step 2: Insert additional candidates Step 2.1: Creating bidirectional prediction candidates Step 2.2: Inserting zero motion candidates

[0042] These steps are also shown schematically in Figure 1. For spatial merge candidate derivation, up to four merge candidates are selected from candidates at five different locations. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since the decoder assumes a fixed number of candidates per PU, additional candidates are generated if the number of candidates obtained in step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the slice header. Since the number of candidates is fixed, truncated unary binarization (TU) is used to encode the index of the best merge candidate. If the CU size is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list for the 2N × 2N prediction units.

[0043] The operations associated with the above steps are now described in detail.

[0044] FIG. 1 illustrates an example of a derivation process for building a merge candidate list.

[0045] 2.1.2.2 Deriving Spatial Candidates

[0046] In deriving spatial merge candidates, up to four merge candidates are selected from the candidates located at the positions shown in FIG. 2. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any of the PUs at positions A1, B1, B0, and A0 are unavailable (e.g., because they belong to another slice or tile) or are intra-coded. After adding the candidate at position A1, the remaining candidates are subjected to a redundancy check, which ensures that candidates with identical motion information are removed from the list, improving coding efficiency. To reduce computational complexity, the aforementioned redundancy check does not consider all possible candidate pairs. Instead, only pairs linked by arrows in FIG. 3 are considered, and a candidate is added to the list only if the corresponding candidate used for the redundancy check does not have the same motion information. Another source of overlapping motion information is a "second PU" associated with a partition different from 2N×2N. As an example, FIG. 4 shows the second PU for N×2N and 2N×N cases, respectively. When dividing the current PU into Nx2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would lead to two prediction units with the same motion information, which is redundant since there is only one PU in the coding unit. Similarly, when dividing the current PU into 2NxN, position B1 is not considered.

[0047] FIG. 2 shows examples of spatial merge candidate locations.

[0048] FIG. 3 shows an example of a pair of candidates that are considered for the redundancy check of spatial merge candidates.

[0049] FIG. 4 shows examples of the location of the second PU for N×2N and 2N×N partitions.

[0050] 2.1.2.3 Temporal Candidate Derivation

[0051] In this step, only one candidate is added to the list. Specifically, in deriving this temporal merge candidate, a scaled motion vector is derived based on the co-located PU belonging to the picture with the smallest POC difference with the current picture in a given reference picture list. The reference picture list used to derive the co-located PU is explicitly signaled in the slice header. As shown by the dotted line in Figure 5, the scaled motion vector of the temporal merge candidate is obtained, which is scaled from the motion vector of the co-located PU using POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located PU and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. The practical implementation of the scaling process is described in the HEVC specification. For a B slice, two motion vectors are obtained, one for reference picture list 0 and one for reference picture list 1, and combined to form a bi-predictive merge candidate.

[0052] FIG. 5 is an illustration of scaling of motion vectors for temporal merge candidates.

[0053] For the co-located PU(Y) belonging to the reference frame, the location of the temporal candidate is selected between candidate C0 and candidate C1, as shown in Figure 6. If the PU at location C0 is unavailable, intra-coded, or outside the current Coding Tree Unit (CTU) (also known as Largest Coding Unit (LCU)) row, location C1 is used. Otherwise, location C0 is used to derive the temporal merge candidate.

[0054] FIG. 6 shows example candidate locations for temporal merge candidates, C0 and C1.

[0055] 2.1.2.4 Inserting additional candidates

[0056] In addition to spatial-temporal merge candidates, there are two additional types of merge candidates: joint bidirectional prediction merge candidates and zero merge candidates. The spatial-temporal merge candidates are used to generate joint bidirectional prediction merge candidates. Joint bidirectional prediction merge candidates are used only for B slices. A joint bidirectional prediction candidate is generated by combining the first reference picture list motion parameters of a first candidate with the second reference picture list motion parameters of another candidate. If these two tuples provide different motion hypotheses, they form a new bidirectional prediction candidate. As an example, Figure 7 shows the case where two candidates with mvL0 and refIdxL0 or mvL1 and refIdxL1 in the original list (left) are used to generate a joint bidirectional prediction merge candidate that is added to the final list (right). There are various rules for the combinations considered to generate these additional merge candidates.

[0057] FIG. 7 shows an example of a joint bi-predictive merge candidate.

[0058] Zero motion candidates are inserted to fill the remaining entries in the merge candidate list, thereby hitting the MaxNumMergeCand capacity. These candidates have a spatial displacement of zero and a reference picture index that starts at zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.

[0059] 2.1.3 AMVP

[0060] AMVP exploits the spatial-temporal correlation between motion vectors and neighboring PUs and uses it to explicitly transmit motion parameters. For each reference picture list, it first checks the availability of neighboring PU positions on the left and above, removes redundant candidates, and adds zero vectors to keep the length of the candidate list constant, thereby constructing a motion vector candidate list. The encoder can then select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to merge index signaling, the index of the best motion vector candidate is coded using a truncated unary encoding. In this case, the maximum value to be coded is 2 (see Figure 8). The following sections provide details on the process of deriving motion vector prediction candidates.

[0061] 2.1.3.1 Derivation of AMVP Candidates

[0062] FIG. 8 summarizes the process of deriving motion vector prediction candidates.

[0063] FIG. 8 shows an example of a process for deriving motion vector prediction candidates.

[0064] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. To derive spatial motion vector candidates, as shown in Figure 2, two motion vector candidates are ultimately derived based on the motion vectors of each PU at five different positions.

[0065] To derive a temporal motion vector candidate, select one motion vector candidate from two candidates derived based on two different co-located positions. After creating an initial list of spatial-temporal candidates, remove duplicate motion vector candidates in the list. If the number of candidates is greater than two, remove motion vector candidates whose reference picture index in the associated reference picture list is greater than 1 from the list. If the number of spatial-temporal motion vector candidates is less than two, add an additional zero motion vector candidate to the list.

[0066] 2.1.3.2 Spatial Motion Vector Candidates

[0067] In deriving spatial motion vector candidates, up to two candidates that are in the same position as the motion merge are considered among five possible candidates derived from PUs positioned as shown in Figure 2. The derivation order for the left side of the current PU is specified as A0, A1, scaled A0, scaled A1. The derivation order for the top side of the current PU is specified as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each edge, there are four possible cases for motion vector candidates: two cases that do not require the use of spatial scaling and two cases that do use spatial scaling. The four different cases can be summarized as follows: No spatial scaling -(1) Same reference picture list and same reference picture index (same POC) -(2) Different reference picture lists and the same reference picture (same POC) Spatial scaling (3) Same reference picture list but different reference pictures (different POC) (4) Different reference picture lists and different reference pictures (different POCs)

[0068] First check the non-spatial scaling case, then perform spatial scaling. Regardless of the reference picture list, if the POC differs between the reference picture of the neighboring PU and the reference picture of the current PU, spatial scaling is considered. If all PUs of the left candidate are not available or are intra-coded, scaling of the upper motion vector helps to derive the left and upper MV candidates in parallel. Otherwise, spatial scaling is not allowed for the upper motion vector.

[0069] FIG. 9 shows a description of motion vector scaling for spatial motion vector candidates.

[0070] In the spatial scaling process, the motion vectors of neighboring PUs are scaled in the same manner as in temporal scaling, as shown in Figure 9. The main difference is that the reference picture list and index of the current PU are given as input, and the actual scaling process is the same as in temporal scaling.

[0071] 2.1.3.3 Temporal Motion Vector Candidates

[0072] Except for deriving the reference picture index, all the processes for deriving temporal merge candidates are the same as the processes for deriving spatial motion vector candidates (see FIG. 6). The reference picture index is signaled to the decoder.

[0073] 2.2 Inter-Prediction Method in VVC

[0074] New coding tools to improve inter prediction include Adaptive Motion Vector difference Resolution (AMVR) for signaling MVD, Merge with Motion Vector Differences (MMVD), Triangular Prediction Mode (TPM), Combined Intra-Inter Prediction (CIIP), Advanced TMVP (ATMVP) (also known as SbTMVP), Affine Prediction Mode, Generalized BI-prediction (GBI), Decoder-side Motion Vector Refinement (DMVR), and BI-Optical flow (BIO) (also known as BDOF).

[0075] There are three different merge list building processes supported by VVC: 1) Sub-block merge candidate list: Contains ATMVP and affine merge candidates. One merge list construction process is shared for both affine and ATMVP modes. Note that ATMVP and affine merge candidates may be added sequentially. The size of the sub-block merge list is signaled in the slice header, with a maximum value of 5. 2) Regular merge list: For the remaining coding blocks, one merge list construction process is shared. Here, spatial / temporal / HMVP, pairwise combined bidirectional predictive merge candidates, and zero motion candidates may be inserted in order. The size of the regular merge list is signaled in the slice header, and its maximum value is 6. MMVD, TPM, and CIIP depend on the regular merge list. 3) IBC Merge List: Performed in the same way as a regular merge list.

[0076] Similarly, the AMVP lists supported by VVC are: 1) Affine AMVP candidate list 2) Regular AMVP candidate list 3) IBC AMVP candidate list: Due to the adoption of JVET-N0843, the construction process is the same as the IBC merge list.

[0077] 2.2.1 Coding Block Structure in VVC

[0078] In VVC, a quadtree / binarytree / ternarytree (QT / BT / TT) structure is adopted to divide the picture into square or rectangular blocks.

[0079] Besides QT / BT / TT, a separate tree (also known as a dual coding tree) is also adopted in VVC for I-frames. Using separate trees, the coding block structure is signaled separately for luma and chroma components.

[0080] In addition, CU is set to PU and TU except for blocks coded using two specific coding methods (e.g., intra subpartition prediction, where PU is equal to TU but smaller than CU, and subblock transformation of inter-coded blocks, where PU is equal to CU but TU is smaller than PU).

[0081] 2.2.2 Affine Prediction Mode

[0082] In HEVC, only a translational motion model is applied for MCP (Motion Compensation Prediction). Meanwhile, in the real world, there are various types of motion, such as zoom-in / zoom-out, rotation, perspective motion, and other irregular motions. In VVC, a 4-parameter affine model and a 6-parameter affine model are used to apply simple affine transformation motion compensation prediction. As shown in Figures 10A and 10B, the affine motion region of a block is represented by two CPMVs (Control Point Motion Vectors) in the case of the 4-parameter affine model, and by three CPMVs in the case of the 6-parameter affine model.

[0083] Figures 10A and 10B show: 10A: Simple affine motion model - parameter affine, 10B: 6-parameter affine mode.

[0084] The MVF (Motion Vector Field) of a block is expressed by the following equations using the four-parameter affine model in equation (1) (where the four parameters are defined as variables a, b, e, and f) and the six-parameter affine model in equation (2) (where the four parameters are defined as variables a, b, c, d, e, and f), respectively.

[0085]

number

[0086]

number

[0087] where (mv h 0,mv h 0) is the motion vector of the control point in the upper left corner, and (mv h 1,mv h 1) is the motion vector of the control point in the upper right corner, and (mv h 2,MV h 2) is the motion vector of the control point in the lower left corner, and all three motion vectors are called CPMV (Control Point Motion Vector), (x,y) represents the coordinates of the representative point relative to the upper left sample in the current block, and (mv h (x,y),mv v (x,y) is the motion vector derived for the sample located at (x,y). The CP motion vector may be signaled (as in Affine AMVP mode) or derived on the fly (as in Affine Merge mode). w and h are the width and height of the current block. In practice, this division is performed by a right shift with rounding operations. In VTM, the representative point is defined as the center position of a sub-block. For example, if the coordinates of the upper left corner of a sub-block relative to the upper left sample in the current block are (xs,ys), the coordinates of the representative point are defined as (xs+2,ys+2). For each sub-block (i.e., 4x4 in VTM), the representative point is used to derive the motion vector for the entire sub-block.

[0088] To further simplify motion compensation prediction, subblock-based affine transformation prediction is applied. To derive the motion vector of each M×N subblock (in current VVC, both M and N are set to 4), the motion vector of the center sample of each subblock is calculated according to Equation (1) and Equation (2), as shown in Figure 11, and rounded to 1 / 16 decimal precision. Then, a 1 / 16 pixel motion compensation interpolation filter is applied, and the derived motion vector is used to generate a prediction for each subblock. The 1 / 16 pixel interpolation filter is implemented in affine mode.

[0089] FIG. 11 shows an example of affine MVF for each sub-block.

[0090] After MCP, the high-precision motion vectors of each sub-block are rounded and stored with the same precision as the normal motion vectors.

[0091] 2.2.3 Merging Whole Blocks 2.2.3.1 Merge list construction in translational normal merge mode 2.2.3.1.1 HMVP(History-based Motion Vector Prediction)

[0092] Unlike the merge list design, VVC employs the History-based Motion Vector Prediction (HMVP) method.

[0093] HMVP stores previously coded motion information. The motion information of previously coded blocks is defined as HMVP candidates. Multiple HMVP candidates are stored in a table called the HMVP table, which is maintained on-the-fly during the encoding / decoding process. When starting encoding / decoding of a new tile / LCU row / slice, the HMVP table is emptied. Whenever there are inter-coded blocks and non-sub-blocks in non-TPM mode, the associated motion information is added as a new HMVP candidate to the last entry of the table. The overall coding flow is shown in Figure 12.

[0094] 2.2.3.1.2 Normal Merge List Construction Process

[0095] The construction of a typical merge list (for translation) can be summarized as the following sequence of steps: ● Step 1: Derive spatial candidates ●Step 2: Inserting HMVP candidates Step 3: Insert pairwise average candidates ● Step 4: Default movement candidates

[0096] HMVP candidates can be used for both AMVP and merge candidate list construction processes. Figure 13 illustrates the modified merge candidate list construction process (highlighted in blue). If the merge candidate list is not full after inserting the TMVP candidates, the HMVP candidates stored in the HMVP table can be used to fill the merge candidate list. Considering that a block usually has a high correlation with its nearest neighbors in terms of motion information, the HMVP candidates in the table are inserted in descending order of index. The last entry in the table is added to the list first, followed by the first entry. Similarly, redundancy elimination is applied to the HMVP candidates. The merge candidate list construction process terminates when the total number of available merge candidates reaches the signaled maximum number of allowed merge candidates.

[0097] Note that all spatial / temporal / HMVP candidates are coded in non-IBC mode, otherwise they are not allowed to be added to the regular merge candidate list.

[0098] The HMVP table contains up to five canonical motion candidates, each unique.

[0099] 2.2.3.2 TPM (Triangular Prediction Mode)

[0100] In VTM4, triangulation mode is supported for inter prediction. Triangulation mode applies only to CUs that are 8x8 or larger, are coded in merge mode, and are not coded in MMVD or CIIP mode. For CUs that meet these conditions, a CU-level flag is signaled to indicate whether triangulation mode applies.

[0101] When using this mode, a CU is divided into two equal triangular partitions using either a diagonal or anti-diagonal partition, as shown in Figure 13. Each triangular partition in a CU is inter-predicted using its own motion, and only uni-prediction is allowed for each partition. That is, each partition has one motion vector and one reference index. Similar to traditional bi-directional prediction, a uni-prediction motion constraint is applied to require only two motion-compensated predictions for each CU.

[0102] FIG. 14 shows an example of triangulation-based inter prediction.

[0103] If the CU-level flag indicates that the current CU is coded in triangular partition mode, a flag indicating the triangular partition direction (diagonal or anti-diagonal) and two merge indices (one for each partition) are further signaled. After predicting each triangular partition, sample values ​​along the diagonal or anti-diagonal edges are adjusted using a blending process with adaptive weights. This is the prediction signal for the entire CU, and transformation and quantization processes are applied to the entire CU, as with other prediction modes. Finally, the motion field of the CU predicted using the triangular partition mode is stored in 4x4 units.

[0104] The normal merge candidate list is reused for triangulation merge prediction without excessive pruning of motion vectors. For each merge candidate in the normal merge candidate list, only one of its L0 or L1 motion vectors is used for triangulation prediction. The order of selecting L0 vs. L1 motion vectors is based on their merge index parity. This scheme allows the normal merge list to be used directly.

[0105] 2.2.3.3 MMVD

[0106] JVET-L0054 proposes the Ultimate Motion Vector Expression (UMVE), also known as MMVD, which is a proposed motion vector representation method used in either skip or merge modes.

[0107] UMVE reuses merge candidates similar to those included in the regular merge candidate list in VVC. Among the merge candidates, base candidates can be selected, which is further extended by the proposed motion vector representation method.

[0108] UMVE provides a new MVD (Motion Vector Difference) representation method, which uses a starting point, a motion magnitude, and a motion direction to represent one MVD.

[0109] FIG. 15 shows an example of the UMVE search process.

[0110] FIG. 16 shows an example of UMVE search points.

[0111] The proposed technique uses the merge candidate list as is, but for the extension of UMVE, only candidates with the default merge type (MRG_TYPE_DEFAULT_N) are considered.

[0112] The base candidate index defines the starting point, which indicates the best candidate among the candidates in the list as follows:

[0113] [Table 1]

[0114] If the number of base candidates is equal to 1, the base candidate IDX is not signaled.

[0115] The distance index is information on the magnitude of the movement. The distance index indicates a predefined distance from the start point information. The predefined distances are as follows:

[0116] [Table 2]

[0117] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown below.

[0118] [Table 3]

[0119] The UMVE flag is signaled immediately after sending the skip or merge flag. If the skip or merge flag is true, the UMVE flag is parsed. If the UMVE flag is equal to 1, the UMVE syntax is parsed. But if it is not 1, the AFFINE flag is parsed. If the AFFINE flag is equal to 1, i.e., in AFFINE mode, but not 1, the skip / merge index is parsed for the VTM's skip / merge mode.

[0120] No additional line buffers are required due to UMVE candidates, because the software skip / merge candidates are used as base candidates. The input UMVE index is used to determine the MV completion just before motion compensation, so there is no need to maintain a long line buffer.

[0121] Under the current common test conditions, either the first or second merge candidate in the merge candidate list may be selected as the base candidate.

[0122] UMVE is also known as MMVD (Merge with MV Differences).

[0123] 2.2.3.4 CIIP(Combined Intra-Inter Prediction)

[0124] In JVET-L0100, multiple multiple hypotheses are proposed, and combined intra and inter prediction is one way to generate multiple hypotheses.

[0125] When applying multi-hypothesis prediction to fine-tune intra modes, multi-hypothesis prediction combines one intra prediction and one merge index prediction. In a merge CU, if a flag is true, one flag is signaled for merge mode, and an intra mode is selected from the intra candidate list. For the luma component, the intra candidate list is derived from only one intra prediction mode, i.e., planar mode. The weights applied to a predicted block from intra prediction and inter prediction are determined by the coded modes (intra or non-intra) of the two neighboring blocks (A1 and B1).

[0126] 2.2.4 Merging for Sub-Block Based Techniques

[0127] Note that in addition to the regular merge list of non-subblock merge candidates, it is proposed that all subblock-related motion candidates be put into a separate merge list.

[0128] Sub-block related motion candidates are put into a separate merge list called the "sub-block merge candidate list."

[0129] In one example, the sub-block merge candidate list includes ATMVP candidates and affine merge candidates.

[0130] The sub-block merge candidate list is filled with candidates in the following order: a. ATMVP candidates (possibly available or unavailable) b. Affine merge list (including inheritance affine candidates and construction affine candidates) c. Padding as a zero MV 4-parameter affine model

[0131] 2.2.4.1.1 ATMVP (also known as Sub-Block Temporal Motion Vector Predictor, SbTMVP)

[0132] The basic idea of ​​ATMVP is to derive multiple temporal motion vector predictors for a block. Each sub-block is assigned a set of motion information. When generating ATMVP merge candidates, motion compensation is performed at the 8x8 level instead of the whole block level.

[0133] 2.2.5 Normal Inter Mode (AMVP) 2.2.5.1 AMVP Motion Candidate List

[0134] Similar to the AMVP design in HEVC, up to two AMVP candidates may be derived. However, HMVP candidates may be added after the TMVP candidates. HMVP candidates in the HMVP table are traversed in ascending order of index (i.e., starting from the oldest index equal to 0). Up to four HMVP candidates may be checked to see if their reference picture is the same as the target reference picture (i.e., has the same POC value).

[0135] 2.2.5.2 AMVR

[0136] In HEVC, when use_integer_mv_flag is equal to 0 in the slice header, MVD (Motion Vector Difference) (the difference between the motion vector of the PU and the predicted motion vector) is signaled in units of 1 / 4 luma sample. In VVC, local AMVR (Adaptive Motion Vector Resolution) is introduced. In VVC, MVD can be coded in units of 1 / 4 luma sample, integer luma sample, or 4 luma sample (i.e., 1 / 4 pixel, 1 pixel, 4 pixel). MVD resolution is controlled at the CU (Coding Unit) level, and an MVD resolution flag is conditionally signaled for each CU that has at least one non-zero MVD component.

[0137] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MV precision is used in the CU. If the first flag (equal to 1) indicates that quarter luma sample MV precision is not used, another flag is signaled to indicate whether integer luma sample MV precision or 4 luma sample MV precision is used.

[0138] If the first MVD resolution flag of a CU is zero or not coded for the CU (i.e., all MVDs in the CU are zero), 1 / 4 luma sample MV resolution is used for the CU. If the CU uses integer luma sample MV precision or 4 luma sample MV precision, round the MVPs in the CU's AMVP candidate list to the corresponding precision.

[0139] 2.2.5.3 Symmetric Motion Vector Differentials in JVET-N1001-v2

[0140] In JVET-N1001-v2, SMVD (Symmetric Motion Vector Difference) is applied to motion information coding in bidirectional prediction.

[0141] First, at the slice level, the variables RefIdxSyml0 and RefIdxSyml1, which indicate the reference picture indexes of lists 0 / 1 used in SMVD mode, respectively, are derived in the following steps as specified in N1001-v2: If at least one of the two variables is equal to -1, SMVD mode shall be disabled.

[0142] 2.2.6 Improvement of motion information 2.2.6.1 DMVR (Decoder-side Motion Vector Refinement)

[0143] In bidirectional prediction operation, to predict one block area, two prediction blocks formed using the MV (Motion Vector) of list0 and the MV of list1 respectively are combined to form one prediction signal. In the Decoder-side Motion Vector Refinement (DMVR) method, the two motion vectors of bidirectional prediction are further refined.

[0144] For DMVR in VVC, as shown in FIG. 17, we assume mirroring between list 0 and list 1, and perform bilateral matching to refine the MV, i.e., find the best MVD among several MVD candidates. The MVs of two reference picture lists are denoted by MVL0(L0X, L0Y) and MVL1(L1X, L1Y). The MVD denoted by (MvdX, MvdY) for list 0 that can minimize a cost function (e.g., SAD) is defined as the best MVD. The SAD function is defined as the SAD between the reference block of list 0 derived by the motion vector (L0X+MvdX, L0Y+MvdY) in the reference picture of list 0 and the reference block of list 1 derived by the motion vector (L1X-MvdX, L1Y-MvdY) in the reference picture of list 1.

[0145] The motion vector refinement process may be iterated twice. In each iteration, up to six MVDs (integer pixel accuracy) may be checked in two steps, as shown in FIG. 18. In the first step, MVDs (0,0), (-1,0), (1,0), (0,-1), and (0,1) are checked. In the second step, one of MVDs (-1,-1), (-1,1), (1,-1), or (1,-1) may be selected and checked further. The function Sad(x,y) returns the SAD value of MVD(x,y). The MVD, represented by (MvdX,MvdY) checked in the second step, is determined as follows: MvdX=-1; MvdY=-1; If(Sad(1,0) <Sad(-1,0)) MvdX=1; If(Sad(0,1) <Sad(0,-1)) MvdY=1;

[0146] In the first iteration, the starting point is the signaled MV, and in the second iteration, the starting point is the signaled MV plus the best MVD selected in the first iteration. DMVR is applied only if one reference picture is an earlier picture and the other reference picture is a later picture, and the two reference pictures have the same picture order count distance from the current picture.

[0147] FIG. 17 shows an example of an MVD(0,1) mirrored between list 0 and list 1 in a DMVR.

[0148] FIG. 18 shows an example of MVs that may be checked in one iteration.

[0149] To further simplify the DMVR process, JVET-M0147 proposed several changes to the JEM design. Specifically, the DMVR design adopted for VTM-4.0 (soon to be released) has the following key features: ●If the SAD at the (0,0) position between list0 and list1 is less than a threshold, the process will terminate early. ●If the SAD between list0 and list1 is 0 at a certain position, terminate early. ● Block size for DMVR: W*H>=64&H>=8, where W and H are the width and height of the block. For DMVRs where the size of the CU is greater than 16*16, the CU is divided into multiple 16x16 sub-blocks. If only the width or height of the CU is greater than 16, the CU is divided only vertically or horizontally. ●The size of the reference block is (W+7)*(H+7) (for luminance). 25-point SAD-based integer pixel search (i.e., (+-) two refinement search ranges, single stage) ●Bilinear interpolation based DMVR. ● "Parametric Error Surface Equation" based sub-pixel refinement. This procedure is performed only if the minimum SAD cost is not equal to zero and the best MVD is (0,0) in the previous MV refinement iteration. ● Reference block padding for luma / chroma MC (if necessary). ● Improved MV used only for MC and TMVP.

[0150] 2.2.6.1.1 How to Use DMVR

[0151] The DMVR may be enabled if all of the following conditions are true: - The DMVR enabled flag in the SPS (i.e., sp_dmvr_enabled_flag) is equal to 1. - The TPM flag, the inter-affine flag, the sub-block merge flag (ATMVP or affine merge), and the MMVD flag are all equal to 0. -Merge flag equals 1. The current block is bi-directionally predicted and the POC distance between the current picture and the reference picture in list 1 is equal to the POC distance between the reference picture in list 0 and the current picture. - The height of the current CU is 8 or more. -The number of luminance samples (CU width * height) is 64 or more.

[0152] 2.2.6.1.2 "Parametric Error Surface Equation" Based Sub-Pixel Improvement

[0153] The method is summarized below. 1. Compute the parametric error surface fit only if the center location is the best cost location at a given iteration. 2. Fit the following 2-D parabolic error surface equation using the center location cost and the costs at locations (-1,0), (0,-1), (1,0), and (0,1) from the center: E(x,y)=A(x-x0) 2 +B(y-y0) 2 +C where (x0,y0) corresponds to the position with the minimum cost and C corresponds to the minimum cost value. By solving five equations in five unknowns, (x0,y0) is calculated as follows: x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) y0=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (x0,y0) can be calculated with any required sub-pixel precision by adjusting the precision with which the division is performed (i.e., how many bits of the quotient are calculated). For 1 / 16-pixel precision, the absolute value of the quotient only needs to be calculated with 4 bits, which is suitable for a fast shift-and-subtract based implementation requiring two divisions per CU. 3. The calculated (x0, y0) is added to the integer distance refinement MV to get the sub-pixel accurate refinement difference MV.

[0154] 2.2.6.2 BDOF (Bi-Directional Optical Flow) 2.3 Intra-block copy

[0155] Intra Block Copy (IBC), also known as current picture reference, is adopted in HEVC-Content Coding extensions (HEVC-SCC) and the current VVC Test Model (VTM-4.0). IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. As shown in Figure 18, when IBC is applied, the current block is predicted by a reference block within the same picture. Before encoding or decoding the current block, samples in the reference block must already be reconstructed. While IBC is not very efficient for most camera-captured sequences, it exhibits significant coding gains for screen content. This is because screen content pictures often contain repetitive patterns, such as icons and characters. IBC can effectively remove redundancies between these repetitive patterns. In HEVC-SCC, an inter-coded coding unit (CU) can apply IBC if it selects the current picture as its reference picture. In this case, MV is renamed to Block Vector (BV), and BV always has integer pixel precision. To conform to the Main Profile HEVC, the current picture is marked as a "long-term" reference picture in the Decoded Picture Buffer (DPB), just as inter-view reference pictures are also marked as "long-term" reference pictures in multiple view / 3D video coding standards.

[0156] If we find the reference block following the BV, we can generate a prediction by copying the reference block. The residual is obtained by subtracting the reference pixels from the original signal. Then, we can apply transforms and quantization, just like in other coding modes.

[0157] FIG. 19 shows an example of intra block copying.

[0158] However, if the reference block is outside the picture, or overlaps with the current block, or is outside the reconstructed region, or is outside the valid region limited by some constraint, some or all of the pixel values ​​may be undefined. Essentially, there are two solutions to deal with such problems. One is to not allow such situations, e.g., bitstream conformance. The other is to apply padding to these undefined pixel values. The following subsections explain the solutions in detail.

[0159] 2.3.1 IBC in the VVC Test Model (VTM4.0)

[0160] In the current VVC test model, i.e., VTM-4.0 design, the entire reference block should have the current Coding Tree Unit (CTU) and not overlap with the current block. Therefore, there is no need to pad the reference or predicted block. The IBC flag is coded as the prediction mode of the current CU. Thus, there are a total of three prediction modes for each CU: MODE_INTRA, MODE_INTER, and MODE_IBC.

[0161] 2.3.1.1 IBC merge mode

[0162] In IBC merge mode, indices that point to entries in the IBC merge candidate list are parsed from the bitstream. The construction of the IBC merge list can be summarized according to the following sequence of steps: ● Step 1: Derive spatial candidates ●Step 2: Inserting HMVP candidates Step 3: Insert pairwise average candidates

[0163] In the derivation of spatial merge candidates, up to four merge candidates are selected from the candidates at positions A1, B1, B0, A0, and B2 shown in FIG. 2. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, and A0 is unavailable (e.g., because it belongs to another slice or tile) or is not coded in IBC mode. After adding the candidate at position A1, the insertion of the remaining candidates is subject to a redundancy check to ensure that candidates with the same motion information can be removed from the list, so that coding efficiency is improved.

[0164] After inserting the spatial candidate, if the IBC merge list size is still less than the maximum IBC merge list size, an IBC candidate from the HMVP table may be inserted. Upon inserting the HMVP candidate, a redundancy check is performed.

[0165] Finally, the pairwise average candidates are inserted into the IBC merge list.

[0166] If the reference block identified by the merge candidate is outside the picture, or overlaps with the current block, or is outside the reconstructed region, or is outside the valid region limited by some constraint, the merge candidate is called an invalid merge candidate.

[0167] Note that invalid merge candidates may be inserted into the IBC merge list.

[0168] 2.3.1.2 IBC AMVP Mode

[0169] In IBC AMVP mode, an AMVP index that points to an entry in an IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized according to the following sequence of steps: ● Step 1: Derive spatial candidates ○Check A0, A1 until a usable candidate is found. ○Check B0, B1, B2 until a usable candidate is found. ●Step 2: Inserting HMVP candidates ● Step 3: Inserting zero candidates

[0170] If, after inserting the spatial candidates, the IBC AMVP list size is still less than the maximum IBC AMVP list size, then an IBC candidate from the HMVP table may be inserted.

[0171] Finally, the zero candidate is inserted into the IBC AMVP list.

[0172] 2.3.1.3 Saturation IBC Mode

[0173] In the current VVC, motion compensation in chroma IBC mode is performed at the subblock level. A chroma block is divided into multiple subblocks. Each subblock determines whether the corresponding luma block has a block vector, and if so, whether it is valid. The current VTM has an encoder constraint that tests chroma IBC mode to determine whether all subblocks in the current chroma CU have valid luma block vectors. For example, in a YUV420 video, the chroma block is NxM, and the co-located luma region is 2Nx2M. The subblock size of a chroma block is 2x2. There are several steps to derive the chroma mv and then perform the block copy process. 1) The chroma block is first divided into (N>>1)*(M>>1) sub-blocks. 2) For each sub-block whose top-left sample is located at (x,y), fetch the corresponding luma block containing the same top-left sample located at (2x,2y). 3) The encoder checks the block vector (bv) of the fetched luma block. If it meets one of the following conditions, bv is considered invalid: a. There is no corresponding luminance block bv. The predicted block identified by b.bv has not yet been reconstructed. The predicted block identified in c.bv partially or completely overlaps with the current block. 4) The chroma motion vector of a sub-block is set to the motion vector of the corresponding luma sub-block.

[0174] If all sub-blocks find a valid bv, IBC mode is enabled in the encoder.

[0175] The decoding process for IBC blocks is shown below. The parts related to the derivation of chroma motion vectors in IBC mode are enclosed in double bold brackets, i.e., {{a}} indicates that "a" is related to the derivation of chroma motion vectors in IBC mode.

[0176] 8.6.1 General Decoding Process for Coding Units Coded with IBC Prediction The inputs to this process are: a luminance position (xCb, yCb) that defines the top-left sample of the current coding block relative to the top-left luminance sample of the current picture; a variable cbWidth that specifies the width of the current coding block in luminance samples, a variable cbHeight that specifies the height of the current coding block in luminance samples; A variable treeType that specifies whether a single or dual tree is used, and if a dual tree is used, whether the current tree corresponds to the luma or chroma component.

[0177] The output of this process is a modified reconstructed image before in-loop filtering. The quantization parameter derivation process specified in Section 8.7.1 is called with the luma position (xCb, yCb), the width of the current coding block in luma samples, cbWidth, the height of the current coding block in luma samples, cbHeight, and the variable treeType as input. The decoding process for a coding unit coded in ibc prediction mode consists of the following ordered steps:

[0178] 1. The motion vector components of the current coding unit are derived as follows: 1. If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the following applies: The motion vector component derivation process specified in Section 8.6.2.1 is called with the position (xCb, yCb) of the luminance coding block, the width cbWidth of the luminance coding block, and the height cbHeight of the luminance coding block as input, and the luminance motion vector mvL[0][0] as output. If -treeType is equal to SINGLE_TREE, the chroma motion vector derivation process of section 8.6.2.9 is called with the luma motion vector mvL[0][0] as input and the chroma motion vector mvC[0][0] as output. The number of luma coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY are both set equal to 1. 1. Otherwise, if treeType is equal to DUAL_TREE_CHROMA, the following applies: The number of LUMA coding sub-blocks in the horizontal direction NUMSBX and vertical direction NUMBY is derived as follows: numSbX=(cbWidth>>2) (8-886) numSbY=(cbHeight>>2) (8-887) The chroma motion vector mvC[xSbIdx][ySbIdx] is derived as follows when xSbIdx=0..numSbX-1, ySbIdx=0..numSbY-1. The luminance motion vector mvL[xSbIdx][ySbIdx] is derived as follows: The position (xCuY, yCuY) of the co-located luminance coding unit is derived as follows: xCuY=xCb+xSbIdx*4 (8-888) yCuY=yCb+ySbIdx*4 (8-889) - If CuPredMode[xCuY][yCuY] is equal to MODE_INTRA, then the following applies: mvL[xSbIdx][ySbIdx][0]=0 (8-890) mvL[xSbIdx][ySbIdx][1]=0 (8-891) predFlagL0[xSbIdx][ySbIdx]=0 (8-892) predFlagL1[xSbIdx][ySbIdx]=0 (8-893) - Otherwise (if CuPredMode[xCuY][yCuY] is equal to MODE_IBC), the following applies: mvL[xSbIdx][ySbIdx][0]=MvL0[xCuY][yCuY][0] (8-894) mvL[xSbIdx][ySbIdx][1]=MvL0[xCuY][yCuY][1] (8-895) predFlagL0[xSbIdx][ySbIdx]=1 (8-896) predFlagL1[xSbIdx][ySbIdx]=0 (8-897)}} The chroma motion vector derivation process in Section -8.6.2.9 is called with mvL[xSbIdx][ySbIdx] as input and mvC[xSbIdx][ySbIdx] as output. It is a bitstream conformance requirement that the chroma motion vectors mvC[xSbIdx][ySbIdx] obey the following constraints: - When the block availability derivation process [ED.(BB): Neighborhood Block Availability Check Process tbd] specified in Section 6.4.X is called with inputs the current chroma position (xCurr, yCurr) equal to (xCb / SubWidthC, yCb / SubHeightC), the neighboring chroma positions (xCb / SubWidthC + (mvC[xSbIdx][ySbIdx][0] >> 5), yCb / SubHeightC + (mvC[xSbIdx][ySbIdx][1] >> 5)), the output is equal to TRUE. - When the block availability derivation process [ED.(BB): Neighborhood Block Availability Check Process tbd] specified in Section 6.4.X is called with inputs the current chroma position (xCurr, yCurr) equal to (xCb / SubWidthC, yCb / SubHeightC) and the neighboring chroma position (xCb / SubWidthC + (mvC[xSbIdx][ySbIdx][0]>>5) + cbWidth / SubWidthC-1, yCb / SubHeightC + (mvC[xSbIdx][ySbIdx][1]>>5) + cbHeight / SubHeightC-1), the output shall be equal to TRUE. -One or both of the following conditions are true: -(mvC[xSbIdx][ySbIdx][0]>>5)+xSbIdx*2+2 is less than or equal to 0. -(mvC[xSbIdx][ySbIdx][1]>>5)+ySbIdx*2+2 is less than or equal to 0.

[0179] 2. The predicted samples for the current coding unit are derived as follows: If treeType is equal to SINGLE_TREE or DUAL_TREE_LUMA, the predicted samples for the current coding unit are derived as follows: The decoding process for the ibc block, as specified in clause 8.6.3.1, takes as input the position of the luma coding block (xCb, yCb), the width of the luma coding block cbWidth, the height of the luma coding block cbHeight, the number of luma coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, the luma motion vector mvL[xSbIdx][ySbIdx], where xSbIdx=0..numSbX-1 and SbIdx=0..numSbY-1, and the variable cIdx set equal to 0, and returns the (cbWidth) × (cbHeight) array of predicted luma samples predSamples L It is called with the ibc predicted samples (predSamples) as output. Alternatively, if treeType is equal to SINGLE_TREE or DUAL_TREE_CHROMA, the predicted samples for the current coding unit are derived as follows: The decoding process for the ibc block, as specified in subclause 8.6.3.1, takes as input the position of the luma coding block (xCb, yCb), the width cbWidth of the luma coding block, the height cbHeight of the luma coding block, the number of luma coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, the chroma motion vector mvC[xSbIdx] with xSbIdx=0..numSbX-1 and ySbIdx=0..numSbY-1, and the variable cIdx set equal to 1, and returns a (cbWidth / 2) × (cbHeight / 2) array predSamples of predicted chroma samples for the chroma component Cb. Cb It is called with the ibc predicted samples (predSamples) as output. The decoding process for the ibc block, as specified in subclause 8.6.3.1, takes as input the position of the luma coding block (xCb, yCb), the width of the luma coding block cbWidth, the height of the luma coding block cbHeight, the number of luma coding sub-blocks in the horizontal direction numSbX and the vertical direction numSbY, the chroma motion vector mvC[xSbIdx] with xSbIdx=0..numSbX-1 and ySbIdx=0..numSbY-1, and the variable cIdx set equal to 2, and returns a (cbWidth / 2) × (cbHeight / 2) array predSamples of predicted chroma samples for the chroma component Cr. Cr It is called with the ibc predicted samples (predSamples) as output.

[0180] 3. The variables NumSbX[xCb][yCb] and NumSbY[xCb][yCb] are set equal to numSbX and numSbY, respectively.

[0181] 4. The residual samples of the current coding unit are derived as follows: When treeType is equal to SINGLE_TREE or when treeType is equal to DUAL_TREE_LUMA, the decoding process of the residual signal of the coding block coded in the inter prediction mode specified in Section 8.5.8 uses the position (xTb0, yTb0) set equal to the luminance position (xCb, yCb), the width nTbW set equal to the luminance coding block width cbWidth, the height nTbH set equal to the luminance coding block height cbHeight, and the variable cldxset set equal to 0 as input, and the array resamples L is called with the output When treeType is equal to SINGLE_TREE or when treeType is equal to DUAL_TREE_CHROMA, the decoding process of the residual signal of a coding block coded in the inter prediction mode specified in Section 8.5.8 uses the position (xTb0, yTb0) set equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set equal to the chroma coding block width cbWidth / 2, the height nTbH set equal to the chroma coding block height cbHeight / 2, and the variable cldxset set equal to 1 as input, and the array resamples Cb is called with the output When treeType is equal to SINGLE_TREE or when treeType is equal to DUAL_TREE_CHROMA, the decoding process of the residual signal of a coding block coded in the inter prediction mode specified in Section 8.5.8 uses the position (xTb0, yTb0) set equal to the chroma position (xCb / 2, yCb / 2), the width nTbW set equal to the chroma coding block width cbWidth / 2, the height nTbH set equal to the chroma coding block height cbHeight / 2, and the variable cldxset set equal to 2 as input, and the array resSamples Cr is called with the output

[0182] 5. The reconstructed samples of the current coding unit are derived as follows: - If treeType is equal to SINGLE_TREE or if treeType is equal to DUAL_TREE_LUMA, the color component image reconstruction process specified in Section 8.7.5 shall be performed with the block position (xB, yB) set equal to (xCb, yCb), the block width bWidth set equal to cbWidth, the block height bHeight set equal to cbHeight, and the variables cIdx, predSamples set equal to 0. L The (cbWidth) x (cbHeight) arrays predSamples and resSamples are set equal to LIt is called with input the (cbWidth) × (cbHeight) array resSamples set equal to , and the output is the modified reconstructed image before in-loop filtering. - If treeType is equal to SINGLE_TREE or if treeType is equal to DUAL_TREE_CHROMA, the image reconstruction process for color components specified in Section 8.7.5 shall be performed with the block position (xB, yB) set equal to (xCb / 2, yCb / 2), the block width bWidth set equal to cbWidth / 2, the block height bHeight set equal to cbHeight / 2, and the variables cIdx, predSamples set equal to 1. Cb The (cbWidth / 2) x (cbHeight / 2) arrays predSamples and resSamples are set equal to Cb and the output is the modified reconstructed image before in-loop filtering. - If treeType is equal to SINGLE_TREE or if treeType is equal to DUAL_TREE_CHROMA, the image reconstruction process for color components specified in Section 8.7.5 shall be performed with the block position (xB, yB) set equal to (xCb / 2, yCb / 2), the block width bWidth set equal to cbWidth / 2, the block height bHeight set equal to cbHeight / 2, and the variables cIdx, predSamples set equal to 2. Cr The (cbWidth / 2) × (cbHeight / 2) arrays predSamples and resSamples are set equal to Cr and the output is the modified reconstructed image before in-loop filtering.

[0183] 2.3.2 Recent Trends in IBC (VTM 5.0 version) 2.3.2.1 Single BV List

[0184] JVET-N0843 is adopted for VVC. In JVET-N0843, the BV predictors for merge mode and AMVP mode in IBC share a common predictor list consisting of the following elements: ●Two spatially adjacent positions (A1 and B1 in Figure 2) ● 5 HMVP entries ●By default, the zero vector

[0185] The number of candidates in the list is controlled by a variable derived from the slice header. In merge mode, up to the first six entries of this list are used, and in AMVP mode, the first two entries of this list are used. And the list complies with the shared merge list space requirement (sharing the same list within an SMR).

[0186] In addition to the BV predictor candidate list mentioned above, JVET-N0843 also proposes simplifying the pruning process between HMVP candidates and existing merge candidates (A1, B1). In this simplified case, the first HMVP candidate is only compared with one or more spatial merge candidates, so a maximum of two pruning operations are required.

[0187] 2.3.2.1.1 Decryption Process

[0188] 8.6.2.2 Derivation process for IBC luma motion vector prediction This process is only invoked if CuPredMode[xCb][yCb] is equal to MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture. The inputs to this process are: the luminance position (xCb, yCb) of the top left sample of the current luma coding block relative to the top left luma sample of the current picture, a variable cbWidth that specifies the width of the current coding block in luminance samples, A variable cbHeight that defines the height of the current coding block in luma samples. The output of this process is: - Luminance motion vectors at 1 / 16 fractional sample precision mvL.

[0189] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows: xSmr=IsInSmr[xCb][yCb]?SmrX[xCb][yCb]:xCb (8-910) ySmr=IsInSmr[xCb][yCb]?SmrY[xCb][yCb]:yCb (8-911) smrWidth=IsInSmr[xCb][yCb]?SmrW[xCb][yCb]:cbWidth (8-912) smrHeight=IsInSmr[xCb][yCb]?SmrH[xCb][yCb]:cbHeight (8-913) smrNumHmvpIbcCand=IsInSmr[xCb][yCb]?NumHmvpSmrIbcCand:NumHmvpIbcCand (8-914)

[0190] The luminance motion vector mvL is derived by the following sequential steps: The process of deriving spatial motion vector candidates from neighboring coding units, as specified in section 1.8.6.2.3, is invoked with inputs the position (xCb, yCb) of the luma coding block set equal to (xSmr, ySmr), the width cbWidth of the luma coding block and the height cbHeight of the luma coding block set equal to smrWidth and smrHeight, and the outputs are the availability flags availableFlagA1, availableFlagB1, and the motion vectors mvA1 and mvB1.

[0191] 2. The motion vector candidate list mvCandList is configured as follows: i=0 if(availableFlagA1) mvCandList[i++]=mvA1 if(availableFlagB1) mvCandList[i++]=mvB1(8-915)

[0192] 3. The variable numCurrCand is set equal to the number of merge candidates in the mvCandList.

[0193] 4. If numCurrCand is less than MaxNumMergeCand and smrNumHmvpIbcCand is greater than 0, the IBC history-based motion vector candidate derivation process specified in 8.6.2.4 is invoked with mvCandList, isInSmr set equal to isInSmr[xCb][yCb], and numCurrCand as inputs, and the modified mvCandList and numCurrCand as outputs.

[0194] 5. If numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand is equal to MaxNumMergeCand: 1. mvCandList[numCurrCand][0] is set equal to 0. 2. mvCandList[numCurrCand][1] is set equal to 0. 3.numCurrCand is incremented by 1.

[0195] 6. The variable mvIdx is derived as follows: mvIdx=general_merge_flag[xCb][yCb]?merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-916)

[0196] 7. The following assignments will be made: mvL[0]=mergeCandList[mvIdx][0] (8-917) mvL[1]=mergeCandList[mvIdx][1] (8-918)

[0197] 2.3.2.2 IBC Size Limits

[0198] In recent VVC and VTM5, in previous VTM and VVC versions, it was proposed to explicitly use a syntax constraint to disable IBC mode for 128x128 in addition to the current bitstream constraint, so that the presence of the IBC flag would depend on CU size < 128x128.

[0199] 2.3.2.3 Shared Merge List for IBC

[0200] To reduce decoder complexity and support parallel coding, JVET-M0147 proposes sharing the same merge candidate list for all leaf Coding Units (CUs) of an ancestor node in a CU partition tree to enable parallel processing of small skip / merge coded CUs. The ancestor node is called a merge-share node. A shared merge candidate list is generated at the merge-share node, making it appear as if it is a leaf CU.

[0201] Specifically, the following may apply: - Use merge lists between very small blocks (e.g., two adjacent 4x4 blocks) if the block has 32 or fewer luma samples and is split into two 4x4 child blocks. However, if the luminance samples of a block are greater than 32, after splitting at least one child block will be smaller than the threshold (32) and all of these split children will share the same merge list (e.g. 16x4 or 4x16 split ternary, or 8x8 split 4 times).

[0202] Such restrictions apply only to IBC merge mode.

[0203] 2.4 Syntax Tables and Semantics for Coding Units and Merge Modes

[0204] 7.3.5.1 General Slice Segment Header Syntax

[0205] [Table 4]

[0206] 7.3.7.5 Coding Unit Syntax

[0207] [Table 5]

[0208] [Table 6]

[0209] 7.3.7.7 Merge Data Syntax

[0210] [Table 7]

[0211] [Table 8]

[0212] 7.4.6.1 General slice header semantics six_minus_max_num_merge_cand specifies the maximum number of merge Motion Vector Prediction (MVP) candidates supported by a slice, subtracted from 6. The maximum number of merge MVP candidates, MaxNumMergeCand, is derived as follows: MaxNumMergeCand=6-six_minus_max_num_merge_cand (7-57) The value of MaxNumMergeCand is in the range of 1 to 6. five_minus_max_num_subblock_merge_cand specifies the maximum number of subblock-based merge MVP (Motion Vector Prediction) candidates supported in a slice, subtracted from 5. If five_minus_max_num_subblock_merge_cand is not present, it is inferred to be equal to 5 - sps_sbtmvp_enabled_flag. The maximum number of subblock-based merge MVP candidates, MaxNumSubblockMergeCand, is derived as follows: MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand (7-58) The value of MaxNumSubblockMergeCand is in the range of 1 to 5.

[0213] 7.4.8.5 Coding Unit Syntax pred_mode_flag equal to 0 specifies that the current coding unit is coded in inter prediction mode. pred_mode_flag equal to 1 specifies that the current coding unit is coded in intra prediction mode. If pred_mode_flag is not present, it is inferred as follows: - If cbWidth is equal to 4 and cbHeight is equal to 4, pred_mode_flag is inferred to be equal to 1. Otherwise, pred_mode_flag is inferred to be equal to 1 when decoding an I slice, and to be equal to 0 when decoding a P or B slice, respectively. The variables CuPredMode[x][y] are derived as follows for x=x0..x0+cbWidth-1 and y=y0..y0+cbHeight-1: If -pred_mode_flag is equal to 0, CuPredMode[x][y] is set equal to MODE_INTER. - Or (if pred_mode_flag is equal to 1), CuPredMode[x][y] is set equal to MODE_INTRA.

[0214] If pred_mode_ibc_flag is equal to 1, it specifies that the current coding unit is coded in IBC prediction mode. If pred_mode_ibc_flag is equal to 0, it specifies that the current coding unit is not coded in IBC prediction mode. If pred_mode_ibc_flag is not present, it is inferred as follows: - If cu_skip_flag[x0][y0] is equal to 1, and cbWidth is equal to 4, and cbHeight is equal to 4, then pred_mode_ibc_flag is inferred to be equal to 1. Alternatively, if both cbWidth and cbHeight are equal to 128, pred_mode_ibc_flag is inferred to be 0. Alternatively, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag when decoding an I slice, and 0 when decoding a P or B slice, respectively. If pred_mode_ibc_flag is equal to 1, the variable CuPredMode[x][y] is set equal to MODE_IBC for x=x0..x0+cbWidth-1 and y=y0..y0+cbHeight-1.

[0215] general_merge_flag[x0][y0] specifies whether to infer inter prediction parameters in the current coding unit from neighboring inter prediction partitions. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If general_merge_flag[x0][y0] is not present, it is inferred as follows: - If cu_skip_flag[x0][y0] is equal to 1, general_merge_flag[x0][y0] is inferred to be equal to 1. - Otherwise, general_merge_flag[x0][y0] is inferred to be equal to 0.

[0216] mvp_l0_flag[x0][y0] specifies the motion vector predictor index of list0, where x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If mvp_l0_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0217] mvp_l1_flag[x0][y0] has the same semantics as mvp_l0_flag, with l0 and list0 replaced by l1 and list1, respectively. inter_pred_idc[x0][y0] specifies whether list0, list1, or bidirectional prediction is used for the current coding unit according to Table 7-10. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0218] [Table 9]

[0219] If inter_pred_idc[x0][y0] is not present, it is inferred to be equal to PRED_L0.

[0220] 7.4.8.7 Merge Data Semantics regular_merge_flag[x0][y0], when equal to 1, specifies that the regular merge mode is used to generate inter prediction parameters for the current coding unit. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If regular_merge_flag[x0][y0] is not present, it is inferred as follows: - regular_merge_flag[x0][y0] is inferred to be equal to 1 if all of the following conditions are true: -sps_mmvd_enabled_flag equals 0. -general_merge_flag[x0][y0] equals 1. -cbWidth*cbHeight equals 32. - Otherwise, regular_merge_flag[x0][y0] is inferred to be equal to 0.

[0221] mmvd_merge_flag[x0][y0] equal to 1 specifies that the merge mode with motion vector differential is used to generate inter prediction parameters for the current coding unit. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If mmvd_merge_flag[x0][y0] does not exist, it is inferred as follows: - mmvd_merge_flag[x0][y0] is inferred to be equal to 1 if all of the following conditions are true: -sps_mmvd_enabled_flag equals 1. -general_merge_flag[x0][y0] equals 1. -cbWidth*cbHeight equals 32. -regular_merge_flag[x0][y0] equals 0. - Otherwise, mmvd_merge_flag[x0][y0] is inferred to be equal to 0.

[0222] mmvd_cand_flag[x0][y0] specifies whether to use the first (0) or second (1) candidate in the merge candidate list in the motion vector differential derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array index x0,y0 specifies the position (x0,y0) of the top-left luminance sample of the considered coding block relative to the top-left luminance sample of the picture. If mmvd_cand_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0223] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 7-12. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0224] [Table 10]

[0225] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 7-13. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture.

[0226] [Table 11]

[0227] Both components of the merge+MVD offset MmvdOffset[x0][y0] are derived as follows: MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] (7-124) MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] (7-125)

[0228] merge_subblock_flag[x0][y0] specifies whether to infer subblock-based inter prediction parameters in the current coding unit from neighboring blocks. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If merge_subblock_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0229] merge_subblock_idx[x0][y0] specifies the merge candidate index in the subblock-based merge candidate list, where x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If merge_subblock_idx[x0][y0] is not present, it is inferred to be equal to 0.

[0230] ciip_flag[x0][y0] specifies whether to combine inter-picture merging and intra-picture prediction for the current coding unit. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If ciip_flag[x0][y0] is not present, it is inferred to be equal to 0. If ciip_flag[x0][y0] is equal to 1, the variable IntraPredModeY[x][y], where x=xCb..xCb+cbWidth-1 and y=yCb..yCb+cbHeight-1, is set equal to INTRA_PLANAR. When decoding a B slice, the variable MergeTriangleFlag[x0][y0], which specifies whether to use triangle-based motion compensation to generate prediction samples for the current coding unit, is derived as follows: - MergeTriangleFlag[x0][y0] is set equal to 1 if all of the following conditions are true: -sps_triangle_enabled_flag equals 1. -slice_type equals B. -general_merge_flag[x0][y0] equals 1. -MaxNumTriangleMergeCand is 2 or greater. -cbWidth*cbHeight is 64 or greater. -regular_merge_flag[x0][y0] equals 0. -mmvd_merge_flag[x0][y0] equals 0. -merge_subblock_flag[x0][y0] equals 0. -ciip_flag[x0][y0] equals 0. Otherwise, MergeTriangleFlag[x0][y0] is set equal to 0.

[0231] merge_triangle_split_dir[x0][y0] specifies the split direction for the merge triangle mode. The array index x0,y0 specifies the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If merge_triangle_split_dir[x0][y0] is not present, it is inferred to be equal to 0.

[0232] merge_triangle_idx0[x0][y0] specifies the first merge candidate index in the triangle-based motion compensation candidate list, where x0, y0 specifies the position (x0, y0) of the top-left luminance sample of the considered coding block relative to the top-left luminance sample of the picture. If merge_triangle_idx0[x0][y0] is not present, it is inferred to be equal to 0.

[0233] merge_triangle_idx1[x0][y0] specifies the second merge candidate index in the triangle-based motion compensation candidate list, where x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If merge_triangle_idx1[x0][y0] is not present, it is inferred to be equal to 0.

[0234] merge_idx[x0][y0] specifies the merge candidate index in the merge candidate list, where x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If merge_idx[x0][y0] does not exist, it is inferred as follows: - If mmvd_merge_flag[x0][y0] is equal to 1, merge_idx[x0][y0] is inferred to be equal to mmvdmmvd_cand_flag[x0][y0]. - Otherwise (if mmvd_merge_flag[x0][y0] is equal to 0), merge_idx[x0][y0] is inferred to be equal to 0.

[0235] 3. Examples of technical problems that the embodiments aim to solve

[0236] The current IBC may have the following problems: 1. For P / B slices, the number of IBC motion (BV) candidates that can be added to the IBC motion list is set to the same size as the normal merge list. Therefore, if the normal merge mode is disabled, the IBC merge mode is also disabled. However, it is recommended to enable the IBC merge mode even when the normal merge mode is disabled. 2. IBC AMVP and IBC merge modes share the same motion (BV) candidate list construction process. The size of the list, indicated by the maximum number of merge candidates (e.g., MaxNumMergeCand), is signaled in the slice header. For IBC AMVP, the BV predictor can only choose from one of the two IBC motion candidates. a. If the size of the BV candidate list is 1, then in this case, signaling of the BV predictor index is not required. b. If the BV candidate list size is 0, both IBC merging and AMVP mode are disabled. However, in the current design, the BV predictor index is still signaled. 3. IBC may be enabled at the sequence level, but due to the size of the signaled BV list, IBC may be disabled, but the bit-wasting IBC mode and associated syntax indications may still be signaled.

[0237] 4. Exemplary Techniques and Embodiments

[0238] The following detailed inventions should be considered as examples to illustrate the general concept. These inventions should not be construed in a narrow sense. Furthermore, these inventions can be combined in any way.

[0239] In the present invention, DMVD (decoder side Motion Vector Derivation) includes methods such as DMVR and FRUC that perform motion estimation to derive or improve motion information of blocks / sub-blocks, and BIO that performs sample-wise motion improvement.

[0240] maxIBCCandNum indicates the maximum number of BV candidates in the BV candidate list (i.e., the size of the BV candidate list) that allows a BV to be derived or predicted in the BV candidate list.

[0241] Denote the maximum number of IBC merge candidates as maxIBCMrgNum, and the maximum number of IBCAMVP candidates as maxIBCAMVPN. Note that skip mode may be treated as a special merge mode with coefficients equal to all zeros. 1. IBC (e.g., IBC AMVP and / or IBC merge) mode may be disabled for a picture / slice / tile / tile group / brick or other video unit even if IBC is enabled for the sequence. a. In one example, an indication of whether IBC AMVP and / or IBC merge mode is enabled may be signaled at the picture / slice / tile / tile group / brick or other video unit level (PPS, APS, slice header, picture header, etc.). i. In one example, the indication may be a flag. ii. In one example, the indication may be the number of allowed BV predictors, where the number of allowed BV predictors equals 0, indicating that IBC AMVP and / or merge mode are disabled. 1) Alternatively, the number of allowed BV predictors may be signaled in the prediction scheme. For example, X minus the number of allowed BV predictors may be signaled, where X is a fixed number such as 2. iii. Additionally, alternatively, the indication may be conditionally signaled based on whether the video content is screen content or the sps_ibc_enabled_flag. b. In one example, if the video content is not screen content (such as camera-captured content or mixed content), IBC mode may be disabled. c. In one example, if the video content is camera-captured content, IBC mode may be disabled.

[0242] 2. Whether to signal the use of IBC (eg, pred_mode_ibc_flag) and IBC-related syntax elements may depend on maxIBCCandNum (eg, maxIBCCandNum is set equal to MaxNumMergeCand). a. In one example, if maxIBCCandNum is equal to 0, signaling of the use of IBC skip mode (eg, cu_skip_flag) for I-slices may be skipped and IBC is inferred to be disabled. b. In one example, if maxIBCMrgNum is equal to 0, signaling of the use of IBC skip mode (eg, cu_skip_flag) for I-slices may be skipped, and IBC skip mode is inferred to be disabled. c. In one example, if maxIBCCandNum is equal to 0, signaling of use of IBC merge / IBC AMVP mode (e.g., pred_mode_ibc_flag) may be skipped and IBC is inferred to be disabled. d. In one example, if maxIBCCandNum is equal to 0 and the current block is coded in IBC mode, merge mode signaling (e.g., general_merge_flag) may be skipped and IBC merge mode is inferred to be disabled. i. Alternatively, the current block is inferred to be coded in IBC AMVP mode. e. In one example, if maxIBCMrgNum is equal to 0 and the current block is coded in IBC mode, merge mode signaling (e.g., general_merge_flag) may be skipped and IBC merge mode is inferred to be disabled. i. Alternatively, the current block is inferred to be coded in IBC AMVP mode.

[0243] 3. Whether to signal syntax elements related to motion vector differentials for IBC AMVP mode may depend on maxIBCC and Num (eg, maxIBCC and Num are set equal to MaxNumMergeCand). a. In one example, if maxIBCCandNum is equal to 0, signaling of motion vector predictor index for IBC AMVP mode (eg, mvp_l0_flag) may be skipped and IBC AMVP mode is inferred to be disabled. b. In one example, if maxIBCCandNum is equal to 0, motion vector differential signaling (eg, mvd_coding) may be skipped and IBC AMVP mode is inferred to be disabled. c. In one example, the motion vector predictor index and / or the motion vector predictor accuracy and / or the motion vector differential accuracy of IBC AMVP mode may be coded under the condition of maxIBCCandNum being greater than K (e.g., K=0 or 1). i. In one example, if maxIBCCandNum is equal to 1, the motion vector predictor index (eg, mvp_l0_flag) for IBC AMVP mode may not be signaled. 1) In one example, the motion vector predictor index (eg, mvp_l0_flag) for IBC AMVP mode may be inferred to be a value such as 0 in this case. ii. In one example, if maxIBCCandNum is equal to 0, signaling of the precision of the motion vector predictor and / or the precision of the motion vector differential in IBC AMVP mode (eg, amvr_precision_flag) may be skipped. iii. In one example, signaling of the precision of the motion vector predictor and / or the precision of the motion vector differential (eg, amvr_precision_flag) for IBC AMVP mode may be subject to the condition that maxIBCC and Num are greater than 0.

[0244] 4. It is proposed that maxIBCCandNum may be decoupled from the normal maximum number of merge candidates. In one example, maxIBCCandNum may be signaled directly. i. Alternatively, if IBC is enabled for a slice, a conforming bitstream shall satisfy maxIBCCandNum greater than 0. b. In one example, predictive coding of maxIBCCandNum and other syntax elements / fixed values ​​may be signaled. i. In one example, the difference between the normal merge list size and maxIBCCandNum may be coded. ii. In one example, (K-maxIBCCandNum) may be coded as, for example, K=5 or 6. iii. In one example, (maxIBCCandNum-K) may be coded as, for example, K=0 or 2. c. In one example, the maxIBCMrgNum and / or maxIBCAMVPN indications may be signaled according to the methods described above.

[0245] 5. maxIBCCandNum may be set equal to Func(maxIBCMrgNum, maxIBCAMVPN). a. In one example, maxIBCAMVPN is fixed at 2 and maxIBCCandNum is set equal to Func(maxIBCMrgNum, 2). b. In one example, Func(a, b) returns the larger value between two variables a and b.

[0246] 6. maxIBCCandNum may be determined according to the coded mode information of one block, i.e., the number of BV candidates that may be added to the BV candidate list may depend on the mode information of the block. a. In one example, if one block is coded in IBC merge mode, maxIBCCandNum may be set equal to maxIBCMrgNum. b. In one example, if one block is coded in IBC AMVP mode, maxIBCCandNum may be set equal to maxIBCAMVPN.

[0247] 7. A conforming bitstream shall satisfy that the decoded IBCAMVP or IBC merge index is less than maxIBCCandNum. In one example, a conforming bitstream satisfies that the decoded IBCAMVP index is less than maxIBCAMVPN. b. In one example, a conforming bitstream satisfies that the decoded IBC merge index is less than maxIBCMrgNum (eg, 2).

[0248] 8. If the IBC AMVP or IBC merge candidate index fails to identify a BV candidate in the BV candidate list (e.g., if the IBC AMVP or merge candidate index is greater than or equal to maxIBCCandNum, if the decoded IBC AMVP index is greater than or equal to maxIBCAMVPN, or if the decoded IBC merge index is greater than or equal to maxIBCMrgNum), a default prediction block may be used. In one example, all samples of the default predicted intra block are set to (1<<(bit depth-1)). b. In one example, a default BV may be assigned to the block.

[0249] 9. If the IBC AMVP or IBC merge candidate index fails to identify a BV candidate in the BV candidate list (e.g., if the IBC AMVP or merge candidate index is greater than or equal to maxIBCCandNum, if the decoded IBC AMVP index is greater than or equal to maxIBCAMVPN, or if the decoded IBC merge index is greater than or equal to maxIBCMrgNum), the block may be treated as an IBC block with an invalid BV. In one example, the processing applied to blocks with invalid BVs may be applied to the block.

[0250] 10. A supplemental BV candidate list may be constructed under certain conditions, such as when the decoded IBC AMVP or IBC merge index is greater than or equal to maxIBCCandNum. In one example, the supplemental BV candidate list may be constructed in one or more of the following steps (in sequence or in an interleaved manner): i.Addition of HMVP candidates 1) Ascending order of HMVP candidate index (starting from the Kth entry in the HMVP table (e.g., K=0)) 2) Descending order of HMVP candidate index (from the Kth (e.g., K=0 or K=maxIBCCandNum-IBC AMVP / IBC merge index or K=maxIBCCandNum-1-IBC AMVP / IBC merge index) entry to the last entry in the HMVP table) ii. Virtual BV candidates derived from available candidates using a BV candidate list with maxIBCCandNum candidates. 1) In one example, an offset may be added to the offset to the horizontal and / or vertical components of the BV candidate to obtain a virtual BV candidate. 2) In one example, an offset may be added to the offset to the horizontal and / or vertical components of an HMVP candidate in an HMVP table to obtain a virtual BV candidate. iii. Adding a default candidate 1) In one example, (0,0) may be added as a default candidate. 2) In one example, the default candidate may be derived according to the current block size.

[0251] 11. maxIBCAMVPN does not have to be equal to 2. a. Additionally, alternatively, instead of signaling a flag indicating the motion vector predictor index (eg, mvp_l0_flag), an index that may be greater than 1 may be signaled. i. In one example, the index may be binarized using unary / truncated unary / fixed length / exponential-golomb / other binarization methods. ii. In one example, the bins of the binary bin string of the index may be context coded or bypass coded. 1) In one example, the first K (eg, K=1) bins of the binary bin string of the index may be context coded, and the remaining bins may be bypass coded. b. In one example, maxIBCAMVPN may be greater than maxIBCMrgNum. i. Alternatively, the first maxIBCMrgNum BV candidates in the BV candidate list may also be used for the IBC merge-coded block.

[0252] 12. An IBC BV candidate list may be constructed even if maxIBCCandNum is set to 0 (e.g., MaxNumMergeCand=0). In one example, the merge list may be constructed as when IBC AMVP mode is enabled. b. In one example, if IBC merge mode is disabled, the merge list may be constructed to include a maximum of maxIBCAMVPN BV candidates.

[0253] 13. In the above method, the term "maxIBCCandNum" may be replaced with "maxIBCMrgNum" or "maxIBCAMVPN".

[0254] 14. In the above method, the terms "maxIBCCandNum" and / or "maxIBCMrgNum" may be replaced with MaxNumMergeCand, which may represent the maximum number of merge candidates in a regular merge list.

[0255] 5. Embodiments

[0256] New additions over JVET-N1001-v5 are enclosed in double bold curly brackets, i.e., {{a}} indicates that an "a" has been added, and deletions are enclosed in double bold square brackets, i.e., [[a]] indicates that an "a" has been deleted.

[0257] 5.1 Embodiment #1 An indication of the maximum number of IBC motion candidates (eg, in the case of IBC AMVP and / or IBC merge) may be signaled in the slice header / PPS / APS / DPS.

[0258] 7.3.5.1 General Slice Segment Header Syntax

[0259] [Table 12]

[0260] [Table 13]

[0261] {{max_num_merge_cand_minus_max_num_IBC_cand specifies the maximum number of IBC merge mode candidates supported in a slice minus MaxNumMergeCand. The maximum number of IBC merge mode candidates, MaxNumIBCMergeCand, is derived as follows: MaxNumIBCMergeCand=MaxNumMergeCand-max_num_merge_cand_minus_max_num_IBC_cand If max_num_merge_cand_minus_max_num_IBC_cand is present, the value of MaxNumIBCMergeCand shall range from 2 (or 0) to MaxNumMergeCand. If max_num_merge_cand_minus_max_num_IBC_cand is not present, MaxNumIBCMergeCand is set equal to 0. If MaxNumIBCMergeCand is 0, IBC merge mode and IBCAMVP mode are not allowed on the current slice. Alternatively, the signaling of the MaxNumIBCMergeCand indication may be replaced by:

[0262] [Table 14]

[0263] MaxNumIBCMergeCand=5-five_minus_max_num_IBC_cand}}

[0264] 5.2 Embodiment #2 It is proposed to change the maximum number of IBC BV lists from MaxNumMergeCand, which controls both IBC and regular intermode, to a separate variable, MaxNumIBCMergeCand.

[0265] 8.6.2.2 IBC Luminance Motion Vector Prediction Derivation Process This process is only invoked if CuPredMode[xCb][yCb] is equal to MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luma coding block relative to the top-left luma sample of the current picture. The inputs to this process are: the luminance position (xCb, yCb) of the top left sample of the current luma coding block relative to the top left luma sample of the current picture, a variable cbWidth that specifies the width of the current coding block in luminance samples, A variable cbHeight that defines the height of the current coding block in luma samples.

[0266] The output of this process is: - Luminance motion vectors at 1 / 16 fractional sample precision mvL. The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows: xSmr=IsInSmr[xCb][yCb]?SmrX[xCb][yCb]:xCb (8-910) ySmr=IsInSmr[xCb][yCb]?SmrY[xCb][yCb]:yCb (8-911) smrWidth=IsInSmr[xCb][yCb]?SmrW[xCb][yCb]:cbWidth (8-912) smrHeight=IsInSmr[xCb][yCb]?SmrH[xCb][yCb]:cbHeight (8-913) smrNumHmvpIbcCand=IsInSmr[xCb][yCb]?NumHmvpSmrIbcCand:NumHmvpIbcCand (8-914)

[0267] The luminance motion vector mvL is derived by the following sequential steps: The process of deriving spatial motion vector candidates from neighboring coding units, as specified in Section 1.8.6.2.3, is invoked with inputs the position (xCb, yCb) of the luma coding block set equal to (xSmr, ySmr), the width cbWidth of the luma coding block set equal to smrWidth and the height cbHeight of the luma coding block set equal to smrHeight, and the outputs are the availability flags availableFlagA1, availableFlagB1, and the motion vectors mvA1 and mvB1.

[0268] 2. The motion vector candidate list mvCandList is configured as follows: i=0 if(availableFlagA1) mvCandList[i++]=mvA1 if(availableFlagB1) mvCandList[i++]=mvB1(8-915)

[0269] 3. The variable numCurrCand is set equal to the number of merge candidates in the mvCandList.

[0270] 4. If numCurrCand is less than [[MaxNumMergeCand]]{{MaxNumIBCMergeCand}} and smrNumHmvpIbcCand is greater than 0, then the IBC history-based motion vector candidate derivation process as specified in 8.6.2.4 is invoked with mvCandList, isInSmr set equal to IsInSmr[xCb][yCb], and numCurrCand as inputs, and the modified mvCandList and numCurrCand as outputs.

[0271] 5. If numCurrCand is less than [[MaxNumMergeCand]]{{MaxNumIBCMergeCand}}, the following applies until numCurrCand is equal to MaxNumMergeCand: 1. mvCandList[numCurrCand][0] is set equal to 0. 2. mvCandList[numCurrCand][1] is set equal to 0. 3.numCurrCand is incremented by 1.

[0272] 6. The variable mvIdx is derived as follows: mvIdx=general_merge_flag[xCb][yCb]?merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-916)

[0273] 7. The following assignments will be made: mvL[0]=mergeCandList[mvIdx][0] (8-917) mvL[1]=mergeCandList[mvIdx][1] (8-918) In one example, if the current block is in IBC merge mode, MaxNumIBCMergeCand is set to MaxNumMergeCand, and if the current block is in IBC AMVP mode, MaxNumMergeCand is set to 2. In one example, MaxNumIBCMergeCand is derived from signaled information, for example using embodiment #1.

[0274] 5.3 Embodiment #3 In one example, conditional signaling of IBC-related syntax elements according to the maximum number of allowed IBC candidates, MaxNumIBCMergeCand is set to MaxNumMergeCand.

[0275] 7.3.7.5 Coding Unit Syntax

[0276] [Table 15]

[0277] cu_skip_flag[x0][y0] equal to 1 specifies that, for the current coding unit, when decoding a P or B slice, no syntax elements other than the IBC mode flag pred_mode_ibc_flag[x0][y0] {{if MaxNumIBCMergeCand is greater than 0}} and one or more syntax elements of the merge_data() syntax structure are parsed after cu_skip_flag[x0][y0]; when decoding an I slice {{and MaxNumIBCMergeCand is greater than 0}}, no syntax elements other than merge_idx[x0][y0] are parsed after cu_skip_flag[x0][y0]. cu_skip_flag[x0][y0] equal to 0 specifies that the coding unit is not skipped. The array indexes x0,y0 specify the position (x0,y0) of the top-left luma sample of the considered coding block relative to the top-left luma sample of the picture. If cu_skip_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0278] pred_mode_ibc_flag equal to 1 specifies that the current coding unit is coded in IBC prediction mode. pred_mode_ibc_flag equal to 0 specifies that the current coding unit is not coded in IBC prediction mode. If pred_mode_ibc_flag is not present, it is inferred as follows: - If cu_skip_flag[x0][y0] is equal to 1, and cbWidth is equal to 4, and cbHeight is equal to 4, then pred_mode_ibc_flag is inferred to be equal to 1. Alternatively, if both cbWidth and cbHeight are equal to 128, pred_mode_ibc_flag is inferred to be equal to 0. -Alternatively, when decoding an I slice {{and MaxNumIBCMergeCand is greater than 0}}, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag, and when decoding a P slice or a B slice, pred_mode_ibc_flag is 0.

[0279] 5.4 Embodiment #3 Conditional signaling of IBC-related syntax elements according to the maximum number of allowed IBC candidates

[0280] [Table 16]

[0281] In one example, MaxNumIBCAMVPCand may be set to MaxNumIBCMergeCand or MaxNumMergeCand.

[0282] 20A shows an example method 2000 for video processing. The method 2000 includes, at act 2002, performing a conversion between a video domain of a video and a bitstream representation of the video, the bitstream representation selectively including MVD (Motion Vector Difference)-related syntax elements for an IBC Advanced Motion Vector Prediction (AMVP) mode based on a maximum number of first-type IBC (Intra Block Copy) candidates used during the video domain conversion. In some embodiments, when the IBC mode is applied, samples of the video domain are predicted from other samples in a video picture corresponding to the video domain.

[0283] 20B illustrates an example method 2005 for video processing. The method 2005 includes, at act 2007, determining to disable indication of use of Intra Block Copy (IBC) mode for the video domain and enable use of IBC mode at a sequence level of the video for conversion between a video domain of the video and a bitstream representation of the video.

[0284] The method 2005 includes performing a transform based on the determination, at operation 2009. In some embodiments, when an IBC mode is applied, samples of the video region are predicted from other samples in the video picture that correspond to the video region.

[0285] 20C shows an example method 2010 for video processing. The method 2010 includes, at act 2012, performing a conversion between a video domain of a video and a bitstream representation of the video, the bitstream representation selectively including an indication regarding the use of an IBC mode and / or one or more IBC-related syntax elements based on a maximum number of a first type of IBC (Intra Block Copy) candidate used during the video domain conversion. In some embodiments, when the IBC mode is applied, samples of the video domain are predicted from other samples in the video picture that correspond to the video domain.

[0286] 20D shows an example method 2020 for video processing. The method 2020 includes, at operation 2022, performing a conversion between a video domain of a video and a bitstream representation of the video, wherein an indication of a maximum number of Intra Block Copy (IBC) candidates of a first type used during the video domain conversion is signaled in the bitstream representation independent of a maximum number of inter-mode merge candidates used during the conversion. In some embodiments, when an IBC mode is applied, samples of the video domain are predicted from other samples in the video picture that correspond to the video domain.

[0287] 20E shows an example method 2030 for video processing. The method 2030 includes, at act 2032, performing a conversion between a video domain of a video and a bitstream representation of the video, where a maximum number of Intra Block Copy (IBC) motion candidates used during the video domain conversion is a function of a maximum number of IBC merge candidates and a maximum number of Advanced Motion Vector Prediction (AMVP) candidates. In some embodiments, when an IBC mode is applied, samples of the video domain are predicted from other samples in the video picture that correspond to the video domain.

[0288] 20F shows an example method 2040 for video processing. The method 2040 includes, at operation 2042, performing a conversion between a video domain of a video and a bitstream representation of the video, where a maximum number of Intra Block Copy (IBC) motion candidates used during the video domain conversion is based on coded mode information of the video domain.

[0289] 20G illustrates an example method 2050 for video processing. The method 2050 includes, at operation 2052, performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a decoded Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) merge index or a decoded IBC merge index is less than a maximum number of Intra Block Copy (IBC) motion candidates.

[0290] 20H illustrates an example method 2060 for video processing. The method 2060 includes, at act 2062, determining that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index fails to identify a block vector candidate in a block vector candidate list during conversion between a video domain of a video and a bitstream representation of the video.

[0291] The method 2060 includes, at act 2064, using a default prediction block during the transform based on the determination.

[0292] 20I illustrates an example method 2070 for video processing. The method 2070 includes, at act 2072, determining that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index fails to identify a block vector candidate in a block vector candidate list during conversion between a video domain of a video and a bitstream representation of the video.

[0293] The method 2070 includes, at act 2074, performing the transformation by treating the image region as having invalid block vectors based on the determination.

[0294] 20J shows an example method 2080 for video processing. The method 2080 includes, at act 2082, determining that an Intra Block Copy (IBC) alternative motion vector predictor candidate index or an IBC merge candidate index does not satisfy a condition during conversion between a video domain of a video and a bitstream representation of the video.

[0295] The method 2080 includes, at act 2084, generating a supplemental Block Vector (BV) candidate list based on the determination.

[0296] The method 2080 includes, at operation 2086, performing a transformation using the supplemental BV candidate list.

[0297] 20K shows an example method 2090 for video processing. The method 2090 includes, at act 2092, performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the maximum number of Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) candidates is not two.

[0298] FIG. 21 is a block diagram of a video processing device 2100. The device 2100 may be used to implement one or more of the methods described herein. The device 2100 may be implemented in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. The device 2100 may include one or more processors 2102, one or more memories 2104, and video processing hardware 2106. The one or more processors 2102 may be configured to implement one or more methods described herein. The one or more memories 2104 may be used to store data and code used to implement the methods and techniques described herein. The video processing hardware 2106 may be used to implement the techniques described herein in a hardware circuit. The video processing hardware 2106 may be included partially or completely within the processor 2102 in the form of dedicated hardware, or a Graphical Processor Unit (GPU) or dedicated signal processing block.

[0299] In some embodiments, the video coding method may be performed using an apparatus implemented on a hardware platform, such as that described with reference to FIG.

[0300] Some embodiments of the disclosed technology include determining or deciding to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, an encoder uses or implements the tool or mode when processing video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, the conversion from blocks of video to a bitstream representation of video uses the video processing tool or mode when the video processing tool or mode is enabled based on the decision or determination. In another example, when a video processing tool or mode is enabled, a decoder processes the bitstream knowing that the bitstream has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of video to blocks of video is performed using the video processing tool or mode enabled based on the decision or determination.

[0301] Some embodiments of the disclosed techniques include deciding or determining to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, an encoder does not use the tool or mode when converting blocks of video into a bitstream representation of the video. In another example, when a video processing tool or mode is disabled, a decoder processes the bitstream knowing that the bitstream has not been modified using the video processing tool or mode that was enabled based on the decision or determination.

[0302] 22 is a block diagram illustrating an example video processing system 2200 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 2200. System 2200 may include an input 2202 for receiving video content. The video content may be received in a raw or uncompressed format, e.g., 8- or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 2202 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet or Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0303] System 2200 may include a coding component 2204 capable of implementing various encoding or coding methods described herein. Coding component 2204 may reduce the average bit rate of video from input 2202 to the output of coding component 2204, generating a coded representation of the video. Thus, coding techniques are sometimes referred to as video compression or video transcoding techniques. The output of coding component 2204 may be stored or transmitted via a connected communication, as represented by component 2206. The bitstream (or coded) representation of the video received at input 2202, stored, or communicated, may be used by component 2208 to generate pixel values ​​or displayable video that are transmitted to display interface 2210. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, while certain video processing operations are referred to as “coding” operations or tools, it will be understood that coding tools or operations are performed by an encoder and corresponding decoding tools or operations that reverse the decoding results are performed by a decoder.

[0304] Examples of peripheral bus interface units or display interface units may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), DisplayPort, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), PCI, IDE interfaces, etc. The techniques described herein may be implemented in various electronic devices such as mobile phones, laptops, smartphones, or other devices capable of digital data processing and / or video display.

[0305] In some embodiments, the following technical solutions can be implemented.

[0306] A1. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes MVD (Motion Vector Difference) related syntax elements for an IBC AMVP (Advanced Motion Vector Prediction) mode based on a maximum number of IBC (Intra Block Copy) candidates of a first type used during the video domain conversion, and wherein, when applying the IBC mode, samples of the video domain are predicted from other samples in the video picture corresponding to the video domain.

[0307] A2. The method described in Solution A1, wherein the MVD-related syntax elements include at least one of a coded motion vector predictor index, a coded precision of the motion vector predictor, a motion vector differential, and a coded precision of the motion vector differential in IBC AMVP mode.

[0308] A3. The method according to Solution A1, wherein the bitstream representation excludes signaling of IBC AMVP mode and MDV-related syntax elements for IBC AMVP mode, and the IBC AMVP mode is inferred to be disabled based on the maximum number of IBC candidates of the first type being less than or equal to K, where K is an integer.

[0309] A4. The method according to Solution A1, wherein the bitstream representation excludes signaling of IBC AMVP mode and motion vector differentials for IBC AMVP mode, and IBC AMVP mode is inferred to be disabled, based on the maximum number of IBC candidates of the first type being less than or equal to K, where K is an integer.

[0310] A5. The method described in Solution A1, wherein, based on the maximum number of IBC candidates of the first type being greater than K, the bitstream representation selectively includes at least one of a coded motion vector predictor index, a coded precision of the motion vector predictor, and a coded precision of the motion vector differential for IBC AMVP mode, where K is an integer.

[0311] A6. The method according to any of Solutions A3 to A5, wherein K=0 or K=1.

[0312] A7. The method according to Solution A1, wherein the bitstream representation excludes a motion vector predictor index for IBC AMVP mode based on the maximum number of IBC candidates of the first type being 1.

[0313] A8. The method according to Solution A19, wherein the motion vector predictor index for IBC AMVP mode is inferred to be zero.

[0314] A9. The method according to Solution A1, wherein the bitstream representation excludes the precision of the motion vector predictor and / or the precision of the motion vector differential in IBC AMVP mode based on the maximum number of IBC candidates of the first type being equal to zero.

[0315] A10. The method according to Solution A1, wherein the bitstream representation excludes the precision of the motion vector predictor and / or the precision of the motion vector differential in IBC AMVP mode based on the maximum number of IBC candidates of the first type being greater than zero.

[0316] A11. The method according to any of Solutions A1 to A10, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC motion candidates (denoted as maxIBCCandNum).

[0317] A12. The method according to any of Solutions A1 to A10, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC merge candidates (denoted as maxIBCMrgNum).

[0318] A13. The method according to any of Solutions A1 to A10, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC Advanced Motion Vector Prediction (AMVP) candidates (denoted as maxIBCAMVPNum).

[0319] A14. The method according to any of Solutions A1 to A10, wherein the maximum number of IBC candidates of the first type is signaled in the bitstream representation.

[0320] A15. A method of video processing comprising: determining that an indication of the use of IBC (Intra Block Copy) mode is disabled for a video region and that the use of IBC mode is enabled at the video sequence level for conversion between a video region of a video and a bitstream representation of the video; and performing the conversion based on the determination, wherein when the IBC mode is applied, samples of the video region are predicted from other samples in the video picture corresponding to the video region.

[0321] A16. The method of solution A15, wherein the video region corresponds to a picture, a slice, a tile, a tile group, or a brick of a video picture.

[0322] A17. The method of any one of solutions A15 to A16, wherein the bitstream representation has an indication associated with the decision.

[0323] A18. The method according to solution A17, wherein the instruction is signaled in a picture, slice, tile, brick, or APS (Adaptation Parameter Set).

[0324] A19. The method according to solution A18, wherein the instruction is signaled in a Picture Parameter Set (PPS), a slice header, a picture header, a tile header, a tile group header, or a brick header.

[0325] A20. The method according to solution A15, in which the IBC mode is disabled based on the video content different from the screen content.

[0326] A21. The method of solution A15, wherein the IBC mode is disabled based on the video content of the video being camera-captured content.

[0327] A22. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of said video, wherein said bitstream representation selectively includes instructions for the use of an IBC (Intra Block Copy) mode and / or one or more IBC-related syntax elements based on a maximum number of IBC candidates of a first type used during the video domain conversion, and wherein, when an IBC mode is applied, samples of the video domain are predicted from other samples in a video picture corresponding to the video domain.

[0328] A23. The method according to Solution A22, wherein the maximum number of IBC candidates of the first type is set equal to the maximum number of merging candidates for inter-coded blocks.

[0329] A24. The method according to solution A22, wherein the bitstream representation excludes IBC skip mode and IBC skip mode signaling, and IBC skip mode is inferred to be disabled, based on the maximum number of IBC candidates of the first type being equal to zero.

[0330] A25. The method according to solution A22, wherein, based on the maximum number of IBC candidates of the first type being equal to zero, the bitstream representation excludes signaling of IBC merge mode or IBC Advanced Motion Vector Prediction (AMVP) mode, and the IBC mode is inferred to be disabled.

[0331] A26. The method according to solution A22, wherein the bitstream representation excludes merge mode signaling and the IBC merge mode is inferred to be disabled based on the maximum number of IBC candidates of the first type being equal to zero.

[0332] A27. A method according to any of solutions A22 to A26, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC motion candidates (denoted by maxIBCCandNum), the maximum number of IBC merge candidates (denoted by maxIBCMrgNum), or the maximum number of IBC AMVP (Advanced Motion Vector Prediction) candidates (denoted by maxIBCAMVPNum).

[0333] A28. The method according to any of Solutions A22 to A26, wherein the maximum number of IBC candidates of the first type is signaled in the bitstream representation.

[0334] A29. The method according to any of Solutions A1 to A28, wherein the conversion comprises generating pixel values ​​in the video domain from the bitstream representation.

[0335] A30. A method according to any of Solutions A1 to A28, wherein the conversion generates a bitstream representation from pixel values ​​in the video domain.

[0336] A31. A video processing device comprising a processor configured to implement the methods described in one or more of Solutions A-A30.

[0337] A32. A computer-readable medium storing program code that, when executed, causes a processor to implement a method according to one or more of solutions A1-A30.

[0338] In some embodiments, the following technical solutions can be implemented.

[0339] B1. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein an indication of a maximum number of IBC (Intra Block Copy) candidates of a first type to be used during the conversion of the video domain is signaled in the bitstream representation independent of a maximum number of inter-mode merge candidates used in the conversion, and when the IBC mode is applied, samples of the video domain are predicted from other samples in the video picture corresponding to the video domain.

[0340] B2. The method according to Solution B1, wherein the maximum number of IBC candidates of the first type is signaled directly in the bitstream representation.

[0341] B3. The method according to Solution B1, wherein the maximum number of IBC candidates of the first type is greater than zero due to the IBC mode being enabled for video domain transformations.

[0342] B4. The method according to any of Solutions B1 to B3, further comprising predictively coding the maximum number of IBC candidates of the first type (denoted by maxNumIBC) using another value.

[0343] B5. The method according to solution 4, wherein the indication of the maximum number of IBC candidates of the first type is set to the difference after predictive coding.

[0344] B6. The method according to solution B5, wherein the difference is between another value and the maximum number of IBC candidates of the first type.

[0345] B7. The method according to Solution B6, wherein another value is the size of the normal merge list (S), (S-maxNumIBC) being signaled in the bitstream, where S is an integer.

[0346] B8. The method according to Solution B6, wherein another value is a fixed integer (K) and (K-maxNumIBC) is signaled in the bitstream.

[0347] B9. The method according to solution B8, where K=5 or K=6.

[0348] B10. The method according to any of solutions B6 to B9, wherein the maximum number of IBC candidates of the first type is derived to be (K - an indication of the maximum number signaled in the bit stream).

[0349] B11. The method according to Solution B5, wherein the difference is between the maximum number of IBC candidates of the first type and the other value.

[0350] B12. The method according to solution B11, wherein another value is a fixed integer (K), (maxNumIBC-K) being signaled in the bitstream.

[0351] B13. The method according to solution B12, wherein K=0 or K=2.

[0352] B14. The method according to solutions B11 to B13, wherein the maximum number of IBC candidates of the first type is derived to be (maximum number indication signaled in the bitstream + K).

[0353] B15. The method according to any of Solutions B1 to B14, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC motion candidates (denoted by maxIBCCandNum).

[0354] B16. The method according to any of Solutions B1 to B14, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC merge candidates (denoted by maxIBCMrgNum).

[0355] B17. The method according to any of Solutions B1 to B14, wherein the maximum number of IBC candidates of the first type is the maximum number of IBC Advanced Motion Vector Prediction (AMVP) candidates (denoted by maxIBCAMVPNum).

[0356] B18. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the maximum number of IBC (Intra Block Copy) motion candidates (denoted by maxIBCCandNum) used during the video domain conversion is a function of the maximum number of IBC merge candidates (denoted by maxIBCMrgNum) and the maximum number of IBC AMVP (Advanced Motion vector Prediction) candidates (denoted by maxIBCAMVPNum), and when an IBC mode is applied, samples of the video domain are predicted from other samples in the video picture that correspond to the video domain.

[0357] B19.maxIBCAMVPNum is equal to 2, the method described in solution B18.

[0358] B20. The function returns the larger of its arguments, as described in solution B18 or B19.

[0359] B21. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of Intra Block Copy (IBC) motion candidates (denoted maxIBCCandNum) used during said conversion of the video domain is based on coded mode information of the video domain.

[0360] B22. The method according to Solution B21, wherein maxIBCCandNum is set to the maximum number of IBC merge candidates (denoted by maxIBCMrgNum) based on the video region that is coded in IBC merge mode.

[0361] B23. The method according to solution B21, wherein maxIBCCandNum is set to the maximum number of IBC Advanced Motion Vector Prediction (AMVP) candidates (denoted by maxIBCAMVPNum) based on the video region coded in IBC AMVP mode.

[0362] B24. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the decoded IBC (Intra Block Copy) AMVP (Advanced Motion Vector Prediction) merge index or the decoded IBC merge index is smaller than the maximum number of IBC motion candidates (denoted by maxIBCCandNum).

[0363] B25. The method of solution B24, wherein the decoded IBC AMVP merge index is less than the maximum number of IBC AMVP candidates (denoted by maxIBCAMVPNum).

[0364] B26. The method according to Solution B24, wherein the decoded IBC merge index is less than the maximum number of IBC merge candidates (denoted by maxIBCMrgNum).

[0365] The method according to solution B26, where B27.maxIBCMrgNum=2.

[0366] B28. A method of video processing, comprising: determining, during conversion between a video domain of a video and a bitstream representation of the video, that an IBC (Intra Block Copy) AMVP (Alternative Motion Vector Predictor) candidate index or an IBC merge candidate index cannot identify a block vector candidate in a block vector candidate list; and, based on the determination, using a default prediction block during conversion.

[0367] B29. The method according to Solution B28, wherein each sample of the default prediction block is set to (1<<(BitDepth-1)), where BitDepth is a positive integer.

[0368] B30. The method according to Solution B28, wherein a default block vector is assigned to a default predicted block.

[0369] B31. A method of video processing, comprising: determining, during conversion between a video region of a video and a bitstream representation of the video, that an IBC (Intra Block Copy) AMVP (Alternative Motion Vector Predictor) candidate index or an IBC merge candidate index cannot identify a block vector candidate in a block vector candidate list; and, based on the determination, performing the conversion by treating the video region as having an invalid block vector.

[0370] B32. A method of video processing, comprising: determining that an Intra Block Copy (IBC) Alternative Motion Vector Predictor (AMVP) candidate index or an IBC merge candidate index does not satisfy a condition during conversion between a video domain of a video and a bitstream representation of the video; generating a supplemental Block Vector (BV) candidate list based on the determination; and performing the conversion using the supplemental BV candidate list.

[0371] B33. The method of Solution B32, wherein the condition is that the IBC AMVP candidate index or the IBC merge candidate index is greater than or equal to the maximum number of IBC motion candidates for the video region.

[0372] B34. A method according to Solution B32 or B33, wherein the supplemental BV candidate vector list is generated using the steps of adding one or more HMVP (History-based Motion Vector Prediction) candidates, generating one or more virtual BV candidates from other BV candidates, and adding one or more default candidates.

[0373] B35. The method of solution B34, wherein the steps are performed in sequence.

[0374] B36. The method of solution B34, wherein the steps are performed in an interleaved manner.

[0375] B37. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the maximum number of Intra Block Copy (IBC) Advanced Motion Vector Prediction (AMVP) candidates (denoted maxIBCAMVPNum) is not equal to 2.

[0376] B38. The method according to Solution B37, wherein the bitstream representation excludes a flag indicating a motion vector predictor index and includes an index having a value greater than 1.

[0377] B39. The method of Solution B38, wherein the index is binary coded using unary, truncated unary, fixed length, or exponential-Golomb representation.

[0378] B40. The method according to Solution B38, wherein the bins of the binary bin string of the index are context coded or bypass coded.

[0379] B41. The method according to Solution B37, wherein maxIBCAMVPNum is greater than the maximum number of IBC merge candidates (denoted by maxIBCMrgNum).

[0380] B42. A method of video processing comprising determining that a maximum number of IBC (Intra Block Copy) motion candidates is zero during conversion between video domain and bitstream representations of the video domain, and performing the conversion by generating an IBC block vector candidate list during the conversion based on the determination.

[0381] B43. The method of Solution B42, wherein the transform further includes generating the merge list based on IBC Advanced Motion Vector Prediction (AMVP) mode being enabled for the video domain transform.

[0382] B44. The method of Solution B42, wherein the transform further includes generating a merge list having a length up to the maximum number of IBC AMVP candidates based on IBC merge mode being disabled for the video domain transform.

[0383] B45. The method according to any of Solutions B1 to B44, wherein the conversion comprises generating pixel values ​​in the video domain from the bitstream representation.

[0384] B46. The method according to any of Solutions B1 to B44, wherein the conversion comprises generating a bitstream representation from pixel values ​​in the video domain.

[0385] B47. A video processing device comprising a processor configured to implement the methods described in one or more of Solutions B1 to B46.

[0386] B48. A computer-readable medium storing program code that, when executed, causes a processor to implement a method according to one or more of solutions B1-B46.

[0387] In some embodiments, the following technical solutions can be implemented.

[0388] C1. A method of video processing comprising: determining that during conversion between a video region of a video and a bitstream representation of the video region, use of intra block copy mode for the video region is disabled and use of intra block copy mode for other video regions of the video is enabled; and performing conversion based on the determination, wherein in the intra block copy mode, pixels of the video region are predicted from other pixels in a video picture corresponding to the video region.

[0389] C2. A video region is a video picture, or a slice or tile of a video picture. , or the method according to solution C1 corresponding to a tile group or brick.

[0390] C3. The method according to any of solutions C1-C2, wherein the bitstream representation includes an indication regarding the decision.

[0391] C4. The method of solution C3, wherein the instructions are contained at the picture, slice, tile, tile group, or brick level.

[0392] C5. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the bitstream representation selectively includes instructions for determining when the maximum number of IBC motion candidates used during the conversion of the video domain (denoted by maxIBCCandNum) is equal to the maximum number of merge candidates used during the conversion (denoted by MaxNumMergeCand).

[0393] C6. The method according to solution C5, wherein maxIBCCandNum is equal to zero and due to maxIBCCardNum being equal to zero, the bitstream representation excludes signaling of certain IBC information.

[0394] C7. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the bitstream representation includes motion vector differential-related syntax elements for intra block copy alternative motion vector predictors used during the video domain conversion and dependent on the value of a maximum number of intra block copy motion candidates (denoted by maxIBCCandNum) that satisfy a condition, and wherein in an intra block copy mode, pixels of the video domain are predicted from other pixels in a video picture that correspond to the video domain.

[0395] C8. The method according to solution C7, wherein the condition includes that maxIBCCandNum is equal to the maximum number of merge candidates used during the transformation (denoted by MaxNumMergeCand).

[0396] C9. The method according to Solution C8, wherein maxIBCCandNum is equal to 0 and, due to maxIBCCandNum being equal to 0, signaling of motion vector predictor indices for intra block copy alternative motion vector predictors is omitted from the bitstream representation.

[0397] C10. The method of Solution C8, wherein maxIBCCandNum is equal to 0 and, due to maxIBCCandNum being equal to 0, signaling of motion vector differentials for intra block copy alternative motion vector predictors is omitted from the bitstream representation.

[0398] C11. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of intra block copy motion candidates (indicated by maxIBCCandNum) used during the conversion of the video domain is signaled in the bitstream representation independently of a maximum number of merge candidates (indicated by MaxNumMergeCand) used during the conversion, and the intra block copy corresponds to a mode in which pixels of the video domain are predicted from other pixels in a video picture corresponding to the video domain.

[0399] C12. The method according to solution C11, wherein MaxIBCCandNum is greater than zero due to intra block copying being enabled for video domain transformations.

[0400] C13. The method according to any of Solutions C11 to C12, wherein maxIBCCandNum is signaled in the bitstream representation by predictive coding using another value.

[0401] C14. Another value is MaxNumMergeCand, as described in solution C13.

[0402] C15. Another value is the constant K, as in solution C13.

[0403] C16. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein the maximum number of IBC (Intra Block Copy) motion candidates (denoted by maxIBCCandNum) used during the video domain conversion depends on coded mode information of the video domain, and the IBC corresponds to a mode in which pixels of the video domain are predicted from other pixels in a video picture corresponding to the video domain.

[0404] C17. The method of solution C16, wherein if the block is coded in IBC merge mode, maxIBCCandNum is set to the maximum number of IBC merge candidates (denoted by maxIBCMrgNum).

[0405] C18. The method of solution C16, wherein if the block is coded in an IBC Alternative Motion Vector Prediction (AMVP) candidate mode, maxIBCCandNum is set to the maximum number of IBC AMVP numbers (denoted by maxIBCAMVPNum).

[0406] C19.maxIBCMrgNum is signaled in the slice header, solution C17 method.

[0407] The method according to solution C17, where C20.maxIBCMrgNum is set equal to the maximum number of allowed non-IBC translational merge candidates.

[0408] C21.maxIBCAMVPNum is set equal to 2, solution C18 method.

[0409] C22. A method of video processing, comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of IBC (Intra Block Copy) motion candidates (denoted by maxIBCCandNum) used during the video domain conversion is a first function of a maximum number of IBC merge candidates (denoted by maxIBCMrgNum) and a maximum number of IBC alternative motion vector prediction candidates (denoted by maxIBCAMVPNum), and IBC corresponds to a mode in which pixels of the video domain are predicted from other pixels in a video picture corresponding to the video domain.

[0410] C23.maxIBCAMVPNum is equal to 2, the method described in solution C22.

[0411] C24. The method according to solutions C22 to C23, wherein during transformation, the decoded IBC alternative motion vector predictor index is less than maxIBCAMVPNum.

[0412] C25. A method of video processing comprising: determining, during conversion between a video domain of a video and a bitstream representation of the video, that an intra block copy alternative motion vector predictor index or an intra block copy merge candidate index cannot identify a block vector candidate in a block vector candidate list; and using a default prediction block during the conversion based on the determination.

[0413] C26. A method of video processing, comprising: during conversion between a video region of a video and a bitstream representation of the video, determining that an intra block copy alternative motion vector predictor index or an intra block copy merge candidate index fails to identify a block vector candidate in a block vector candidate list; and, based on the determination, performing the conversion by treating the video region as having an invalid block vector.

[0414] C27. A method of video processing, comprising: determining that an intra block copy alternative motion vector predictor index or an intra block copy merge candidate index does not satisfy a condition during conversion between a video domain of a video and a bitstream representation of the video; generating a supplemental BV (Block Vector) candidate list based on the determination; and performing the conversion using the supplemental block vector candidate list.

[0415] C28. The method according to Solution C27, wherein the condition includes that the intra block copy alternative motion vector predictor index or the intra block copy merge candidate index is less than the maximum number of intra block copy candidates for the video region.

[0416] C29. A method according to any of solutions C27 to C28, wherein the supplemental BV candidate vector list is generated using the steps of generating history-based motion vector predictor candidates, generating virtual BV candidates from other BV candidates, and adding default candidates.

[0417] C30. The method of solution C29, wherein the steps are performed in order.

[0418] C31. The method according to solution C29, wherein the steps are performed in an interleaved manner.

[0419] C32. A method of video processing comprising performing a conversion between a video domain of a video and a bitstream representation of the video, wherein a maximum number of IBC alternative motion vector prediction candidates (denoted maxIBCAMVPNum) is not equal to 2, and wherein IBC corresponds to a mode in which pixels of the video domain are predicted from other pixels in the video picture that correspond to the video domain.

[0420] C33. The method according to solution C32, wherein the bitstream representation excludes a first flag indicating a motion vector predictor index and includes an index having a value greater than 1.

[0421] C34. The method according to any of Solutions C32-C33, wherein the index is coded using unary / truncated unary / fixed length / exponential-golomb / binarized in other binary representations.

[0422] C35. A method of video processing, comprising: determining that a maximum number of Intra Block Copy (IBC) candidates is zero during conversion between a video domain and a bitstream representation of the video domain; and performing the conversion by generating an IBC block vector candidate list during the conversion based on the determination, wherein IBC corresponds to a mode in which pixels of the video domain are predicted from other pixels in a video picture corresponding to the video domain.

[0423] C36. The method of Solution C35, wherein the transform further comprises generating the merge list by assuming that an IBC alternative motion vector predictor mode is enabled for the video domain transform.

[0424] C37. The method according to Solution C34, wherein the transform further comprises generating a merge list having a length up to the maximum number of IBC advanced motion vector predictor candidates if IBC merge mode is disabled for the video domain transform.

[0425] C38. The method according to any of Solutions C1 to C37, wherein the conversion comprises generating pixel values ​​in the video domain from the bitstream representation.

[0426] C39. The method according to any of Solutions C1 to C37, wherein the conversion comprises generating a bitstream representation from pixel values ​​in the video domain.

[0427] C40. A video processing device comprising a processor configured to implement the methods described in one or more of solutions C1 to C39.

[0428] C41. A computer-readable medium storing program code that, when executed, causes a processor to implement a method according to one or more of solutions C1-C39.

[0429] Implementations of the disclosed and other solutions, examples, embodiments, modules, and functional operations described herein, including the structures disclosed herein and their structural equivalents, may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, or in one or more combinations thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for implementation by or controlling the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" includes all apparatuses, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may include code that creates an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, for example, a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiving device.

[0430] A computer program (also called a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be recorded as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), may be stored in a single file dedicated to the program, or may be stored in multiple coordinating files (e.g., files containing one or more modules, subprograms, or portions of code). A computer program can be deployed to run on one computer located at a single site or on multiple computers distributed across multiple sites and interconnected by a communications network.

[0431] The processing and logic flows described herein may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processing and logic flows may also be performed by, and devices may be implemented as, special purpose logic circuitry, such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).

[0432] Processors suitable for executing a computer program include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will include one or more mass storage devices, e.g., magnetic, magneto-optical, or optical disks, for storing data, or be operatively coupled to receive data from or transfer data to these mass storage devices. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, EPROMs, EEPROMs, flash storage devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and semiconductor storage devices such as CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0433] While this patent specification contains many details, these should not be construed as limiting the scope of any subject matter or the scope of the claims, but rather as descriptions of features that may be specific to particular embodiments of a particular technology. Certain features described in this patent specification in the context of separate embodiments may also be implemented in combination in a single example. Conversely, various features described in the context of a single example may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as acting in a particular combination and initially claimed as such, one or more features from a claimed combination may, in some cases, be extracted from the combination, and the claimed combination may be directed to subcombinations or variations of the subcombination.

[0434] Similarly, although operations are shown in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order or sequential order shown, or that all of the operations shown be performed, to achieve desired results. Also, the separation of various system components in the examples described in this patent specification should not be understood as requiring such separation in all embodiments.

[0435] Only a few implementations and examples have been described; other embodiments, extensions and variations are possible based on the content described and illustrated in this patent document.

Claims

1. 1. A method of video processing, comprising: determining whether a merge mode is applied to a video block of the video domain and whether an Intra Block Copy (IBC) mode is applied to the video block for conversion between a video domain of a video and a bitstream of the video; if the merge mode is not applied to the video block and the IBC mode is applied to the video block, determining whether to include at least one syntax element for the video block in the bitstream based on a maximum number of IBC candidates used during the transformation of the video domain, the maximum number of IBC candidates being decoupled from a maximum number of regular merge candidates for a block; and When the IBC mode is applied, samples of the video block are predicted from other samples in a video picture that has the video block, and the maximum number of IBC candidates is derived to be equal to (6 - an indication signaled in the bitstream).

2. the at least one syntax element includes a vector predictor index of list 0; The method of claim 1 , wherein list 0 is the reference picture list for the video block with a reference picture list index of 0.

3. based on the maximum number of IBC candidates being greater than K, the bitstream includes the at least one syntax element; The method of claim 1 , wherein the bitstream excludes the at least one syntax element based on the maximum number of IBC candidates being less than or equal to K.

4. A vector predictor index included in the at least one syntax element is removed; The method of claim 3 , wherein the vector predictor index is assumed to be zero.

5. 5. The method of claim 3, wherein K=1.

6. the maximum number of IBC candidates is equal to zero; The method of claim 1 , wherein the IBC mode is disabled.

7. The method of claim 1 , wherein if the IBC mode is enabled for a slice having the video block, the bitstream satisfies that the maximum number of IBC candidates is greater than 0.

8. The method of claim 1 , wherein the maximum number of IBC candidates is a maximum number of IBC merge candidates (denoted maxIBCMrgNum).

9. The method of claim 1 , wherein the converting comprises decoding the video block from the bitstream.

10. The method of claim 1 , wherein the transforming comprises encoding the video blocks into the bitstream.

11. 1. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions, The instructions, when executed by the processor, cause the processor to: determining whether a merge mode is applied to a video block of the video domain and whether an Intra Block Copy (IBC) mode is applied to the video block for conversion between a video domain of a video and a bitstream of the video; if the merge mode is not applied to the video block and the IBC mode is applied to the video block, determining whether to include at least one syntax element for the video block in the bitstream based on a maximum number of IBC candidates used during the transformation of the video domain, the maximum number of IBC candidates being decoupled from a maximum number of regular merge candidates for a block; Execute When the IBC mode is applied, samples of the video block are predicted from other samples in a video picture that includes the video block, and the maximum number of IBC candidates is derived to be equal to (6 - an indication signaled in the bitstream).

12. The processor determining whether a merge mode is applied to a video block of the video domain and whether an Intra Block Copy (IBC) mode is applied to the video block for conversion between a video domain of a video and a bitstream of the video; if the merge mode is not applied to the video block and the IBC mode is applied to the video block, determining whether to include at least one syntax element for the video block in the bitstream based on a maximum number of IBC candidates used during the transformation of the video domain, the maximum number of IBC candidates being decoupled from a maximum number of regular merge candidates for a block; Execute when the IBC mode is applied, samples of the video block are predicted from other samples in a video picture that has the video block, and the maximum number of IBC candidates is derived to be equal to (6 minus an indication signaled in the bitstream). A non-transitory computer-readable storage medium storing instructions.

13. 1. A method for storing a video bitstream, comprising: The method comprises: determining whether a merge mode is applied to a video block of the video domain and whether an Intra Block Copy (IBC) mode is applied to the video block for conversion between a video domain of a video and a bitstream of the video; if the merge mode is not applied to the video block and the IBC mode is applied to the video block, determining whether to include at least one syntax element for the video block in the bitstream based on a maximum number of IBC candidates used during the transformation of the video domain, the maximum number of IBC candidates being decoupled from a maximum number of regular merge candidates for a block; generating the bitstream based on the video region of the video; storing the bitstream on a non-transitory computer-readable recording medium; and When the IBC mode is applied, samples of the video block are predicted from other samples in a video picture that has the video block, and the maximum number of IBC candidates is derived to be equal to (6 - an indication signaled in the bitstream).

Citation Information

Patent Citations

  • Method, video decoder, and computer program for video decoding, and method for video encoding - Patents.com

    JP2022516062A

  • JPP7332721B

  • Method and apparatus for signaling predictor candidate list size

    WO2020227204A1