Intra block copy with triangular partitioning
By dividing video blocks into wedge or triangle shapes and combining different encoding/decoding modes and palette encoding/decoding modes, the problem of low encoding/decoding efficiency of non-rectangular segmented video blocks in existing technologies is solved, achieving more efficient video encoding/decoding effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-02
- Publication Date
- 2026-03-27
AI Technical Summary
Existing video codec standards have room for improvement in efficiency and quality when dealing with non-rectangular segmented video blocks, especially in the case of triangular or wedge-shaped segments, where existing technologies struggle to effectively utilize the non-rectangular characteristics of the blocks for efficient encoding and decoding.
By employing a wedge or triangle segmentation method for video blocks, different encoding/decoding modes or palette encoding/decoding modes are used by determining the transformation method for each sub-region, and by combining intra-frame encoding/decoding modes and motion vector prediction, the encoding/decoding efficiency and quality of video blocks are improved.
It improves the encoding and decoding efficiency and quality of video blocks, especially in the case of triangular or wedge segmentation, achieving more efficient encoding and decoding performance.
Smart Images

Figure CN113924778B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] In accordance with the applicable Patent Law and / or the Paris Convention, this application promptly claims priority and interest in International Patent Application No. PCT / CN2019 / 089795, filed on June 3, 2019. The entire disclosure of International Patent Application No. PCT / CN2019 / 089795 is incorporated herein by reference as a part of this application disclosure. Technical Field
[0003] This article covers video and image encoding and decoding technologies. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] The disclosed techniques can be used in video or image decoder or encoder embodiments where non-rectangular segmentation is used to divide a block into multiple sub-blocks. For example, non-rectangular segmentation may include wedge segmentation or triangular segmentation.
[0006] In one exemplary aspect, a method for processing video is disclosed. The method includes: during the conversion of a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining (2402) that the conversion of each sub-region is based on an encoding / decoding method using a current image as a reference image; based on the determination, determining (2404) motion vectors for IBC encoding / decoding of the multiple sub-regions; and performing (2406) the conversion using two different motion vectors.
[0007] In another exemplary aspect, a method for processing video is disclosed. The method includes: during a conversion between a current video block segmented into at least two sub-regions based on wedge or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on an encoding / decoding method using the current image as a reference image, and the conversion of the other sub-region is based on an alternative encoding / decoding mode different from the IBC encoding / decoding mode; and performing the conversion based on the determination.
[0008] In another exemplary aspect, a method for processing video is disclosed. The method includes: during a conversion between a current video block divided into at least two sub-regions based on wedge or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on an intra-frame encoding / decoding mode, and the conversion of the other sub-region is based on an alternative encoding / decoding mode different from the intra-frame encoding / decoding mode; and performing the conversion based on the determination.
[0009] In another exemplary aspect, a method for processing video is disclosed. The method includes: during a conversion between a current video block segmented into sub-regions based on wedge or triangle segmentation and a bitstream representation of the current video block, determining to encode / decode the sub-regions based on a palette encoding / decoding mode, wherein the palettes of at least two sub-regions are different from each other; and performing a conversion based on the determination.
[0010] In another exemplary aspect, a method for processing video is disclosed. The method includes: during a conversion between a current video block divided into at least two sub-regions based on wedge or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on a palette codec mode, and the conversion of the other sub-region is based on an alternative codec mode different from the palette codec mode; and performing the conversion based on the determination.
[0011] In another exemplary aspect, a method for processing video is disclosed. The method includes performing a conversion between a block of a video component and a bitstream representation of that block using the method described in any of the preceding clauses, such that a conversion between corresponding blocks of another component of the video is performed using a different encoding / decoding technique than that used for the conversion of that block.
[0012] In another exemplary aspect, a method for processing video is disclosed. The method includes: during a conversion between a current video block segmented into at least two sub-regions based on wedge or triangle segmentation and a bitstream representation of the current video block, determining a final prediction for the current video block using predictions for each sub-region; and performing the conversion based on the determination.
[0013] In another exemplary aspect, a method for processing video is disclosed. The method includes: storing motion information of one sub-region as motion information of the current video block during a conversion between a current video block segmented into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block; and using the motion information to convert the current video block or a subsequent video block.
[0014] In another exemplary aspect, a method for processing video is disclosed. The method includes: a conversion between a block of video and a bitstream representation of the block; determining to apply at least one of an intra-block copy (IBC) mode, an intra-frame mode, an inter-frame mode, and a palette mode to a plurality of sub-regions of the block, wherein the block is divided into two or more triangular or wedge-shaped sub-regions; and performing the conversion based on the determination.
[0015] In another exemplary aspect, the above method may be implemented by a video decoder device including a processor.
[0016] In another exemplary aspect, the above method may be implemented by a video encoder device including a processor.
[0017] In yet another exemplary aspect, these methods may be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.
[0018] These and other aspects are further described in this article. Attached Figure Description
[0019] Figure 1 The derivation process for constructing the merge candidate list is shown.
[0020] Figure 2 An example of the location of the spatial merge candidate is shown.
[0021] Figure 3 An example of candidate pairs considered for redundancy checks of spatial merge candidates is shown.
[0022] Figure 4 Example locations of the second PU for N×2N and 2N×N segmentation are shown.
[0023] Figure 5 An illustrated example of motion vector scaling for temporal merge candidates is shown.
[0024] Figure 6 An example of the candidate positions for temporal merge candidates C0 and C1 is shown.
[0025] Figure 7 An example of combined bidirectional prediction merge candidates is shown.
[0026] Figure 8 An example of the derivation process for motion vector prediction candidates is shown.
[0027] Figure 9 An example illustration of motion vector scaling for spatial motion vector candidates is shown.
[0028] Figure 10Simplified affine motion models for a 4-parameter affine mode (left) and a 6-parameter affine model (right) are shown.
[0029] Figure 11 An example of the affine motion vector field for each sub-block is shown.
[0030] Figure 12 Example candidate positions for the affine merge pattern are shown.
[0031] Figure 13 An example of the modified merge list construction process is shown.
[0032] Figure 14 An example of inter-frame prediction based on triangle segmentation is shown.
[0033] Figure 15 An example of the final motion vector representation (UMVE) search process is shown.
[0034] Figure 16 An example of a UMVE search point is shown.
[0035] Figure 17 An example of MVD(0,1) that is a mirror image between list 0 and list 1 in DMVR is shown.
[0036] Figure 18 The MV that can be checked in a single iteration is shown.
[0037] Figure 19 This is a diagram of the Intra-Block Copy (IBC) mode.
[0038] Figure 20 An example of dividing a block into four sub-regions is shown.
[0039] Figure 21 An example of dividing a block into two wedge-shaped sub-regions is shown.
[0040] Figure 22 An example of dividing a block into two wedge-shaped sub-regions is shown.
[0041] Figure 23 This is a block diagram of an example video processing device.
[0042] Figure 24 This is a flowchart illustrating an example of a video processing method.
[0043] Figure 25 This is a flowchart illustrating an example of a video processing method. Detailed Implementation
[0044] This article provides a variety of techniques that decoders of image or video bitstreams can use to improve the quality of decompressing or decoding digital video or images. For the sake of brevity, the term "video" is used throughout this article to include both sequences of images (conventionally referred to as video) and individual images. Furthermore, video encoders can implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0045] Chapter headings are used in this document for ease of understanding and not to limit the embodiments and techniques to the corresponding chapters. Accordingly, embodiments from one chapter may be combined with embodiments from other chapters.
[0046] 1. Brief Summary
[0047] This patent document relates to video codec technology. Specifically, this patent document relates to a palette mode in video codec. It can be applied to existing video codec standards, such as HEVC, or pending standards (Multi-Functional Video Codec). It may also be applicable to future video codec standards or video codecs.
[0048] 2. Preliminary Discussion
[0049] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Experts Team (JVET) was created between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to research a VVC standard with a target bitrate reduction of 50% compared to HEVC.
[0050] 2.1 Inter-frame prediction in HEVC / H.265
[0051] For an inter-frame codec unit (CU), depending on the segmentation mode, it can be encoded and decoded into having one prediction unit (PU) or two PUs. Each inter-frame prediction PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be signaled using `inter_pred_idc`. The motion vectors can be explicitly encoded and decoded as Δ relative to the predictor.
[0052] When encoding and decoding a CU using the skip mode, a PU is associated with that CU and there are no significant residual coefficients, no encoded motion vector Δ, or no reference picture index. The merge mode is specified such that the motion parameters of the current PU are obtained from neighboring PUs, including both spatial and temporal candidates. The merge mode can be applied to any inter-frame prediction PU, not just those in the skip mode. An alternative to the merge mode is explicit transmission of motion parameters, where explicit signaling is given for each PU based on its motion vector (more precisely, the motion vector difference (MVD) relative to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the reference picture list usage. Such a mode is referred to in this disclosure as Advanced Motion Vector Prediction (AMVP).
[0053] When the signaling indicates that one of the two reference image lists will be used, the PU is generated from a sample block. This practice is called "one-way prediction". One-way prediction is available for both P-strips and B-strips.
[0054] When the signaling indicates that both reference image lists are to be used, the PU is generated from two sample blocks.
[0055] This approach is called "bidirectional forecasting". Bidirectional forecasting is only available for the B band.
[0056] The following section provides details about these inter-frame prediction modes specified in HEVC. The description will begin with the merge mode.
[0057] 2.1.1 List of Reference Images
[0058] In HEVC, the term inter-frame prediction is used to refer to predictions derived from data elements (e.g., sample values or motion vectors) of a reference image (rather than the currently decoded image). Similar to H.264 / AVC, an image can be predicted from multiple reference images. The reference images used for inter-frame prediction are organized into one or more reference image lists. A reference index identifies which reference image in the list should be used to create the prediction signal.
[0059] A single list of reference images, List 0, is used for stripe P, and two lists of reference images, List 0 and List 1, are used for stripe B. It should be noted that the reference images included in List 0 / 1 can be derived from past and future images, referring to the capture / display order.
[0060] 2.1.2 Merge Mode
[0061] 2.1.2.1 Derivation of candidate merge modes
[0062] When predicting PU using merge mode, the indices pointing to entries in the merge candidate list are parsed from the bitstream, and these indices are used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:
[0063] Step 1: Initial Candidate Derivation
[0064] Step 1.1: Spatial Candidate Derivation
[0065] Step 1.2: Redundancy check of airspace candidates
[0066] Step 1.3: Time-domain candidate derivation
[0067] Step 2: Additional candidate insertion
[0068] Step 2.1: Creation of bidirectional prediction candidates
[0069] Step 2.2: Insertion of zero-motion candidates
[0070] Still Figure 1 These steps are illustrated schematically. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since a constant number of candidates are taken for each PU at the decoder, additional candidates are generated if the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, truncated univariate binarization (TU) is used to index and encode the best merge candidate. If the size of the CU is equal to 8, then all PUs of the current CU share a single merge candidate list, which is equivalent to the merge candidate list of a 2N×2N prediction unit.
[0071] The operations associated with the foregoing steps will be described in detail below.
[0072] 2.1.2.2 Derivation of Airspace Candidates
[0073] In the derivation of the spatial merge candidate, located in Figure 2 At most four merge candidates are selected from the candidates at the indicated positions. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or is subject to intra-frame encoding / decoding. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only those with... Figure 3 The arrows link pairs of motion information, and a candidate is added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information. Another source of repetitive motion information is a "second PU" associated with a segmentation other than 2N×2N. As an example, Figure 4 The second PUs are shown for N×2N and 2N×N scenarios, respectively. When the current PU is divided into N×2N segments, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would cause the two prediction units to have the same motion information, which is redundant for having only one PU within the encoder-decoder unit. Similarly, when the current PU is divided into 2N×N segments, position B1 is not considered.
[0074] 2.1.2.3 Time-domain candidate derivation
[0075] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaling motion vector is derived based on the juxtaposition PU of the image that belongs to the given list of reference images and has the smallest POC difference with the current image. The list of reference images to be used for the derivation of the juxtaposition PU is explicitly signaled within the strip header. The scaling motion vector used for the temporal merge candidate is as follows: Figure 5 The image, shown by the dashed lines, is obtained by scaling the motion vectors of the juxtaposed PU using the POC differences (i.e., tb and td). Here, tb is defined as the POC difference between the current image and its reference image, and td is defined as the POC difference between the juxtaposed image and its reference image. The reference image index for the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For B-strips, two motion vectors are obtained (one for reference image list 0 and the other for reference image list 1), and they are combined to generate bidirectional predictive merge candidates.
[0076] In the juxtaposed PU(Y) belonging to the reference frame, the position of the temporal candidate is selected between candidate C0 and C1, such as... Figure 6 As shown. If the PU at position C0 is unavailable, is intra-frame encoded or decoded, or is outside the current codec tree unit (CTU, also known as LCU, maximum codec unit) row, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.
[0077] 2.1.2.4 Additional Candidate Insertion
[0078] In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated using both spatial and temporal merge candidates. Combined bidirectional prediction merge candidates are only used for B-strips. They are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of the other. If these two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. As an example, Figure 7 This illustrates the case where two candidates from the original list (left side) (having mvL0 and refIdxL0 or mvL1 and refIdxL1) are used to create combined bidirectional predictive merge candidates that are added to the final list (right side). There are many rules regarding the combinations considered for generating these additional merge candidates.
[0079] Zero-motion candidates are inserted to populate the remaining entries in the merge candidate list, thus hitting the MaxNumMergeCand capacity. These candidates have zero spatial displacement and a reference image index that starts at zero and increases each time a new zero-motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.
[0080] 2.1.3 AMVP
[0081] AMVP utilizes the spatiotemporal correlation between motion vectors and adjacent PUs, which is used for the explicit transmission of motion parameters. For each reference image list, the motion vector candidate list is constructed as follows: first, the availability of the temporally adjacent PU positions on the left and top sides is checked, redundant candidates are removed, and zero vectors are added, thus ensuring a constant length for the candidate list. Then, the encoder can select the best predictor from the candidate list and transmit the corresponding index indicating the selected candidate. Similar to merge index signaling, a truncated unary code is used to encode the index of the best motion vector candidate. In this case, the maximum value to be encoded is 2 (reference...). Figure 8The following sections will provide details on the derivation process of the motion vector prediction candidates.
[0082] 2.1.3.1 Derivation of AMVP Candidates
[0083] Figure 8 The derivation process of motion vector prediction candidates is summarized.
[0084] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, the final result is located in... Figure 2 Based on the motion vectors of each PU at the five different locations shown, two motion vector candidates are derived.
[0085] For temporal motion vector candidate derivation, a motion vector candidate is selected from two candidates derived based on two different juxtaposition positions. After creating a first list of spatiotemporal candidates, duplicate motion vector candidates in this list are removed. If the number of possible candidates is greater than two, motion vector candidates whose reference image index in the associated reference image list is greater than 1 are removed from the list. If the number of spatiotemporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0086] 2.1.3.2 Candidate Spatial Motion Vectors
[0087] In the derivation of the spatial motion vector candidates, at most two candidates are considered from five possible candidates. These five possible candidates are selected from those located in the spatial domain, such as... Figure 2 The PUs at the indicated positions are derived, and these positions are identical to those in the motion merge. The derivation order for the left side of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order for the upper side of the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. For each side, there are therefore four cases that can be used as motion vector candidates, two of which do not require spatial scaling, and two of which do. These four different cases are summarized below:
[0088] • No spatial scaling
[0089] –(1) List of identical reference images and index of identical reference images (identical POCs)
[0090] –(2) Different lists of reference images, but with the same reference image (same POC)
[0091] • Spatial scaling
[0092] –(3) Same list of reference images, but different reference image indices (different POCs)
[0093] –(4) List of different reference images and different reference images (different POCs)
[0094] First, check for cases without spatial scaling, then proceed with spatial scaling. When the POC differs between the reference image of an adjacent PU and the reference image of the current PU, spatial scaling is considered regardless of the reference image list. If all candidate PUs on the left are unavailable or intra-frame encoded / decoded, scaling for the upper motion vector is allowed, thus facilitating parallel derivation of the left and upper MV candidates. Otherwise, spatial scaling for the upper motion vector is not allowed.
[0095] During spatial scaling, the motion vectors of adjacent PUs are scaled in a manner similar to temporal scaling, such as... Figure 9 As shown. The main difference is that the current PU's reference image list and index are given as input; the actual scaling process is the same as temporal scaling.
[0096] 2.1.3.3 Candidate Motion Vectors in the Time Domain
[0097] Aside from the derivation of the reference image index, the entire derivation process of the temporal merge candidate is identical to the derivation of the spatial motion vector candidate (see [link to derivation]). Figure 6 (Same as above.) The reference image index signaling is sent to the decoder.
[0098] 2.2 Inter-frame prediction methods in VVC
[0099] Several new codec tools have been developed for improving inter-frame prediction, such as Adaptive Motion Vector Resolution (AMVR) for Signaling Notification MVD, Merge with Motion Vector Difference (MMVD), Triangle Prediction Mode (TPM), Combined Intra-Inter-Frame Prediction (CIIP), Advanced TMVP (ATMVP, also known as SbTMVP), Affine Prediction Mode, Universal Bidirectional Prediction (GBI), Decoder-Side Motion Vector Refinement (DMVR), and Bidirectional Optical Flow (BIO, also known as BDOF).
[0100] VVC supports three different merge list construction processes:
[0101] 1) Sub-block merge candidate list: Includes ATMVP and affine merge candidates. A merge list construction process is shared for both affine and ATMVP modes. ATMVP and affine merge candidates can be added sequentially. The sub-block merge list size is signaled in the stripe header, with a maximum value of 5.
[0102] 2) Regular merge list: For the remaining codec blocks, a common merge list construction process is used. Here, spatial / temporal / HMVP, paired combined bidirectional prediction merge candidates, and zero motion candidates can be inserted sequentially. The size of the regular merge list is indicated by signaling in the stripe header, with a maximum value of 6. MMVD, TPM, and CIIP rely on the regular merge list.
[0103] 3) IBC merge list: Completed in a similar manner to the regular merge list.
[0104] Similarly, VVC supports three lists of AMVPs:
[0105] 1) Affine AMVP Candidate List
[0106] 2) Regular AMVP Candidate List
[0107] 3) IBC AMVP Candidate List: The same construction process as the IBC merge list.
[0108] 2.2.1 Codec Block Structure in VVC
[0109] In VVC, a quadtree / binary tree / ternary tree (QT / BT / TT) structure is used to divide the image into square or rectangular blocks.
[0110] In addition to QT / BT / TT, VVC also employs a separate tree (also known as a dual codec tree) for I-frames. This separate tree allows for separate signaling of the codec block structure for the luma and chroma components.
[0111] In addition, except for blocks encoded using certain encoding / decoding methods (e.g., intra-frame sub-segmentation prediction, where PU equals TU but is less than CU, and sub-block transformation for inter-frame encoded / decoded blocks, where PU equals CU but TU is less than PU), CU is set to equal PU and TU.
[0112] 2.2.2 Affine Prediction Mode
[0113] In HEVC, only a translational motion model is applied for motion compensation prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, simplified affine transformation motion compensation prediction is applied using 4-parameter and 6-parameter affine models. Figure 10 As shown, the affine motion field is described by the two control point motion vectors (CPMV) of the 4-parameter affine model and the three CPMV blocks of the 6-parameter affine model.
[0114] The motion vector field (MVF) of the block is described by the following equations, where the 4-parameter affine model (where the 4 parameters are defined as variables a, b, e, and f) is in equation (1), and the 6-parameter affine model (where the 6 parameters are defined as a, b, c, d, e, and f) is in equation (2):
[0115]
[0116]
[0117] Among them, (mv h 0,mv h 0) is the motion vector of the top left control point, (mv h 1,mv h 1) is the motion vector of the upper right control point, (mv h 2,mv h 2) is the motion vector of the lower left control point. All three motion vectors are called control point motion vectors (CPMV). (x, y) represents the coordinates of the point relative to the upper left sample point within the current block. (mv) h (x,y),mv v (x, y) is the motion vector derived for the sample point located at (x, y). The Cp motion vector can be known through signaling (e.g., in affine AMVP mode) or derived on the fly (e.g., in affine merge mode). w and h are the width and height of the current block. In practice, this division is performed by right shifting along with a rounding operation. In VTM, the representative point is defined as the center position of the sub-block. For example, if the coordinates of the top-left corner of the sub-block relative to the top-left sample point within the current block are (xs, ys), then the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (e.g., 4x4 in VTM), the motion vector of the entire sub-block is derived using this representative point.
[0118] To further simplify motion compensation prediction, a sub-block-based affine transformation prediction is applied. To derive the motion vector for each M×N (where both M and N are set to 4 in the current VVC) sub-block, the motion vector of the center sample point of each sub-block is calculated according to equations (1) and (2). Figure 11 (As shown), and is rounded to achieve a fractional precision of 1 / 16. Then, a motion-compensated interpolation filter for 1 / 16 pixels is applied to generate a prediction for each sub-block using the derived motion vectors. The 1 / 16 pixel interpolation filter is introduced by an affine pattern.
[0119] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the normal motion vector.
[0120] 2.2.3 MERGE for the entire block
[0121] 2.2.3.1 Constructing the merge list in the regular merge pattern
[0122] 2.2.3.1.1 History-Based Motion Vector Prediction (HMVP)
[0123] Unlike the merge list design, VVC uses a history-based motion vector prediction (HMVP) method.
[0124] In HMVP, motion information from previously encoded / decoded blocks is stored. Motion information from previously encoded / decoded blocks is defined as HMVP candidates. Multiple HMVP candidates are stored in a table called the HMVP table, which is maintained during the encoding / decoding process. The HMVP table is cleared when encoding / decoding a new slice / LCU line / strip begins. Whenever there is an inter-frame encoded / decoded block or a non-sub-block non-TPM mode, the associated motion information is added to the last entry of the table as a new HMVP candidate. Figure 12 The overall encoding and decoding process is illustrated in the diagram.
[0125] 2.2.3.1.2 Standard merge list construction process
[0126] The construction of a regular merge list (for translational motion) can be summarized according to the following steps:
[0127] Step 1: Derivation of Airspace Candidates
[0128] Step 2: Insertion of HMVP candidates
[0129] Step 3: Insertion of Pairwise Average Candidates
[0130] Step 4: Default motion candidates
[0131] HMVP candidates can be used in both the AMVP and merge candidate list construction processes. Figure 13 The modified merge candidate list construction process is illustrated (highlighted in blue). When the merge candidate list is incomplete after inserting TMVP candidates, HMVP candidates stored in the HVMP table can be used to populate the merge candidate list. Considering that a block is generally more relevant to its nearest neighbor in terms of motion information, HMVP candidates in the table are inserted in descending order of their indexes. The last entry in the table is inserted into the list first, and the last entry is added last. Similarly, redundancy elimination is applied to the HMVP candidates. Once the total number of available merge candidates reaches the maximum number of merge candidates allowed for signaling notification, the merge candidate list construction process terminates.
[0132] It should be noted that all spatial / temporal / HMVP candidates should be encoded and decoded using non-IBC mode. Otherwise, they are not allowed to be added to the regular merge candidate list.
[0133] The HMVP table contains up to 5 regular motion candidates, each of which is unique.
[0134] 2.2.3.2 Triangle Prediction Model (TPM)
[0135] In VTM4, triangular splitting mode is supported for inter-frame prediction. Triangular splitting mode is only applied to CUs of 8x8 size or larger that are encoded / decoded in merge mode, not MMVD or CIIP mode. For CUs meeting these conditions, signaling informs CU-level flags indicating whether triangular splitting mode is applied.
[0136] When using this mode, such as Figure 13 As shown, the CU is uniformly divided into two triangular segments using either diagonal or anti-diagonal partitioning. Each triangular segment in the CU utilizes its own motion for inter-frame prediction; for each segment, only unidirectional prediction is allowed, i.e., each segment has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, each CU requires only two motion-compensated predictions.
[0137] If a CU-level flag indicates that the current CU is encoded and decoded using the triangle segmentation mode, further signaling indicates a flag indicating the direction of the triangle segmentation (diagonal or anti-diagonal) and two merge indices (each for one segmentation). After predicting each triangle segmentation, a blending process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the prediction signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangle segmentation mode is stored in each 4x4 cell.
[0138] The regular merge candidate list is reused for triangle segmentation merge prediction without additional motion vector pruning. For each merge candidate in the regular merge candidate list, exactly one of its L0 or L1 motion vectors is used for triangle prediction. Furthermore, the order in which the L0 and L1 motion vectors are selected is based on the parity of their merge indices. Using this scheme, the regular merge list can be used directly.
[0139] 2.2.3.3 MMVD
[0140] The final motion vector representation (UMVE, also known as MMVD) will be introduced. UMVE is used in conjunction with the proposed motion vector representation method in skip or merge modes.
[0141] UMVE reuses the same merge candidates as those included in the regular merge candidate list in VVC. Among these merge candidates, a base candidate can be selected and further expanded using the proposed motion vector representation method.
[0142] UMVE provides a new method for representing motion vector difference (MVD), which uses the starting point, motion amplitude, and motion direction to represent MVD.
[0143] The proposed technique uses the merge candidate list as is. However, it only considers candidates of the default merge type (MRG_TYPE_DEFAULT_N) for UMVE expansion.
[0144] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as described below.
[0145] Table 1. Basic Candidate IDX
[0146] Basic candidate IDX 0 1 2 3 Nth MVP 1st MVP 2nd MVP 3rd MVP 4th MVP
[0147] If the number of basic candidates is equal to 1, then no signaling notification will be sent to the basic candidate IDX.
[0148] The distance index is motion amplitude information. The distance index indicates a predefined distance from the starting point information. The predefined distances are as follows:
[0149] Table 2. Distance from IDX
[0150]
[0151] The direction index represents the direction of MVD relative to the starting point. The direction index can represent the four directions shown below.
[0152] Table 3. Directional IDX
[0153] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0154] Immediately after sending the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, then the UMVE flag is parsed. If the UMVE flag equals 1, then the UMVE syntax is parsed. However, if it is not equal to 1, then the AFFINE flag is parsed. If the AFFINE flag equals 1, it is in AFFINE mode; but if it is not equal to 1, then the skip / merge index is parsed according to the VTM's skip / merge mode.
[0155] No additional line caching is needed due to UMVE candidates, as the software's skip / merge candidates are directly used as the base candidates. When using the input UMVE index, the MV supplementation is determined just before motion compensation. There is no need to maintain a long line cache for this.
[0156] Under the current normal testing conditions, you can choose the first or second merge candidate in the merge candidate list as the base candidate.
[0157] UMVE is also known as Merge using MV difference (MMVD).
[0158] 2.2.3.4 Combined Intra-Frame and Inter-Frame Prediction (CIIP)
[0159] A multi-hypothesis prediction method is proposed, in which combined intra-frame and inter-frame prediction is a way to generate multiple hypotheses.
[0160] When applying multi-hypothesis prediction to improve intra-frame modes, multi-hypothesis prediction combines intra-frame prediction and merge index prediction. In the merge CU, a flag is notified for the merge mode signaling, and an intra-frame mode is selected from the intra-frame candidate list when this flag is true. For the luma component, the intra-frame candidate list is derived from only one intra-frame prediction mode, namely the planar mode. The weights applied to the prediction blocks from intra-frame and inter-frame predictions are determined by the encoding / decoding modes (intra-frame or non-intra-frame) of two adjacent blocks (A1 and B1).
[0161] 2.2.4 MERGE for Sub-Block Based Techniques
[0162] It is recommended that, in addition to the regular merge list used for non-sub-block merge candidates, all sub-block-related motion candidates be placed in a separate merge list.
[0163] The independent merge list of motion candidates related to placing sub-blocks is called the "sub-block merge candidate list".
[0164] In one example, the list of sub-block merge candidates includes ATMVP candidates and affine merge candidates.
[0165] Merge the candidate list using candidate fill sub-blocks in the following order:
[0166] a. ATMVP candidate (may be available or not);
[0167] b. Affine merge list (including inherited affine candidates; and constructed affine candidates)
[0168] c. Zero-filled MV 4-parameter affine model
[0169] 2.2.4.1.1 ATMVP (also known as Sub-Block Temporal Motion Vector Predictor, SbTMVP)
[0170] The basic idea of ATMVP is to derive multiple sets of temporal motion vector predictors for a block. A set of motion information is assigned to each sub-block. When generating ATMVP merge candidates, motion compensation is performed at the 8×8 level rather than the level of the entire block.
[0171] 2.2.5 Standard Inter-Frame Mode (AMVP)
[0172] 2.2.5.1 AMVP Sports Candidate List
[0173] Similar to the AMVP design in HEVC, up to two AMVP candidates can be derived. However, HMVP candidates can also be added after TMVP candidates. The HMVP candidates in the HMVP table are traversed in ascending index order (i.e., starting from the earliest index, equal to 0). Up to four HMVP candidates can be checked to determine if their reference images are the same as the target reference image (i.e., the same POC value).
[0174] 2.2.5.2 AMVR
[0175] In HEVC, when `use_integer_mv_flag` is equal to 0 in the strip header, the motion vector difference (MVD) is signaled in units of quarter-luminance samples (between the PU's motion vector and the predicted motion vector). In VVC, Local Adaptive Motion Vector Resolution (AMVR) is introduced. In VVC, MVD can be encoded and decoded in units of quarter-luminance samples, whole-luminance samples, and four-luminance samples (i.e., 1 / 4 pixel, 1 pixel, and 4 pixels). This MVD resolution is controlled at the codec unit (CU) level, and the MVD resolution flag is signaled in a conventional manner for each CU with at least one non-zero MVD component.
[0176] For a CU with at least one non-zero MVD component, a signaling signal notifies a first flag to indicate whether quarter-luminance sample MV accuracy is used in the CU. When the first flag (equal to 1) indicates that quarter-luminance sample MV accuracy is not used, a signaling signal notifies another flag to indicate whether whole-luminance sample MV accuracy or four-luminance sample MV accuracy is used.
[0177] When the first MVD resolution flag of the CU is zero, or when the CU is not encoded or decoded (meaning all MVDs within the CU are zero), a quarter-luminance sample MV resolution is used for the CU. When the CU uses integer luminance sample MV precision or four-luminance sample MV precision, the MVPs in the CU's AMVP candidate list are rounded relative to the corresponding precision.
[0178] 2.2.5.3 Symmetrical Motion Vector Difference
[0179] Symmetric Motion Vector Difference (SMVD) is applied to the encoding and decoding of motion information in bidirectional prediction.
[0180] First, as specified in N1001-v2, at the stripe level, the variables RefIdxSymL0 and RefIdxSymL1, which respectively indicate the reference image indices of list 0 / 1 used in SMVD mode, are derived using the following steps. SMVD mode should be disabled when at least one of these variables is equal to -1.
[0181] 2.2.6 Refinement of Motion Information
[0182] 2.2.6.1 Decoder-Side Motion Vector Refinement (DMVR)
[0183] In bidirectional prediction, for the prediction of a block, two prediction blocks are combined, formed using motion vectors (MV) from list 0 and MV from list 1 respectively, to create a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.
[0184] For DMVR in VVC, take the MVD mirror between list 0 and list 1, such as Figure 17As shown, bilateral matching is performed to refine the MV, i.e., to find the optimal MVD among several MVD candidates. The MVs of two reference image lists are represented by MVL0(L0X,L0Y) and MVL1(L1X,L1Y). The MVD of list 0 that minimizes the cost function (e.g., SAD), represented by (MvdX,MvdY), is defined as the optimal MVD. For the SAD function, it is defined as the SAD between the reference block of list 0 and the reference block of list 1, where the reference block of list 0 is derived using the motion vectors (L0X+MvdX,L0Y+MvdY) in the reference image of list 0, and the reference block of list 1 is derived using the motion vectors (L1X-MvdX,L1Y-MvdY) within the reference image of list 1.
[0185] The motion vector thinning process can be iterated twice. In each iteration, up to six MVDs (with integer pixel accuracy) can be checked in two steps, such as... Figure 18 As shown, the first step checks MVD(0,0), (-1,0), (1,0), (0,-1), and (0,1). In the second step, you can choose to check one of MVD(-1,-1), (-1,1), (1,-1), or (1,1) and perform further checks on it. Assume the function Sad(x,y) returns the SAD value of MVD(x,y). The MVD represented by (MvdX,MvdY) to be checked in the second step is determined as follows:
[0186] MvdX = -1;
[0187] MvdY = -1;
[0188] If (Sad(1,0)) <Sad(-1,0))
[0189] MvdX = 1;
[0190] If (Sad(0,1) <Sad(0,-1))
[0191] MvdY = 1;
[0192] In the first iteration, the starting point is the signaling notification MV, and in the second iteration, the starting point is the signaling notification MV plus the selected best MVD from the first iteration. DMVR is only applicable when one reference picture is the preceding picture and the other reference picture is the subsequent picture, and both reference pictures have the same picture order count distance relative to the current picture.
[0193] To further simplify the DMVR process, several changes to the design in JEM are proposed. More specifically, the DMVR design adopted by VTM-4.0 (to be released soon) has the following main features:
[0194] • Terminate early when SAD is less than the threshold at position (0,0) between list 0 and list 1.
[0195] • Terminate early when the SAD between list 0 and list 1 is zero for a given position.
[0196] • DMVR block size: W*H>=64&&H>=8, where W and H are the width and height of the block.
[0197] For DMVRs with CU dimensions > 16x16, divide the CU into multiple 16x16 sub-blocks. When only the width or height of the CU is greater than 16, divide it only in the vertical or horizontal direction.
[0198] • Reference block size (W+7)*(H+7) (for brightness).
[0199] • Integer pixel search based on 25-point SAD (i.e., (+-)2 refinement of the search range, single stage)
[0200] • DMVR based on bilinear interpolation.
[0201] • Subpixel thinning based on the "parameter error surface equation". This process is performed only if the lowest SAD cost is not equal to zero and the optimal MVD is (0,0) in the last MV thinning iteration.
[0202] • Luminosity / Chromaticity MC w / Reference Block Fill (if needed).
[0203] • A refined MV for MC and TMVP only.
[0204] 2.2.6.1.1 Use of DMVR
[0205] DMVR can be enabled when all of the following conditions are true:
[0206] – The DMVR enable flag in SPS (i.e., sps_dmvr_enabled_flag) is equal to 1.
[0207] The TPM flag, inter-frame affine flag, and sub-block merge flag (either ATMVP merge or affine merge) are all equal to 0.
[0208] The `-merge` flag is equal to 1.
[0209] - Perform bidirectional prediction on the current block, and the POC distance between the current image and the reference images in list 1 is equal to the POC distance between the reference images in list 0 and the current image.
[0210] - Current CU height is greater than or equal to 8
[0211] – The number of luminance samples (CU width * height) is greater than or equal to 64
[0212] 2.2.6.1.2 Sub-pixel thinning based on "parameter error surface equation"
[0213] The method will be summarized below:
[0214] 1. The fitted parameter error surface is calculated only if the center position is the optimal position in a given iteration.
[0215] 2. The center location cost and the cost at positions (-1,0), (0,-1), (1,0), and (0,1) from the center are used to fit a 2-D parabolic error surface equation of the following form.
[0216] E(x,y)=A(x-x0) 2 +B(y-y0) 2 +C
[0217] Here, (x0, y0) corresponds to the position with the lowest cost, and C corresponds to the lowest cost value. By solving the five equations containing five unknowns, (x0, y0) is calculated as follows:
[0218] x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0)))
[0219] y0=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))
[0220] (x0, y0) can be calculated to the desired subpixel precision by adjusting the precision on which the division is performed (i.e., how many bits are used to calculate the quotient). For 1 / 16 pixel precision, only 4 bits of the absolute value of the quotient need to be calculated, making it suitable for a fast shift subtraction-based implementation that requires 2 divisions per CU.
[0221] 3. The calculated (x0, y0) is added to the integer pixel distance thinning MV to obtain the sub-pixel precise thinning ΔMV.
[0222] 2.2.6.2 Bidirectional Optical Flow (BDOF)
[0223] 2.3 Intra-frame block copying
[0224] Intra-block copy (IBC), also known as current image reference, has been adopted by HEVC Screen Content Codec Extension (HEVC-SCC) and the current VVC test model (VTM-4.0). IBC extends the concept of motion compensation from inter-frame codec to intra-frame codec. Figure 18 As shown, when applying IBC, the current block is predicted using a reference block in the same image. Before encoding / decoding or decoding the current block, the samples in the reference block must have been reconstructed. Although IBC is not highly efficient for most camera capture sequences, it exhibits significant encoding / decoding gains for screen content. This is because screen content images contain numerous repeating patterns, such as icons and text characters. IBC effectively eliminates redundancy between these repeating patterns. In HEVC-SCC, if IBC selects the current image as its reference image, the codec unit (CU) for inter-frame encoding / decoding can apply IBC. In this case, the MV is renamed to a block vector (BV), which always has integer pixel precision. For compatibility with the major HEVC standard, the current image is marked as the "long-term" reference image in the decoded image buffer (DPB). It should be noted that, similarly, in various view / 3D video codec standards, inter-view reference images are also marked as "long-term" reference images.
[0225] By following the BV (Browser-Volume) pattern to find its reference block, predictions can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixel from the initial signal. Then, transforms and quantization can be applied as in other codec modes.
[0226] However, when the reference block is outside the image, overlaps with the current block, is outside the reconstructed region, or is outside the valid region limited by certain constraints, some or all pixel values are not defined. Basically, there are two approaches to handle this type of problem. One approach is to disallow such situations, for example, in bitstream consistency. The other approach is to apply padding for those undefined pixel values. The following subsections describe the approaches in detail.
[0227] 2.3.1 IBC (VTM4.0) in the VCC test model
[0228] In the current VCC test model, specifically the VTM-4.0 design, the entire reference block should utilize the current codec tree unit (CTU) and not overlap with the current block. Therefore, there is no need to pad the reference block or prediction block. The IBC flag is encoded and decoded into the prediction mode of the current CU. Thus, there are three prediction modes for each CU: MODE_INTRA, MODE_INTER, and MODE_IBC.
[0229] 2.3.1.1 IBC Merge Mode
[0230] In IBC Merge mode, the index pointing to an entry in the IBC merge candidate list is parsed from the bitstream. The construction of the IBC merge list can be summarized according to the following steps:
[0231] Step 1: Derivation of Airspace Candidates
[0232] Step 2: Insertion of HMVP candidates
[0233] Step 3: Insertion of Pairwise Average Candidates
[0234] In the derivation of the spatial merge candidate, located in Figure 2 Up to four merge candidates are selected from the candidates shown at positions A1, B1, B0, A0, and B2. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or if encoding / decoding is not performed using IBC mode. After the candidate at position A1 is added, the insertion of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency.
[0235] After inserting a spatial candidate, if the IBC merge list size is still smaller than the maximum IBC merge list size, an IBC candidate from the HMVP table can be inserted. Redundancy checks are performed when inserting an HMVP candidate.
[0236] Finally, pairs of average candidates are inserted into the IBC merge list.
[0237] A merge candidate is considered an invalid merge candidate if the reference block of the merge candidate identifier is outside the image, overlaps with the current block, is outside the reconstructed region, or is outside the valid region of certain constraints.
[0238] It should be noted that invalid merge candidates can be inserted into the IBC merge list.
[0239] 2.3.1.2 IBC AMVP Mode
[0240] In IBC AMVP mode, the AMVP index pointing to an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized according to the following steps:
[0241] Step 1: Derivation of Airspace Candidates
[0242] Check A0 and A1 until a usable candidate is found.
[0243] Check B0, B1, and B2 until a usable candidate is found.
[0244] Step 2: Insertion of HMVP candidates
[0245] Step 3: Insertion with zero candidates
[0246] After inserting a spatial candidate, if the size of the IBC AMVP list is still smaller than the size of the maximum IBC AMVP list, then an IBC candidate from the HMVP table can be inserted.
[0247] Finally, zero candidates were added to the IBC AMVP list.
[0248] 2.3.1.3 Chromaticity IBC Mode
[0249] In the current VVC, motion compensation in the chroma IBC mode is performed at the sub-block level. The chroma block is divided into several sub-blocks. Each sub-block checks whether its corresponding luma block has a block vector, and if so, determines its validity. Encoder constraints exist in the current VTM, where the chroma IBC mode is tested if all sub-blocks in the current chroma CU have valid luma block vectors. For example, in YUV 420 video, if the chroma block is N×M, then the co-occurring luma region is 2N×2M. The sub-block size of the chroma block is 2×2. Several steps are involved in performing chroma mv derivation, followed by the block copying process.
[0250] 1) The chroma block will first be divided into (N>>1)*(M>>1) sub-blocks.
[0251] 2) For each sub-block with the top left sample point coordinates (x, y), obtain the corresponding brightness block covering the same top left sample point with coordinates (2x, 2y).
[0252] 3) The encoder checks the block vector (bv) of the acquired luminance block. If one of the following conditions is met, the bv is considered invalid.
[0253] a. The bv corresponding to the luminance block does not exist.
[0254] b. The prediction block identified by bv has not yet been reconstructed.
[0255] c. The predicted block identified by bv partially or entirely overlaps with the current block.
[0256] 4) The chromaticity motion vector of the sub-block is set to the motion vector of the corresponding luminance sub-block.
[0257] When a valid bv is found in all sub-blocks, IBC mode is enabled at the encoder.
[0258] The decoding process for an IBC block is described below. The parts related to the chroma mv derivation in IBC mode are highlighted. grey .
[0259] 8.6.1 General Decoding Process of Encoding / Decoding Units Used in IBC Prediction
[0260] The input for this process is:
[0261] – Luminance position (xCb, yCb), relative to the top-left luminance sample of the current image, specifies the top-left sample of the current codec block.
[0262] – The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0263] – The variable cbHeight specifies the height of the current codec block in luminance samples.
[0264] - The variable treeType specifies whether to use a unary tree or a binary tree. If a binary tree is used, it specifies whether the current tree corresponds to the luminance component or the chrominance component.
[0265] The output of this process is the modified and reconstructed image before loop filtering.
[0266] Invoke the derivation process for the quantization parameters as specified in Clause 8.7.1, taking the luminance position (xCb, yCb), the width cbWidth of the current codec block in luminance samples, the height cbHeight of the current codec block in luminance samples, and the variable treeType as input.
[0267] In IBC prediction mode, the decoding process of the codec unit consists of the following ordered steps:
[0268] 1. The motion vector components of the current encoder / decoder unit are derived as follows:
[0269] 1. If treeType equals SINGLE_TREE or DUAL_TREE_LUMA, the following applies:
[0270] - Call the derivation process of the motion vector components specified in Clause 8.6.2.1, with the luminance codec block position (xCb, yCb), luminance codec block width cbWidth and luminance codec block height cbHeight as inputs, and the luminance motion vector mvL[0][0] as output.
[0271] - When treeType equals SINGLE_TREE, the derivation process of the chroma motion vector in clause 8.6.2.9 is invoked, with the luminance motion vector mvL[0][0] as input and the chroma motion vector mvC[0][0] as output.
[0272] The number of luminance codec sub-blocks in the horizontal direction (numSbX) and the number of luminance codec sub-blocks in the vertical direction (numSbY) are both set to 1.
[0273] 1. Otherwise, if treeType equals DUAL_TREE_CHROMA, the following applies:
[0274] The following derivation shows the number of luminance codec sub-blocks numSbX in the horizontal direction and the number numSbY in the vertical direction:
[0275] numSbX=(cbWidth>>2) (8886)
[0276] numSbY=(cbHeight>>2) (8887)
[0277] - For xSbIdx=0..BumSbX-1 and ySbIdx=0..numSbY-1, the chromaticity motion vector mvC is derived as follows. [xSbIdx][ySbIdx]:
[0278] - The brightness motion vector mvL[xSbIdx][ySbIdx] is derived as follows:
[0279] - The position of the co-bit luminance codec unit (xCuY, yCuY) is derived as follows:
[0280] xCuY=xCb+xSbIdx*4 (8888)
[0281] yCuY=yCb+ySbIdx*4 (8889)
[0282] - If CuPredMode[xCuY][yCuY] equals MODE INTRA, the following applies:
[0283] mvL[xSbIdx][ySbIdx][0]=0 (8890)
[0284] mvL[xSbIdx][ySbIdx][1]=0 (8891)
[0285] predFlagL0[xSbIdx][ySbIdx]=0 (8892)
[0286] predFlagL1[xSbIdx][ySbIdx]=0 (8893)
[0287] – Otherwise (CuPredMode[xCuY][yCuY] equals MODE_IBC), the following applies:
[0288] mvL[xSbIdx][ySbIdx][0]=MvL0[xCuY][yCuY][0] (8894)
[0289] mvL[xSbIdx][ySbIdx][1]=MvL0[xCuY][yCuY][1] (8895)
[0290] predFlagL0[xSbIdx][ySbIdx]=1 (8896)
[0291] predFlagL1[xSbIdx][ySbIdx]=0 (8897)
[0292] – Call the derivation of the chromaticity motion vector in Clause 8.6.2.9, with mvL[xSbIdx][ySbIdx] as input and mvC[xSbIdx][ySbIdx] as output.
[0293] – The requirement for bitstream consistency is that the chroma motion vector mvC[xSbIdx][ySbIdx] should comply with the following constraints:
[0294] – When invoking the derivation of block availability as specified in Clause 6.4.X [Ed.(BB): Adjacent Block Availability Check Procedure tbd], the current chromaticity position (xCurr, yCurr) and adjacent chromaticity positions (xCb / SubWidthC+(mvC[xSbIdx][ySbIdx][0]>>5), yCb / SubHeightC+(mvC[xSbIdx][ySbIdx][1]>>5)) set to (xCb / SubWidthC, yCb / SubHeightC) should be set to true as input.
[0295] – When invoking the derivation of block availability as specified in Clause 6.4.X [Ed.(BB): Adjacent Block Availability Check Procedure tbd], the current chromaticity position (xCurr, yCurr) and adjacent chromaticity positions (xCb / SubWidthC+(mvC[xSbIdx][ySbIdx][0]>>5)+cbWidth / SubWidthC-1, yCb / SubHeightC+(mvC[xSbIdx][ySbIdx][1]>>5)+cbHeight / SubHeightC-1) set to (xCb / SubWidthC, yCb / SubHeightC) should be set to true as input.
[0296] –One or both of the following conditions must be true:
[0297] –(mvC[xSbIdx][ySbIdx][0]>>5)+xSbIdx*2+2 is less than or equal to 0.
[0298] –(mvC[xSbIdx][ySbIdx][1]>>5)+ySbIdx*2+2 is less than or equal to 0.
[0299] 2. The prediction samples of the current encoding / decoding unit are derived as follows:
[0300] – If treeType equals SINGLE_TREE or DUAL_TREE_LUMA, then the prediction samples of the current encoder / decoder unit are derived as follows:
[0301] – The decoding process of the ibc block as specified in Clause 8.6.3.1 is invoked, taking the luma codec block position (xCb, yCb), luma codec block width cbWidth and luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the luma motion vector mvL[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and the variable cIdx set to 0 as input, and the (cbWidth)x(cbHeight) array predSamples as the predicted luma samples. L The IBC prediction samples (predSamples) are used as the output.
[0302] Otherwise, if treeType equals SINGLE_TREE or DUAL_TREE_CHROMA, the prediction samples of the current encoder / decoder unit are derived as follows:
[0303] – The decoding process of the ibc block as specified in Clause 8.6.3.1 is invoked, taking the luma codec block position (xCb, yCb), luma codec block width cbWidth and luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the chromaticity motion vector mvC[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and the variable cIdx set to 1 as input, and the (cbWidth / 2)x(cbHeight / 2) array predSamples as the predicted chromaticity samples for the chromaticity component Cb. Cb The IBC prediction samples (predSamples) are used as the output.
[0304] – The decoding process of the ibc block as specified in Clause 8.6.3.1 is invoked, taking the luma codec block position (xCb, yCb), luma codec block width cbWidth and luma codec block height cbHeight, the number of luma codec sub-blocks in the horizontal direction numSbX and the number of luma codec sub-blocks in the vertical direction numSbY, the chromaticity motion vector mvC[xSbIdx][ySbIdx] where xSbIdx = 0..numSbX-1 and ySbIdx = 0..numSbY-1, and the variable cIdx set to equal 2 as input, and the (cbWidth / 2)x(cbHeight / 2) array predSamples as the predicted chromaticity samples for the chromaticity component Cr. Cr The IBC prediction samples (predSamples) are used as the output.
[0305] 3. The variables NumSbX[xCb][yCb] and NumSbY[xCb][yCb] are set to be equal to numSbX and numSbY, respectively.
[0306] 4. The residual samples of the current encoding / decoding unit are derived as follows:
[0307] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_LUMA, invoke the decoding process of the residual signal of the codec block encoded and decoded in inter-frame prediction mode as specified in Clause 8.5.8, with the position (xTb0, yTb0) set to be equal to the luma position (xCb, yCb), the width nTbW set to be equal to the luma codec block width cbWidth, the height nTbH set to be equal to the luma codec block height cbHeight, and the variable cIdx set to be equal to 0 as input, and with the array resSamples L As output.
[0308] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_CHROMA, invoke the decoding process of the residual signal of the codec block encoded and decoded in inter-frame prediction mode as specified in Clause 8.5.8, with the position (xTb0, yTb0) set to equal the chroma position (xCb / 2, yCb / 2), the width nTbW set to equal the width cbWidth / 2 of the chroma codec block, the height nTbH set to equal the height cbHeight / 2 of the chroma codec block, and the variable cIdx set to equal to 1 as input, and with the array resSamples Cb As output.
[0309] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_CHROMA, invoke the decoding process of the residual signal of the codec block encoded and decoded in inter-frame prediction mode as specified in Clause 8.5.8, with the position (xTb0, yTb0) set to equal the chroma position (xCb / 2, yCb / 2), the width nTbW set to equal the width cbWidth / 2 of the chroma codec block, the height nTbH set to equal the height cbHeight / 2 of the chroma codec block, and the variable cIdx set to equal to 2 as input, and with the array resSamples Cr As output.
[0310] 5. The reconstruction sample points of the current encoding / decoding unit are derived as follows:
[0311] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_LUMA, invoke the image reconstruction procedure for color components as specified in Item 8.7.5, with the block position (xB, yB) set to (xCb, yCb), the block width bWidth set to cbWidth, the block height bHeight set to cbHeight, the variable cIdx set to 0, and the variable predSamples set to... L The (cbWidth) x (cbHeight) array predSamples is set to equal resSamples. L The (cbWidth)x(cbHeight) array resSamples is used as input, and the output is the modified reconstructed image before loop filtering.
[0312] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_CHROMA, invoke the image reconstruction procedure for color components as specified in Item 8.7.5, with the block position (xB, yB) set to (xCb / 2, yCb / 2), the block width bWidth set to cbWidth / 2, the block height bHeight set to cbHeight / 2, the variable cIdx set to 1, and the variable predSamples set to... Cb The (cbWidth / 2) x (cbHeight / 2) array predSamples is set to equal resSamples. CbThe (cbWidth / 2)x(cbHeight / 2) array resSamples is used as input, and the output is the modified reconstructed image before loop filtering.
[0313] – When treeType equals SINGLE_TREE or treeType equals DUAL_TREE_CHROMA, invoke the image reconstruction procedure for color components as specified in Item 8.7.5, with the block position (xB, yB) set to (xCb / 2, yCb / 2), the block width bWidth set to cbWidth / 2, the block height bHeight set to cbHeight / 2, the variable cIdx set to 2, and the variable predSamples set to 2. Cr The (cbWidth / 2) x (cbHeight / 2) array predSamples is set to equal resSamples. Cr The (cbWidth / 2)x(cbHeight / 2) array resSamples is used as input, and the output is the modified reconstructed image before loop filtering.
[0314] 2.3.2 Recent Developments in IBC (in VTM5.0)
[0315] 2.3.2.1 Single BV List
[0316] In IBC, BV predictors used in merge mode and AMVP mode will share a common predictor list, which consists of the following elements:
[0317] • The location of two adjacent airspaces ( Figure 2 (A1, B1)
[0318] · 5 HMVP entries
[0319] • Zero vectors by default
[0320] The number of candidates in the list is controlled by a variable derived from the strip header. For merge mode, up to the first 6 entries of this list will be used; for AMVP mode, the first 2 entries will be used. This list must conform to the shared merge list region requirement (share the same list within SMR).
[0321] In addition to the aforementioned BV prediction sub-candidate list, a simplified pruning operation is proposed between the HMVP candidate and the existing merge candidate (A1, B1). In this simplification, there will be a maximum of two pruning operations, as it only compares the first HMVP candidate with the spatial merge candidate.
[0322] 2.3.2.1.1 Decoding process
[0323] 8.6.2.2 Derivation of the method used for IBC brightness motion vector prediction
[0324] This process is called only when CuPredMode[xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top-left sample of the current luminance codec block relative to the top-left luminance sample of the current image.
[0325] The input for this process is:
[0326] – The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.
[0327] – The variable cbWidth specifies the width of the current codec block in units of luminance samples.
[0328] – The variable cbHeight specifies the height of the current codec block in luminance samples. The output of this process is:
[0329] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[0330] The following is a derivation of the variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand:
[0331] xSmr=IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:xCb (8910)
[0332] ySmr=IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:yCb (8911)
[0333] smrWidth=IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:cbWidth(8912)
[0334] smrHeight=IsInSmr[xCb][yCb]? SmrH[xCb][yCb]:cbHeight(8913)
[0335] smrNumHmvpIbcCand=IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand:NumHmvpIbcCand(8914)
[0336] The brightness motion vector mvL is derived through the following ordered steps:
[0337] 1. Invoke the derivation process of spatial motion vector candidates from adjacent codec units as specified in Clause 8.6.2.3, taking the luma codec block position (xCb, yCb) set to be equal to (xSmr, ySmr), the luma codec block width cbWidth and the luma codec block height cbHeight set to be equal to smrWidth and smrHeight as input, and outputting availability flags availableFlagA1, availableFlagB1 and motion vectors mvA1 and mvB1.
[0338] 2. Construct the motion vector candidate list mvCandList as follows:
[0339] i=0
[0340] if(availableFlagA1)
[0341] mvCandList[i++] = mvA1 (8915)
[0342] if(availableFlagB1)
[0343] mvCandList[i++] = mvB1
[0344] 3. The variable numCurrCand is set to be equal to the number of merging candidates in mvCandList.
[0345] 4. When numCurrCand is less than MaxNumMergeCandand and smrNumHmvpIbcCand is greater than 0, call the derivation process of motion vector candidates based on IBC history as specified in 8.6.2.4, taking mvCandList, isInSmr set to be equal to IsInSmr[xCb][yCb] and numCurrCand as input, and taking the modified mvCandList and numCurrCand as output.
[0346] 5. When numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand equals MaxNumMergeCand:
[0347] 1. Set mvCandList[numCurrCand][0] to equal 0.
[0348] 2. Set mvCandList[numCurrCand][1] to equal 0.
[0349] 3. Increase numCurrCand by 1.
[0350] 6. The following is the derivation of the variable mvIdx:
[0351] mvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8916)
[0352] 7. Assign the following value:
[0353] mvL[0]=mergeCandList[mvIdx][0] (8917)
[0354] mvL[1]=mergeCandList[mvIdx][1] (8918)
[0355] 2.3.2.2 Size Limitations of IBC
[0356] In the latest VVC and VTM5, it is proposed to explicitly use syntax constraints to disable the 128x128 IBC mode on top of the current bitstream constraints in previous VTM and VVC versions. This makes the presence of the IBC flag dependent on the CU size <128x128.
[0357] 2.3.2.3 Shared merge list for IBC
[0358] To reduce decoder complexity and support parallel encoding, a proposal is made to share the same merging candidate list among all leaf codec units (CUs) of an ancestor node in the CU partitioning tree, enabling parallel processing of small CUs for skip / merge encoding / decoding. The ancestor node is named the merge-shared node. A shared merging candidate list is generated at the merge-shared node, assuming the merge-shared node is a leaf CU.
[0359] More specifically, the following may apply:
[0360] - If a block has no more than 32 luminance samples and is divided into two 4×4 sub-blocks, a shared merge list is used between very small blocks (e.g., two adjacent 4×4 blocks).
[0361] However, if the block has more than 32 luminance samples, and after partitioning, at least one sub-block is less than the threshold (32), then all sub-blocks of the partition share the same merge list (e.g., 16×4 or 4×16 ternary partitioning or 8×8 quadrilateral partitioning).
[0362] This limitation applies only to IBC merge mode.
[0363] 2.4 Syntax and Semantics for Encoding / Decoding Units and Merge Modes
[0364] 7.3.5.1 General Strip Header Syntax
[0365]
[0366]
[0367] 7.3.7.5 Coding unit syntax
[0368]
[0369]
[0370]
[0371] 7.3.7.7 Merge data syntax
[0372]
[0373]
[0374]
[0375] 7.4.6.1 General Strip Header Semantics
[0376] `six_minus_max_num_merge_cand` specifies the maximum number of merging motion vector prediction (MVP) candidates supported from the strips subtracted from 6. The maximum number of merging MVP candidates, `MaxNumMergeCand`, is derived as follows:
[0377] MaxNumMergeCand=6-six_minus_max_num_merge_cand (757)
[0378] The value of MaxNumMergeCand should be in the range of 1 to 6, inclusive.
[0379] `five_minus_max_num_subblock_merge_cand` specifies the maximum number of subblock-based merging motion vector prediction (MVP) candidates supported from the stripes subtracted from 5. If `five_minus_max_num_subblock_merge_cand` is not present, it is inferred to be equal to 5 - `sps_sbtmvp_enabled_flag`. The maximum number of subblock-based merging MVP candidates, `MaxNumSubblockMergeCand`, is derived as follows:
[0380] MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand(758)
[0381] The value of MaxNumSubblockMergeCand should be in the range of 0 to 5, inclusive.
[0382] 7.4.8.5 Semantics of Encoding / Decoding Units
[0383] A pred_mode_flag value of 0 indicates that the current codec unit is encoded and decoded in inter-frame prediction mode. A pred_mode_flag value of 1 indicates that the current codec unit is encoded and decoded in intra-frame prediction mode.
[0384] When pred_mode_flag does not exist, the following inference is made:
[0385] – If cbWidth equals 4 and cbHeight equals 4, then pred_mode_flag is inferred to be equal to 1.
[0386] Otherwise, when decoding I stripes, pred_mode_flag is inferred to be equal to 1, and when decoding P or B stripes, it is inferred to be equal to 0.
[0387] For x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the variable CuPredMode[x][y] is derived as follows:
[0388] – If pred_mode_flag equals 0, then CuPredMode[x][y] is set to equal MODE_INTER.
[0389] Otherwise (pred_mode_flag equals 1), CuPredMode[x][y] is set to equal MODE_INTRA.
[0390] A pred_mode_ibc_flag value of 1 indicates that the current codec unit is encoded and decoded in IBC predictive mode. A pred_mode_ibc_flag value of 0 indicates that the current codec unit is not encoded and decoded in IBC predictive mode.
[0391] When pred_mode_ibc_flag does not exist, the following inference is made:
[0392] – If cu_skip_flag[x0][y0] equals 1, cbWidth equals 4, and cbHeight equals 4, then pred_mode_ibc_flag equals 1.
[0393] Otherwise, if both cbWidth and cbHeight are equal to 128, it is inferred that pred_mode_ibc_flag is equal to 0.
[0394] Otherwise, when decoding I-stripes, it is inferred that pred_mode_ibc_flag is equal to the value of sps_ibc_enabled_flag, and when decoding P or B-stripes, it is inferred that it is equal to 0.
[0395] When pred_mode_ibc_flag equals 1, for x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1, the variable CuPredMode[x][y] is set to equal MODE_IBC.
[0396] `general_merge_flag[x0][y0]` specifies whether the inter-frame prediction parameters used for the current codec unit are inferred from adjacent inter-frame prediction segments. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0397] When general_merge_flag[x0][y0] does not exist, the following inference is made:
[0398] – If cu_skip_flag[x0][y0] equals 1, then it is inferred that general_merge_flag[x0][y0] equals 1.
[0399] Otherwise, infer that general_merge_flag[x0][y0] equals 0.
[0400] mvp_l0_flag[x0][y0] specifies the motion vector prediction sub-index of list 0, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0401] If mvp_l0_flag[x0][y0] does not exist, it is inferred that it is equal to 0.
[0402] mvp_l1_flag[x0][y0] has the same semantics as mvp_l0_flag, where l0 and list 0 are replaced by l1 and list 1, respectively.
[0403] inter_pred_idc[x0][y0] specifies whether to use list 0, list 1, or bidirectional prediction for the current codec unit according to Table 710. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0404] Table 7-10 — Association with the Names of Inter-Frame Prediction Modes
[0405]
[0406] If inter_pred_idc[x0][y0] does not exist, it is inferred that it is equal to PRED_L0.
[0407] 7.4.8.7 Merge Data Semantics
[0408] The regular_merge_flag[x0][y0] value of 1 specifies that the inter-frame prediction parameters for the current codec unit are generated using the regular merge mode. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0409] When regular_merge_flag[x0][y0] does not exist, the following inference is made:
[0410] – If all of the following conditions are true, then it is inferred that regular_merge_flag[x0][y0] equals 1:
[0411] –sps_mmvd_enabled_flag equals 0.
[0412] –general_merge_flag[x0][y0] equals 1.
[0413] –cbWidth*cbHeight equals 32.
[0414] Otherwise, infer that regular_merge_flag[x0][y0] equals 0.
[0415] The `mmvd_merge_flag[x0][y0]` value of 1 specifies that the inter-frame prediction parameters for the current codec unit are generated using the merge mode with motion vector differences. The array indices `x0` and `y0` specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0416] When mmvd_merge_flag[x0][y0] does not exist, the following inference is made:
[0417] – If all of the following conditions are true, then it is inferred that mmvd_merge_flag[x0][y0] equals 1:
[0418] –sps_mmvd_enabled_flag equals 1.
[0419] –general_merge_flag[x0][y0] equals 1.
[0420] –cbWidth*cbHeight equals 32.
[0421] –regular_merge_flag[x0][y0] equals 0.
[0422] Otherwise, infer that mmvd_merge_flag[x0][y0] equals 0.
[0423] mmvd_cand_flag[x0][y0] specifies whether to use the first (0) or second (1) candidate in the merging candidate list with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0424] If mmvd_cand_flag[x0][y0] does not exist, it is inferred that it is equal to 0.
[0425] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 712. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0426] Table 7-12 — Specification of MmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0].
[0427]
[0428] mmvd_direction_idx[x0][y0] specifies the index used to derive MmvdSign[x0][y0] as specified in Table 713. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0429] Table 7-13 — Specification of MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]
[0430]
[0431]
[0432] The following is a derivation of the two components of merge plus MVD offset MmvdOffset[x0][y0]:
[0433] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0](7124)
[0434] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1](7125)
[0435] `merge_subblock_flag[x0][y0]` specifies whether the sub-block-based inter-frame prediction parameters used for the current codec unit are inferred from neighboring blocks. Array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the considered codec block relative to the top-left luminance sample of the image. If `merge_subblock_flag[x0][y0]` does not exist, it is inferred to be equal to 0.
[0436] merge_subblock_idx[x0][y0] specifies the merging candidate index of the merging candidate list based on subblocks, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0437] If merge_subblock_idx[x0][y0] does not exist, it is inferred that it is equal to 0.
[0438] ciip_flag[x0][y0] specifies whether to apply combined inter-frame merge and intra-frame prediction to the current codec unit. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0439] If ciip_flag[x0][y0] does not exist, it is inferred that it is equal to 0.
[0440] When ciip_flag[x0][y0] equals 1, the variable IntraPredModeY[x][y] is set to equal INTRA_PLANAR, where x = xCb..xCb + cbWidth - 1 and y = yCb..yCb + cbHeight - 1.
[0441] The variable MergeTriangleFlag[x0][y0] specifies whether to use triangle-shaped motion compensation to generate prediction samples for the current codec unit when decoding B-stripes. The derivation of this variable is as follows:
[0442] – If all of the following conditions are true, then set MergeTriangleFlag[x0][y0] to equal 1:
[0443] –sps_triangle_enabled_flag equals 1.
[0444] –slice_type equals B.
[0445] –general_merge_flag[x0][y0] equals 1.
[0446] –MaxNumTriangleMergeCand is greater than or equal to 2.
[0447] –cbWidth*cbHeight is greater than or equal to 64.
[0448] –regular_merge_flag[x0][y0] equals 0.
[0449] –mmvd_merge_flag[x0][y0] equals 0.
[0450] –merge_subblock_flag[x0][y0] equals 0.
[0451] –ciip_flag[x0][y0] equals 0.
[0452] Otherwise, set MergeTriangleFlag[x0][y0] to 0.
[0453] `merge_triangle_split_dir[x0][y0]` specifies the direction of the merge triangle pattern. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0454] If merge_triangle_split_dir[x0][y0] does not exist, it is inferred that it is equal to 0.
[0455] merge_triangle_idx0[x0][y0] specifies a list of motion compensation candidates based on the first merging candidate index of the triangle shape, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0456] If merge_triangle_idx0[x0][y0] does not exist, it is inferred that it is equal to 0.
[0457] merge_triangle_idx1[x0][y0] specifies a motion compensation candidate list based on the second merging candidate index of the triangle shape, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0458] If merge_triangle_idx1[x0][y0] does not exist, it is inferred that it is equal to 0.
[0459] merge_idx[x0][y0] specifies the merging candidate index in the merging candidate list, where x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.
[0460] When merge_idx[x0][y0] does not exist, the following inference is made:
[0461] – If mmvd_merge_flag[x0][y0] equals 1, infer merge_idx[x0][y0]
[0462] It equals mmvd_cand_flag[x0][y0].
[0463] Otherwise, if mmvd_merge_flag[x0][y0] equals 0, it is inferred that merge_idx[x0][y0] equals 0.
[0464] Examples of technical problems solved by the 3 embodiments
[0465] The current IBC may have the following problems:
[0466] 1. IBC is only applicable to square or rectangular codec units. Using techniques similar to TPM on IBC codec blocks may result in additional codec gains.
[0467] 4. Various technologies and embodiments
[0468] The detailed techniques described below should be considered as examples to illustrate general concepts. These techniques should not be interpreted narrowly. Furthermore, these inventions can be combined in any way.
[0469] In this section, intra-block copying (IBC) may not be limited to current IBC techniques, but can be interpreted as a technique that uses reference samples within the current strip / piece / brick / picture / other video unit (e.g., CTU line).
[0470] IBC / intraframe can be applied to N (N>1) sub-regions within a block that is divided into two or more triangular or wedge-shaped sub-regions.
[0471] 1. The following are some examples of sub-regions generated by triangle / wedge segmentation.
[0472] a. Figure 20 The image shows an example of a block divided into two triangular regions.
[0473] b. Figure 21 The image shows an example of a block divided into four triangular regions.
[0474] c. Figure 22 The image shows an example of a block divided into two triangular regions.
[0475] Pattern combination for sub-regions
[0476] 2. When a block is divided into multiple sub-regions, all sub-regions can be encoded and decoded in IBC mode, however, at least two of them are encoded and decoded using different motion vectors (aka BV).
[0477] a. In one example, motion vectors (MV) or motion vector differences (MVD) derived from the MV predictor of a subregion can be encoded or decoded.
[0478] i. In one example, predictive encoding and decoding of the motion vector of a subregion relative to the motion vector can be used.
[0479] ii. In one example, the motion vector predictor can be derived from the regular IBC AMVP candidate list.
[0480] iii. In one example, predictive encoding and decoding of motion vectors of one sub-region relative to another sub-region can be used.
[0481] iv. In one example, sub-regions can share a single AMVP candidate list.
[0482] v. In one example, a sub-region can utilize a different list of AMVP candidates.
[0483] b. In one example, the motion vector of a subregion can be inherited or derived from the MV candidate index.
[0484] i. In one example, the motion vector of a sub-region can be derived from a regular IBC merge candidate list.
[0485] ii. In one example, the motion vector of a subregion may be derived from the MV candidate list, which may differ from the regular IBC merge candidate list construction process.
[0486] 1. In one example, you can check different spatial adjacent (adjacent or non-adjacent) blocks in the MV candidate list.
[0487] iii. Can be signaled or candidate indexes can be derived.
[0488] 1. Predictive encoding and decoding of candidate indices can be applied. For example, a candidate index for one sub-region can be derived from a candidate index for another sub-region.
[0489] iv. In one example, sub-regions can share a single merge candidate list.
[0490] v. In one example, different sub-regions can utilize different lists of merge candidates.
[0491] c. In one example, the motion vector of one sub-region may be inherited or derived from the MV candidate index; the MV / MVD of another sub-region may be obtained through encoding and decoding.
[0492] d. In one example, the motion vector (MV) of a subregion can be derived from a list of MV candidates (e.g., a regular IBCmerge candidate list).
[0493] i. In one example, the first M (e.g., M = N) candidates in the candidate list can be used.
[0494] 3. When a block is divided into multiple sub-regions, at least one of the sub-regions may be encoded and decoded using IBC mode, and another sub-region may be encoded and decoded using non-IBC mode.
[0495] a. In one example, the other sub-region could be encoded or decoded using intra-frame mode.
[0496] b. In one example, the other sub-region could be encoded and decoded using inter-frame mode.
[0497] c. In one example, this other sub-region could be encoded or decoded using a palette mode.
[0498] d. In one example, the other sub-region could be encoded and decoded using PCM mode.
[0499] e. In one example, the other sub-region could be encoded and decoded using the RDPCM mode.
[0500] f. The motion vector of a sub-region of the IBC codec can be obtained using the sub-bullet in bullet 2.
[0501] 4. When a block is divided into multiple sub-regions, at least one of the sub-regions may be encoded and decoded using intra-frame mode, and another sub-region may be encoded and decoded using non-intra-frame mode.
[0502] a. In one example, the other sub-region could be encoded and decoded using inter-frame mode.
[0503] b. Motion vectors of sub-regions of inter-frame encoding / decoding can be obtained using the same method as for regular inter-frame encoding / decoding blocks.
[0504] 5. When a block is divided into multiple sub-regions, all sub-regions can be encoded and decoded using a palette mode, however, at least two of them must be encoded and decoded using different palettes.
[0505] 6. When a block is divided into multiple sub-regions, at least one of the sub-regions may be encoded and decoded using a palette mode, and another sub-region may be encoded and decoded using a non-palette mode.
[0506] 7. The proposed method can be applied to only one component (or some components), such as the luminance component.
[0507] a. Optionally, when the proposed method is applied to one color component but not to another, the corresponding block in the other color component can be encoded and decoded as an undivided whole block.
[0508] i. Optionally, the encoding / decoding method used for the entire block in the other color component can be predefined, such as IBC / inter-frame / intra-frame methods.
[0509] b. Alternatively, the proposed method can be applied to all components.
[0510] i. In one example, the chromaticity blocks can be divided using the same dividing pattern as the luminance components.
[0511] ii. In one example, a different partitioning pattern than that used for the luminance component can be used to partition the chroma blocks.
[0512] Generation of the final prediction block
[0513] 8. When dividing a block into multiple sub-regions, you can first generate a prediction for each sub-region and then use all the predictions to obtain the final predicted block.
[0514] a. Alternatively, intermediate prediction blocks can be generated using information from each sub-region. The final prediction block can be obtained by weighted averaging of the intermediate prediction blocks.
[0515] i. In one example, equal weights can be used.
[0516] ii. In one example, unequal weights can be applied, including (0,…,0,1) weights.
[0517] iii. In one example, one or more sets of weights can be predefined to combine intermediate prediction blocks.
[0518] iv. In one example, IBC can be treated as an inter-frame mode, and CIIP weights can be applied to combine intra-frame prediction and IBC prediction.
[0519] v. In one example, the weight on a sample point can depend on the sample point's relative position within the current block.
[0520] vi. In one example, the weight on a sample point can depend on the sample point's location relative to the edge of the sub-region.
[0521] vii. In one example, the weights can depend on the encoding / decoding information of the current block, such as the intra-prediction mode, block dimension, color components, and color format.
[0522] Storage of encoded and decoded information
[0523] 9. When applying the above methods, the motion information of a sub-region can be stored as the motion information of the entire block.
[0524] a. Alternatively, motion information can be stored for each basic unit (e.g., the smallest CU size), as in normal inter-frame mode.
[0525] i. For a basic unit that covers multiple sub-regions, a set of motion information can be selected / derived and stored.
[0526] b. In one example, the stored motion information can be utilized in a loop filtering process (e.g., a deblocking filter).
[0527] c. In one example, the stored motion information can be utilized in the encoding and decoding of subsequent blocks.
[0528] Interaction with other tools
[0529] 10. When applying the above methods, it is not necessary to update the IBC HMVP table.
[0530] a. Alternatively, one or more motion vectors from the motion vectors of the sub-regions used for IBC encoding and decoding can be used to update the IBC HMVP table.
[0531] 11. When applying the above methods, it is not necessary to update the non-IBC HMVP tables.
[0532] a. Alternatively, one or more motion vectors from the motion vectors of the sub-regions used for inter-frame encoding and decoding can be used to update the non-IBC HMVP table.
[0533] 12. Loop filtering processes (e.g., deblocking processes) may depend on the use of the methods described above.
[0534] a. In one example, it is not necessary to filter the samples in the block encoded or decoded using the above method.
[0535] b. In one example, blocks encoded or decoded using the above method can be treated in a similar manner to blocks encoded or decoded using regular IBC encoding or decoding.
[0536] 13. For blocks encoded and decoded using the above methods, certain encoding and decoding methods can be disabled (e.g., sub-block transform, affine motion prediction, multi-reference line intra-prediction, matrix-based intra-prediction, symmetric MVD encoding and decoding, merge with MVD decoder-side motion derivation / refinement, bidirectional optical flow, simplified quadratic transform, multi-transform set, etc.).
[0537] 14. Signaling notifications or instructions for the use of the above methods and / or weighting values can be derived at the sequence / image / strip / slice group / piece / brick / CTU / CTB / CU / PU / TU / other video unit level.
[0538] a. In one example, the above method can be treated as a special IBC mode.
[0539] i. Optionally, if a block is encoded or decoded in accordance with the IBC mode, further instructions may be signaled or deduced to use the conventional block-based IBC method or the methods described above.
[0540] ii. In one example, subsequent IBC codec blocks can utilize the motion information of the current block as MV predictors.
[0541] 1. Alternatively, subsequent IBC encoding / decoding blocks may not be allowed to use the motion information of the current block as MV predictors.
[0542] b. In one example, the above method can be treated as a special triangle pattern. That is, if a block is encoded and decoded according to the triangle pattern, further instructions to use the regular TPM pattern or the above method can be derived through signaling.
[0543] c. In one example, the above method can be treated as a new prediction mode. That is, the allowed modes can be further expanded, such as intra-frame, inter-frame, and IBC modes, to include this new mode.
[0544] 15. Whether and / or how to apply the above methods may depend on the following information:
[0545] a. Signaling notification messages in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Piece Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video Codec Unit
[0546] b. Location of CU / PU / TU / block / video codec unit
[0547] c. Block dimensions of the current block and / or its neighboring blocks
[0548] d. The block shape of the current block and / or its neighboring blocks
[0549] e. Intra-frame mode of the current block and / or its adjacent blocks
[0550] f. The motion / block vector of its adjacent blocks
[0551] g. Indication of color format (e.g., 4:2:0, 4:4:4)
[0552] h. Encoding / decoding tree structure
[0553] i. Strip / group type and / or image type
[0554] j. Color components (e.g., can be applied only to the chromaticity component or the luminance component)
[0555] k. Time-domain layer ID
[0556] l. Standard profile / level / layer
[0557] Figure 23 This is a block diagram of a video processing apparatus 2300. Apparatus 2300 can be used to implement one or more of the methods described herein. Apparatus 2300 can be incorporated into smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 2300 may include one or more processors 2302, one or more memories 2304, and video processing hardware 2306. Processor 2302 can be configured to implement one or more methods described herein. Memory(s) 2304 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 2306 can be used to implement some of the techniques described herein in hardware circuitry. Video processing hardware 2306 may be partially or wholly included within processor 2302, which may be in the form of dedicated hardware or a graphics processing unit (GPU) or dedicated signal processing block.
[0558] The following terms describe some embodiments and techniques. Figures 20 to 22 Examples of wedge and triangular segmentation are given. For example, wedge segmentation can divide a block into multiple parts, where the boundaries of the parts form a stepped pattern, or have vertical and horizontal parts that converge at the corners. For example, triangular segmentation can include dividing into sub-blocks such that adjacent sub-blocks are separated by non-horizontal or non-vertical segmentation.
[0559] 1. A video processing method (e.g., Figure 24 The method 2400 shown includes: during the conversion of a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and the bitstream representation of the current video block, determining (2402) that the conversion of each sub-region is based on an encoding / decoding method using the current picture as a reference picture; based on the determination, determining (2404) motion vectors for IBC encoding / decoding of the multiple sub-regions; and performing (2406) the conversion using two different motion vectors.
[0560] 2. The method according to Clause 1, wherein the bit stream represents an indication of motion vectors for a given sub-region including multiple sub-regions, or motion vector differences for a given sub-region derived from corresponding motion vector predictors.
[0561] 3. The method according to Clause 2, wherein the motion vector predictor is derived from the IBC motion vector predictor list.
[0562] 4. The method according to any one of Clauses 1-2, wherein the motion vector is determined using an inheritance operation or a derivation operation, wherein the motion vector is inherited or derived.
[0563] 5. The method according to Clause 4, wherein the derivation operation includes deriving motion vectors for subregions based on the IBC merge candidate list.
[0564] 6. The method according to any one of Clauses 1-2, wherein the motion vector of a subset of the subregion is inherited or derived, and the motion vector of another subset of the subregion is signaled in the bitstream representation.
[0565] 7. The method according to Clause 1, wherein the encoding / decoding method includes intra-block copying (IBC).
[0566] Item 2 in the previous section provides additional examples and variations of the above terms.
[0567] 8. A video processing method comprising: during a conversion between a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on an encoding / decoding method using the current image as a reference image, and the conversion of the other sub-region is based on an alternative encoding / decoding mode different from the IBC encoding / decoding mode; and performing the conversion based on the determination.
[0568] 9. The method according to Clause 1, wherein the other encoding / decoding mode is an intra-frame encoding / decoding mode.
[0569] 10. The method according to Clause 1, wherein the other encoding / decoding mode is an inter-frame encoding / decoding mode.
[0570] 11. The method according to Clause 1, wherein the other encoding / decoding mode is a palette encoding / decoding mode.
[0571] 12. The method according to Clauses 8-11, wherein the encoding / decoding method includes (IBC).
[0572] Item 3 in the previous section provides additional examples and variations of the above terms.
[0573] 12. A video processing method comprising: during a conversion between a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on an intra-frame encoding / decoding mode and the conversion of the other sub-region is based on an alternative encoding / decoding mode different from the intra-frame encoding / decoding mode; and performing the conversion based on the determination.
[0574] 13. The method according to Clause 12, wherein the other encoding / decoding mode is an inter-frame encoding / decoding mode.
[0575] Item 4 in the previous section provides additional examples and variations of the above terms.
[0576] 14. A video processing method comprising: during a conversion between a current video block segmented into sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining to encode / decode the sub-regions based on a palette encoding / decoding mode, wherein the palettes of at least two sub-regions are different from each other; and performing a conversion based on the determination.
[0577] Item 5 in the previous section provides additional examples and variations of the above terms.
[0578] 15. A video processing method comprising: during a conversion between a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining that the conversion of one sub-region is based on a palette encoding / decoding mode and the conversion of the other sub-region is based on an alternative encoding / decoding mode different from the palette encoding / decoding mode; and performing the conversion based on the determination.
[0579] Item 6 in the previous section provides additional examples and variations of the above terms.
[0580] 16. A video processing method comprising: performing a conversion between a block of a video component and a bitstream representation of the block using the method described in any one of the preceding clauses, such that a conversion between corresponding blocks of another component of the video is performed using a different encoding / decoding technique than that used for the conversion of the block.
[0581] Item 7 in the previous section provides additional examples and variations of the above terms.
[0582] 17. A video processing method comprising: during a conversion between a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block, determining a final prediction for the current video block using predictions for each sub-region; and performing the conversion based on the determination.
[0583] 18. The method according to Clause 17, wherein the determination is based on a weighted prediction of the prediction for each sub-region.
[0584] 19. The method according to Clause 18, wherein equal weights are applied to the predictions for each sub-region to determine the weighted predictions.
[0585] Item 8 in the previous section provides additional examples and variations of the above terms.
[0586] 20. A video processing method, comprising: storing motion information of a sub-region as motion information of the current video block during a conversion between a current video block divided into at least two sub-regions based on wedge segmentation or triangle segmentation and a bitstream representation of the current video block; and using the motion information to convert the current video block or a subsequent video block.
[0587] 21. The method according to Clause 20, wherein using the motion information includes using the motion information to perform loop filtering during conversion.
[0588] Item 9 in the previous section provides additional examples and variations of the above terms.
[0589] 22. The method according to any one of clauses 1 to 21 further comprises: after performing the transformation of the current video block, selectively updating the motion vector prediction sub-table based on the update rule.
[0590] 23. The method described in Clause 22, wherein the rule specifies that motion vector predictors based on intra-block copy history are not updated.
[0591] 24. The method according to Clause 22, wherein the rule specifies that motion vector predictions based on intra-block copy history are updated using motion vectors of a specific sub-region.
[0592] Item 10 in the previous section provides additional examples and variations of the above terms.
[0593] 25. The method described in Clause 22, wherein the rule specifies that the history-based motion vector prediction subtable used for non-intra-block copy mode shall not be updated.
[0594] Item 11 in the previous section provides additional examples and variations of the above terms.
[0595] 26. The method according to any one of Clauses 1 to 25, wherein the conversion includes suppressing deblocking of the current video block during the conversion.
[0596] 27. The method according to any one of Clauses 1 to 25, wherein the conversion is performed by disabling another video encoding / decoding method during the execution of the method.
[0597] 28. The method according to Clause 27, wherein the other video coding / decoding method includes sub-block transform method, affine motion prediction method, multi-reference line intra-frame prediction method, matrix-based intra-frame prediction method, symmetric motion vector difference (MVD) coding / decoding, merge with MVD decoder-side motion derivation / refinement, bidirectional optical flow, simplified quadratic transform or multi-transform set coding / decoding method.
[0598] 29. The method according to any one of Clauses 1 to 28, wherein the use of the method during the conversion corresponds to signaling at the stripe level, picture level, sequence level, slice group level, slice level, brick level, codec tree unit level, codec tree block level, codec unit level, prediction unit level, or transform unit level in the bitstream representation.
[0599] Items 12, 13, and 14 in the previous section provide additional examples and variations of the above terms.
[0600] 30. The method according to any one of clauses 1 to 29, wherein wedge segmentation includes dividing the current video block into a plurality of portions, the boundaries of which have at least one horizontal portion and at least one vertical portion converging at the corners.
[0601] 31. The method according to any one of clauses 1 to 30, wherein triangular segmentation includes dividing the current video block into a plurality of parts having boundaries along non-horizontal and non-vertical directions.
[0602] 32. The method according to any one of clauses 1 to 31, wherein the transformation includes generating a bitstream representation from the current video block.
[0603] 33. The method according to any one of clauses 1 to 31, wherein the transformation includes generating samples of the current video block from the bitstream representation.
[0604] 34. A video processing apparatus, comprising a processor configured to implement the method described in any one or more of clauses 1 to 33.
[0605] 35. A computer-readable medium having code stored thereon, which, when executed, causes a processor to perform the methods described in any one or more of clauses 1 to 33.
[0606] Figure 25This is a flowchart of an exemplary method 2500 for video processing. Method 2500 includes: at 2502, determining, for a block of video and a bitstream representation of that block, applying at least one of an intra-block copy (IBC) mode, an intra-frame mode, an inter-frame mode, and a palette mode to a plurality of sub-regions of the block, wherein the block is divided into two or more triangular or wedge-shaped sub-regions; and at 2504, performing the conversion based on the determination.
[0607] In some examples, the IBC mode uses reference samples from at least one of the current strip, slice, brick, picture, and other video units, including codec tree unit (CTU) rows.
[0608] In some examples, the block is divided into two triangular sub-regions by applying diagonal or anti-diagonal partitioning to the block.
[0609] In some examples, the block is divided into four triangular sub-regions by applying both diagonal and anti-diagonal divisions.
[0610] In some examples, the block is divided into two wedge-shaped sub-regions with a predetermined shape.
[0611] In some examples, when the block is divided into multiple sub-regions, all sub-regions are encoded and decoded using the IBC mode, wherein at least two of the multiple sub-regions are encoded and decoded using different motion vectors (MV).
[0612] In some examples, the motion vector difference (MVD) of the MV or the MV derived from the motion vector prediction sub-region is encoded and decoded.
[0613] In some examples, predictive encoding and decoding of motion vectors of sub-regions relative to motion vectors is used to predict the motion vectors of sub-regions.
[0614] In some examples, motion vector predictors are derived from the regular IBC Advanced Motion Vector Prediction (AMVP) candidate list.
[0615] In some examples, predictive encoding and decoding of motion vectors of one sub-region relative to another sub-region is used.
[0616] In some examples, multiple sub-regions share a single AMVP candidate list.
[0617] In some examples, multiple sub-regions use different AMVP candidate lists.
[0618] In some examples, the motion vectors of subregions are inherited or derived from MV candidate indices.
[0619] In some examples, the motion vectors of the sub-regions are derived from the regular IBC merge candidate list.
[0620] In some examples, the motion vectors of the sub-regions are derived from the MV candidate list, which differs from the regular IBC merge candidate list construction process.
[0621] In some examples, the MV candidate list is checked for adjacent or non-adjacent blocks in different spatial domains.
[0622] In some examples, candidate indices are signaled or inferred.
[0623] In some examples, predictive encoding and decoding of candidate indices are applied.
[0624] In some examples, the candidate index of one subregion is predicted from the candidate index of another subregion.
[0625] In some examples, multiple sub-regions share a single merge candidate list.
[0626] In some examples, multiple sub-regions use different merge candidate lists.
[0627] In some examples, the motion vector of one sub-region is inherited or derived from the MV candidate index; the MV and / or MVD of another sub-region are obtained by encoding and decoding.
[0628] In some examples, the MV for a subregion is derived from a list of MV candidates.
[0629] In some examples, the MV candidate list is the regular IBC merge candidate list.
[0630] In some examples, the top M candidates in the candidate list are used.
[0631] In some examples, when the block is divided into multiple sub-regions, at least one sub-region is encoded and decoded using IBC mode, while another sub-region is encoded and decoded using non-IBC mode.
[0632] In some examples, this other sub-region is encoded and decoded using intra-frame mode.
[0633] In some examples, this other sub-region is encoded and decoded using inter-frame mode.
[0634] In some examples, this other sub-region is encoded and decoded using a palette mode.
[0635] In some examples, this other sub-region is encoded and decoded using Pulse Code Modulation (PCM) mode.
[0636] In some examples, this other sub-region is encoded and decoded using Residual Differential Pulse Codec Modulation (RDPCM) mode.
[0637] In some examples, the motion vectors of the sub-regions encoded and decoded using the IBC mode are obtained in the same way as those of the upper sub-regions encoded and decoded using the IBC mode.
[0638] In some examples, when the block is divided into multiple sub-regions, at least one of the sub-regions is encoded and decoded using intra-frame mode, while another sub-region is encoded and decoded using non-intra-frame mode.
[0639] In some examples, this other sub-region is encoded and decoded using inter-frame mode.
[0640] In some examples, motion vectors for sub-regions of inter-frame encoding are obtained using a block-based approach for regular inter-frame encoding.
[0641] In some examples, when the block is divided into multiple sub-regions, all sub-regions are encoded and decoded using a palette pattern, wherein at least two of the sub-regions are encoded and decoded using different palettes.
[0642] In some examples, when the block is divided into multiple sub-regions, at least one sub-region is encoded and decoded using a palette mode, while another sub-region is encoded and decoded using a non-palette mode.
[0643] In some examples, the method is applied to one or more specific components.
[0644] In some examples, this particular component is the luminance component.
[0645] In some examples, when the method is applied to one color component but not to another, the corresponding block in the other color component is encoded or decoded as an undivided block.
[0646] In some examples, the encoding / decoding method for the entire block in another color component is predefined and is one of IBC mode, inter-frame mode, and intra-frame mode.
[0647] In some examples, the method is applied to all components.
[0648] In some examples, the chroma blocks of the video are divided using the same partitioning pattern as the luminance components.
[0649] In some examples, the chroma blocks of the video are divided using a different partitioning pattern than that of the luminance component.
[0650] In some examples, when dividing a block into multiple sub-regions, a prediction for each sub-region is first generated, and all predictions are used to obtain the final predicted block.
[0651] In some examples, when a block is divided into multiple sub-regions, intermediate prediction blocks using information from each sub-region are generated, and the final prediction block is obtained by weighted averaging of the intermediate prediction blocks.
[0652] In some examples, equal weights are used.
[0653] In some examples, unequal weights are applied.
[0654] In some examples, one or more sets of weights are predefined to combine intermediate prediction blocks.
[0655] In some examples, IBC mode can be treated as inter-frame mode, and combined intra-frame and inter-frame prediction (CIIP) weights can be applied to combine intra-frame prediction and IBC prediction.
[0656] In some examples, the weight of a sample depends on its relative position within the current block.
[0657] In some examples, the weights on a sample point depend on the sample point's location relative to the sub-region's edge.
[0658] In some examples, the weights depend on the codec information of the current block, which includes at least one of the intra-prediction mode, block dimension, color components, and color format.
[0659] In some examples, when applying the above method, the motion information of a sub-region is stored as the motion information of the entire block.
[0660] In some examples, when applying the method, motion information of a sub-region is stored for each basic unit, and the basic unit has a minimum CU size.
[0661] In some examples, for a basic unit that covers multiple sub-regions, a set of motion information is selected or derived and stored.
[0662] In some examples, the stored motion information is utilized during the loop filtering process.
[0663] In some examples, the stored motion information is utilized in the encoding and decoding of subsequent blocks.
[0664] In some examples, the motion vector prediction (HMVP) table based on IBC history cannot be updated when these methods are applied.
[0665] In some examples, one or more motion vectors from the motion vectors of the sub-regions used for IBC encoding and decoding are used to update the IBC HMVP table.
[0666] In some examples, when these methods are applied, the motion vector prediction (HMVP) table, which is not based on IBC history, cannot be updated.
[0667] In some examples, when these methods are applied, one or more motion vectors from the motion vectors of the sub-regions used for inter-frame encoding and decoding are used to update the non-IBC HMVP table.
[0668] In some examples, the loop filtering process used for the block depends on the methods used.
[0669] In some examples, samples in blocks encoded or decoded using these methods are not filtered.
[0670] In some examples, blocks encoded or decoded using these methods are treated in a manner similar to blocks encoded or decoded using regular IBC encoding or decoding.
[0671] In some examples, certain encoding / decoding methods are disabled for blocks that are encoded or decoded using these methods.
[0672] In some examples, these encoding / decoding methods include one or more of the following: subblock transform, affine motion prediction, multi-reference line intra-frame prediction, matrix-based intra-frame prediction, symmetric MVD encoding / decoding, merge with MVD decoder-side motion derivation / refinement, bidirectional optical flow, simplified quadratic transform, and a set of multiple transforms.
[0673] In some examples, signaling notifications or instructions for the use of these methods and / or weighting values can be derived in at least one of the following video unit levels: sequence, picture, strip, slice group, slice, brick, CTU, CTB, CU, PU, TU, and other video unit levels.
[0674] In some examples, the above method is treated as a special IBC mode.
[0675] In some examples, if a block is encoded and decoded in IBC mode, the signaling notification or derivation uses the conventional block-based IBC method or further instructions of the above methods.
[0676] In some examples, subsequent IBC codecs use the motion information of the current block as MV predictors.
[0677] In some examples, subsequent IBC encoding / decoding blocks are not allowed to use the motion information of the current block as MV predictors.
[0678] In some examples, the above method is treated as a special triangle pattern.
[0679] In some examples, if a block is encoded or decoded in a triangular pattern, the signaling notifies or derives further instructions to use the conventional triangular prediction pattern (TPM) method or the above methods.
[0680] In some examples, the above methods are used as new prediction patterns.
[0681] In some examples, the allowed modes are further expanded to include intra-frame mode, inter-frame mode, and IBC mode to include new prediction modes.
[0682] In some examples, whether and / or how to apply the above methods depends on the following information:
[0683] a. Messages notified by signaling in DPS, SPS, VPS, PPS, APS, image header, strip header, slice header, maximum codec unit (LCU), codec unit (CU), LCU line, LCU group, TU, PU block, and video codec unit;
[0684] b. The location of CU, PU, TU, block, and video codec unit;
[0685] c. The block dimension of the current block and / or its neighboring blocks;
[0686] d. The block shape of the current block and / or its neighboring blocks;
[0687] e. Intra-frame mode of the current block and / or its adjacent blocks;
[0688] f. The motion vector or block vector of its adjacent blocks;
[0689] g. Indications including color formats of 4:2:0 or 4:4:4;
[0690] h. Encoding / decoding tree structure; and
[0691] i. Strip or slice type and / or image type.
[0692] In some examples, this conversion is represented by bitstreams to generate video blocks.
[0693] In some examples, the conversion is performed by generating a bitstream representation from blocks of video.
[0694] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagated signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode and decode information for transmission to a suitable receiver device.
[0695] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites and interconnected via a communication network.
[0696] The processes and logic flows described in this specification can be executed by one or more programmable processors executing one or more computer programs, thereby performing functions by manipulating input data and generating outputs. These processes and logic flows can also be executed by dedicated logic circuitry, and the apparatus can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[0697] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any kind of digital computer. Generally, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to one or more mass storage devices, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0698] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of specific features of particular embodiments of a particular technology. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features from the claimed combination may be removed from that combination in certain circumstances, and the claimed combination may involve sub-combinations or variations thereof.
[0699] Similarly, although the operations are shown in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in a sequential order or in the specific order shown, or as requiring all of the shown operations to achieve the desired result. Furthermore, the division of various system components in the embodiments described in this patent document should not be construed as requiring such division in all embodiments.
[0700] Only some implementation methods and examples have been described. Other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.
Claims
1. A method for video processing, comprising: determining, for a conversion between a block of a video and a bitstream of the block, to apply at least one of an intra block copy (IBC) mode, an intra mode, an inter mode, and a palette mode to a plurality of sub-regions of the block, wherein the block is partitioned into two or more triangular or wedge-shaped sub-regions; and performing the conversion based on the determination; wherein, when the block is partitioned into the plurality of sub-regions, all sub-regions are coded with the IBC mode, wherein at least two of the plurality of sub-regions are coded with different motion vectors (MVs); or when the block is partitioned into the plurality of sub-regions, at least one of the plurality of sub-regions is coded with the IBC mode and another sub-region is coded with a non-IBC mode; or when the block is partitioned into the plurality of sub-regions, at least one of the plurality of sub-regions is coded with the intra mode and another sub-region is coded with a non-intra mode; or when the block is partitioned into the plurality of sub-regions, all sub-regions are coded with the palette mode, wherein at least two of the sub-regions are coded with different palettes; or when the block is partitioned into the plurality of sub-regions, at least one of the plurality of sub-regions is coded with the palette mode and another sub-region is coded with a non-palette mode.
2. The method of claim 1, wherein, the IBC mode uses reference samples within at least one of a current slice, tile, brick, picture, other video unit including a coding tree unit (CTU) row.
3. The method of claim 1, wherein, the block is partitioned into two triangular sub-regions by applying a diagonal partition or an anti-diagonal partition to the block.
4. The method of claim 1, wherein, the block is partitioned into four triangular sub-regions by applying both a diagonal partition and an anti-diagonal partition to the block.
5. The method of claim 1, wherein, the block is partitioned into two wedge-shaped sub-regions with a predetermined shape.
6. The method of claim 1, wherein, in response to all sub-regions being coded with the IBC mode, motion vector differences (MVDs) of MVs or MVs derived from motion vector predictors of sub-regions are coded.
7. The method of claim 6, wherein, predictive coding of a motion vector of a sub-region relative to a motion vector predictor is used.
8. The method of claim 6, wherein, the motion vector predictor is derived from a regular IBC advanced motion vector prediction (AMVP) candidate list.
9. The method of claim 6, wherein, predictive coding of a motion vector of one sub-region relative to another sub-region is used.
10. The method of claim 6, wherein, the plurality of sub-regions share a single AMVP candidate list.
11. The method of claim 6, wherein, the plurality of sub-regions use different AMVP candidate lists.
12. The method of claim 1, wherein, in response to all sub-regions being coded with the IBC mode, motion vectors of at least two sub-regions are inherited or derived from a MV candidate index.
13. The method of claim 12, wherein, the motion vectors of the at least two sub-regions are derived from a regular IBC merge candidate list.
14. The method of claim 13, wherein, the motion vectors of the at least two sub-regions are derived from a MV candidate list that is different from the regular IBC merge candidate list construction process.
15. The method of claim 14, wherein, different spatially neighboring, contiguous or non-contiguous blocks are checked in the MV candidate list.
16. The method of claim 12, wherein, the candidate index is signaled or derived.
17. The method of claim 16, wherein, applying the predictive coding of the candidate index.
18. The method of claim 17, wherein, The candidate index of one sub-region is predicted from the candidate index of another sub-region.
19. The method of claim 12, wherein, The plurality of sub-regions share a single merge candidate list.
20. The method of claim 12, wherein, The plurality of sub-regions use different merge candidate lists.
21. The method of claim 1, wherein, In response to all sub-regions being coded with the IBC mode, the motion vector of one sub-region is inherited or derived from a MV candidate index; the MV and / or MVD of another sub-region is coded.
22. The method of claim 1, wherein, In response to all sub-regions being coded with the IBC mode, the MV of at least two sub-regions is derived from a MV candidate list.
23. The method of claim 22, wherein, The MV candidate list is a regular IBC merge candidate list.
24. The method of claim 22, wherein, The first M candidates in the candidate list are used.
25. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the IBC mode, another sub-region being coded with a non-IBC mode, the another sub-region is coded with an intra mode.
26. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the IBC mode, another sub-region being coded with a non-IBC mode, the another sub-region is coded with an inter mode.
27. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the IBC mode, another sub-region being coded with a non-IBC mode, the another sub-region is coded with a palette mode.
28. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the IBC mode, another sub-region being coded with a non-IBC mode, the another sub-region is coded with a pulse coded modulation (PCM) mode.
29. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the IBC mode, another sub-region being coded with a non-IBC mode, the another sub-region is coded with a residual differential pulse coded modulation (RDPCM) mode.
30. The method of any one of claims 24-29, wherein, The motion vector of the sub-region coded with the IBC mode is obtained in the same way as the above sub-region coded with the IBC mode.
31. The method of claim 1, wherein, In response to at least one of the plurality of sub-regions being coded with the intra mode, another sub-region being coded with a non-intra mode, the another sub-region is coded with an inter mode.
32. The method of claim 31, wherein, The motion vector of the sub-region coded with the inter mode is obtained in the way for a block coded with regular inter coding.
33. The method of claim 1, wherein the method is applied to one or more specific components.
34. The method of claim 33, wherein, The specific component is the luma component.
35. The method of claim 33, wherein, In applying the method to one color component and not to another color component, the corresponding block in the another color component is coded as an entire block without partitioning.
36. The method of claim 35, wherein, The coding method for the entire block in the another color component is predefined and is one of the IBC mode, the inter mode and the intra mode.
37. The method of claim 1, wherein the method is applied to all components.
38. The method of claim 34, wherein, The chroma blocks of the video are divided according to the same partitioning pattern as the luminance components.
39. The method of claim 34, wherein, The chroma blocks of the video are divided using a different partitioning mode than that of the luminance component.
40. The method of claim 1, wherein, When dividing a block into the multiple sub-regions, a prediction for each sub-region is first generated, and all the predictions are used to obtain the final predicted block.
41. The method of claim 1, wherein, When dividing a block into the multiple sub-regions, intermediate prediction blocks using information from each sub-region are generated, and a final prediction block is obtained by weighted averaging of the intermediate prediction blocks.
42. The method of claim 41, wherein, Use equal weights.
43. The method of claim 41, wherein, Apply unequal weights.
44. The method of claim 41, wherein, Predefine one or more sets of weights to combine intermediate prediction blocks.
45. The method of claim 41, wherein, The IBC mode is treated as the inter-frame mode, and combined intra-frame and inter-frame prediction CIIP weights are applied to combine intra-frame prediction and IBC prediction.
46. The method of claim 41, wherein, The weight of a sample depends on its relative position within the current block.
47. The method of claim 41, wherein, The weight of a sample depends on its position relative to the edge of the sub-region.
48. The method of claim 42, wherein, The weight depends on the encoding / decoding information of the current block, which includes at least one of the following: intra-prediction mode, block dimension, color components, and color format.
49. The method of claim 1, wherein, When applying the method, the motion information of a sub-region is stored as the motion information of the entire block.
50. The method of claim 1, wherein, When applying the method, motion information of a sub-region is stored for each basic unit, and the basic unit has a minimum CU size.
51. The method of claim 50, wherein, For a basic unit covering the multiple sub-regions, select or derive a set of motion information and store the set of motion information.
52. The method of claim 49, wherein, The stored motion information is utilized during the loop filtering process.
53. The method of claim 49, wherein, The stored motion information is used in the encoding and decoding of subsequent blocks.
54. The method of claim 1, wherein, When applying the method, the motion vector prediction HMVP table based on IBC history cannot be updated.
55. The method of claim 1, wherein, One or more motion vectors from the motion vectors of the sub-regions used for IBC encoding and decoding are used to update the IBC HMVP table.
56. The method of claim 1, wherein, When applying the method, the motion vector prediction HMVP table that is not based on IBC history cannot be updated.
57. The method of claim 1, wherein, When the method is applied, one or more motion vectors from the motion vectors of the sub-regions used for inter-frame encoding and decoding are used to update the non-IBC HMVP table.
58. The method of claim 1, wherein, The loop filtering process used for the block depends on the method used.
59. The method of claim 58, wherein, Samples in blocks encoded or decoded using the method described are not filtered.
60. The method of claim 58, wherein, Blocks encoded or decoded using the method are treated in a manner similar to those used for conventional IBC encoding and decoding.
61. The method of claim 1, wherein, For blocks encoded or decoded using the method described above, certain encoding / decoding methods are disabled.
62. The method of claim 61, wherein, Some of the encoding and decoding methods include one or more of the following: sub-block transform, affine motion prediction, multi-reference line intra-frame prediction, matrix-based intra-frame prediction, symmetric MVD encoding and decoding, merge with MVD decoder-side motion derivation / refinement, bidirectional optical flow, simplified quadratic transform, and multiple transform sets.
63. The method of claim 1, wherein, The signaling notification or indication of the use of the method and / or weighting value is derived in operation in at least one of the following video unit levels: sequence, picture, strip, slice group, slice, brick, CTU, CTB, CU, PU, TU, and other video unit level.
64. The method of claim 63, wherein, The method is treated as a special IBC mode.
65. The method of claim 64, wherein, If a block is coded in the IBC mode, it is signaled or derived that a regular block-based IBC method or a further indication of the method is used.
66. The method of claim 64, wherein, A subsequent IBC coded block is not allowed to use the motion information of the current block as MV predictor.
67. The method of claim 64, wherein, A subsequent IBC coded block is not allowed to use the motion information of the current block as MV predictor.
68. The method of claim 63, wherein, The method is treated as a special triangle mode.
69. The method of claim 68, wherein, If a block is coded in the triangle mode, it is signaled or derived that a regular triangle prediction mode TPM method or a further indication of the method is used.
70. The method of claim 63, wherein, The method is treated as a new prediction mode.
71. The method of claim 70, wherein, The allowed modes including the intra mode, the inter mode and the IBC mode are further extended to include the new prediction mode.
72. The method of claim 1, wherein, Whether and / or how the method is applied depends on the following information: a. a message signaled in a DPS, SPS, VPS, PPS, APS, picture header, slice header, tile group header, largest coding unit (LCU), coding unit (CU), LCU row, LCU group, TU, PU block, video coding unit; b. the location of the CU, PU, TU, block, video coding unit; c. the block dimension of the current block and / or its neighboring blocks; d. the block shape of the current block and / or its neighboring blocks; e. the intra mode of the current block and / or its neighboring blocks; f. the motion vector or block vector of its neighboring blocks; g. an indication of the color format including one of 4:2:0, 4:4:4; h. the coding tree structure; and i. the slice or tile group type and / or the picture type.
73. The method of any one of claims 1-29, 31-72, wherein, The conversion includes generating the bitstream from the blocks of the video.
74. The method of any one of claims 1-29, 31-72, wherein, The conversion includes generating the blocks of the video from the bitstream.
75. An apparatus in a video system, comprising a processor and a non-transitory memory having instructions located thereon, wherein, The instructions, when executed by the processor, cause the processor to implement the method according to any one of claims 1 to 74.
76. A computer program product stored on a non-transitory computer readable medium, wherein, The computer program product includes program code that, when executed by a processor, implements the method according to any one of claims 1 to 74. The computer program product includes program code that, when executed by a processor, implements the method according to any one of claims 1 to 74.