Sub-block-based intra-block copying and interaction between different codec tools
By employing sub-block-level intra-frame block copying and interaction techniques between different encoding and decoding tools at the video block level, the problem of insufficient efficiency in intra-frame block copying and encoding/decoding tool interaction in existing technologies is solved, achieving more efficient video encoding and decoding effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOUYIN VISION CO LTD
- Filing Date
- 2020-06-08
- Publication Date
- 2026-05-26
AI Technical Summary
Existing video codec standards struggle to effectively utilize intra-frame block copying and the interaction between different codec tools when processing video blocks, resulting in insufficient encoding efficiency and decoding quality.
By employing sub-block-level intra-block copying and interaction techniques between different encoding and decoding tools, the video block is divided into multiple sub-blocks, and modified intra-block copying encoding and decoding techniques, inter-frame encoding and decoding techniques, and palette encoding and decoding modes are used, combined with motion vectors and reference sample information to optimize the encoding process.
It improves the efficiency and quality of video encoding and decoding, reduces data redundancy, and enhances video processing performance.
Smart Images

Figure CN113940082B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] In accordance with applicable patent law and / or the rules of the Paris Convention, this application aims to promptly claim priority and benefit to International Patent Application No. PCT / CN2019 / 090409, filed June 6, 2019. For all purposes under U.S. law, the entire disclosure of the foregoing application is incorporated herein by reference as part of the disclosure of this application. Technical Field
[0003] This document covers video and image encoding and decoding technologies. Background Technology
[0004] Digital video accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention
[0005] The disclosed techniques can be used by video or image decoder or encoder embodiments to encode or decode video bitstream intra-block copying segmentation techniques at the sub-block level.
[0006] In one example aspect, a method for video processing is disclosed. The method includes determining that the current block of a video is divided into multiple sub-blocks for a conversion between the current block of the video and its bitstream representation. At least one of the multiple blocks is encoded using a modified intra-block copy (IBC) encoding / decoding technique, wherein the technique uses reference samples from one or more video regions of the current block's current image. The method also includes performing a conversion based on this determination.
[0007] In another example, a method for video processing is disclosed. The method includes determining that the current block of video is divided into multiple sub-blocks for a conversion between the current block and the bitstream representation of the video. Each sub-block is encoded or decoded in a codec representation using a corresponding encoding / decoding technique, depending on the mode. The method also includes performing a conversion based on this determination.
[0008] In another example, a method for video processing is disclosed. This method includes a transformation between a current block of video and a bitstream representation of the video, determining operations associated with a motion candidate list based on conditions relating to the characteristics of the current block. The motion candidate list is constructed either for encoding / decoding techniques or based on information from previously processed blocks of the video. The method also includes performing the transformation based on this determination.
[0009] In another example, a method for video processing is disclosed. The method includes determining, for a conversion between a current block of video and its bitstream representation, that the current block, encoded using a temporally information-based inter-frame encoding / decoding technique, is divided into multiple sub-blocks. At least one of the sub-blocks is encoded / decoded using a modified intra-block copy (IBC) encoding / decoding technique, wherein this technique uses reference samples from one or more video regions including the current block. The method also includes performing a conversion based on this determination.
[0010] In another example, a method for video processing is disclosed. This method includes determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and a bitstream representation of the current video block, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region, wherein the size of the sub-blocks is based on a partitioning rule, and using the sbIBC encoding / decoding mode for multiple sub-blocks to perform the conversion.
[0011] In another example, a method for video processing is disclosed. This method includes determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and its bitstream representation, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region; and performing the conversion using the sbIBC encoding / decoding mode for the multiple sub-blocks, wherein the conversion includes determining an initial motion vector (initMV) for a given sub-block, identifying a reference block from the initMV, and deriving MV information for the given sub-block using the motion vector (MV) information of the reference block.
[0012] In another example aspect, a method for video processing is disclosed. The method includes determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and its bitstream representation, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region; and performing the conversion using the sbIBC encoding / decoding mode for the multiple sub-blocks, wherein the conversion includes generating sub-block IBC candidates.
[0013] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing a first sub-block among the multiple sub-blocks using a sub-block intra-block codec (sbIBC) mode and processing a second sub-block among the multiple sub-blocks using an intra-frame codec mode.
[0014] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing all sub-blocks of the multiple sub-blocks using an intra-frame encoding / decoding mode.
[0015] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing all sub-blocks in the multiple sub-blocks using a palette encoding / decoding mode that uses a palette of representative pixel values for encoding and decoding each sub-block.
[0016] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing the first sub-block using a palette mode that uses a palette of representative pixel values for encoding and decoding, and processing the second sub-block using an intra-block copy encoding and decoding mode.
[0017] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing the first sub-block using a palette mode that uses a palette of representative pixel values for encoding and decoding, and processing the second sub-block using an intra-frame encoding / decoding mode.
[0018] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing a first sub-block among the multiple sub-blocks using a sub-block intra-block coding / decoding (sbIBC) mode and processing a second sub-block among the multiple sub-blocks using an inter-frame coding / decoding mode.
[0019] In another example, a method for video processing is disclosed. This method includes performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing a first sub-block among the multiple sub-blocks using a sub-block intra-frame encoding / decoding mode and processing a second sub-block among the multiple sub-blocks using an inter-frame encoding / decoding mode.
[0020] In another example aspect, a method for video processing is disclosed. The method includes determining to encode a current video block into a bitstream representation using the method described in any one of the preceding claims; and including information indicating the determination in the bitstream representation at the decoder parameter set level or sequence parameter set level or video parameter set level or picture parameter set level or picture header level or strip header level or slice header level or maximum codec unit level or codec unit level or maximum codec unit line level or LCU group level or transform unit level or prediction unit level or video codec unit level.
[0021] In another example aspect, a method for video processing is disclosed. The method includes deciding to encode a current video block into a bitstream representation based on encoding conditions using the method described in any of the preceding claims; and performing encoding using the method described in any of the preceding claims, wherein the conditions are based on one or more of the following: an encoding / decoding unit, a prediction unit, a transform unit, the current video block, or the position of the video encoding / decoding unit of the current video block.
[0022] In another example, a method for video processing is disclosed. This method includes determining a conversion between blocks in a video region and a bitstream representation of the video region using intra-block copy mode and inter-frame prediction mode; and performing the conversion using intra-block copy mode and inter-frame prediction mode for blocks in the video region.
[0023] In another example, a method for video processing is disclosed. This method includes, during the conversion between a current video block and its bitstream representation, performing a motion candidate list construction process and / or a table update process for updating a historically based table of motion vector prediction values, based on encoding / decoding conditions, and performing the conversion based on the motion candidate list construction process and / or the table update process.
[0024] In another example, the above method can be implemented by a video decoder device that includes a processor.
[0025] In another example, the above method can be implemented by a video encoder device that includes a processor.
[0026] In yet another example, these methods can be embodied in the form of processor-executable instructions and stored on a computer-readable program medium.
[0027] These and other aspects are further described in this document. Attached Figure Description
[0028] Figure 1 The derivation process for constructing the Merge candidate list is shown.
[0029] Figure 2An example of the location of the spatial merge candidate is shown.
[0030] Figure 3 An example of candidate pairs is shown that take into account for redundancy checks for spatial merge candidates.
[0031] Figure 4 Example locations of the second PU divided into N×2N and 2N×N segments are shown.
[0032] Figure 5 An example illustration of motion vector scaling for temporal Merge candidates is shown.
[0033] Figure 6 Examples of candidate positions for time-domain Merge candidates C0 and C1 are shown.
[0034] Figure 7 An example of combined bidirectional prediction of Merge candidates is shown.
[0035] Figure 8 An example of the derivation process for motion vector prediction candidates is shown.
[0036] Figure 9 An example illustration of motion vector scaling for spatial motion vector candidates is shown.
[0037] Figure 10 Simplified affine motion models with 4 parameters (left) and 6 parameters (right) are shown as examples.
[0038] Figure 11 An example of the affine motion vector field for each sub-block is shown.
[0039] Figure 12 Example candidate positions for the affine Merge pattern are shown.
[0040] Figure 13 An example of the modified Merge list construction process is shown.
[0041] Figure 14 An example of inter-frame prediction based on triangle segmentation is shown.
[0042] Figure 15 An example of applying the first weighting factor group to CU is shown.
[0043] Figure 16 An example of motion vector storage is shown.
[0044] Figure 17 An example of the final motion vector representation (UMVE) search process is shown.
[0045] Figure 18 An example of a UMVE search point is shown.
[0046] Figure 19 An example of MVD(0,1) mirrored between list 0 and list 1 is shown in DMVR.
[0047] Figure 20 The MV that can be checked in a single iteration is shown.
[0048] Figure 21 This is an example of intra-block copying.
[0049] Figure 22 This is a block diagram of an example video processing device.
[0050] Figure 23 This is a flowchart illustrating an example of a video processing method.
[0051] Figure 24 This is a block diagram of an example video processing system in which the disclosed technology can be implemented.
[0052] Figure 25 This is a flowchart representation of a method for video processing according to the present technology.
[0053] Figure 26 This is a flowchart representation of another method for video processing according to the present technology.
[0054] Figure 27 This is a flowchart representation of another method for video processing according to the present technology.
[0055] Figure 28 This is a flowchart representation of yet another method for video processing based on this technology. Detailed Implementation
[0056] This document provides a variety of techniques that decoders of image or video bitstreams can use to improve the quality of decompressed or decoded digital video or images. For the sake of brevity, the term "video" is used here to include both sequences of pictures (traditionally referred to as video) and individual images. Furthermore, video encoders can also implement these techniques during the encoding process to reconstruct decoded frames for further encoding.
[0057] For ease of understanding, chapter headings are used in this document, and the embodiments and techniques are not limited to the corresponding chapters. Thus, embodiments from one chapter can be combined with embodiments from other chapters.
[0058] 1. Abstract
[0059] This document relates to video codec technology. Specifically, it relates to intra-frame block copying (aka Current Picture Reference, CPR) codec. It can be applied to existing video codec standards (such as HEVC) or standards that are yet to be finalized (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs.
[0060] 2. Background
[0061] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.262 / MPEG-2 video and H.264 / MPEG-4 Advanced Video Coding (AVC) standards, as well as the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, a Joint Video Expert Team (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, with the goal of a 50% bitrate reduction compared to HEVC.
[0062] 2.1 Inter-frame prediction in HEVC / H.265
[0063] For an inter-frame encoding / decoding unit (CU), encoding / decoding can be performed using one prediction unit (PU) or two PUs, depending on the segmentation mode. Each inter-frame prediction PU has motion parameters for one or two lists of reference images. The motion parameters include motion vectors and reference image indices. The use of one of the two lists of reference images can also be signaled using `inter_pred_idc`. The motion vectors can be explicitly encoded as increments relative to the prediction values.
[0064] When encoding and decoding a CU in skip mode, a PU is associated with the CU and there are no significant residual coefficients, no encoded motion vector increments, or reference picture indices. A Merge mode is specified, thereby obtaining the motion parameters of the current PU from neighboring PUs, including spatial and temporal candidates. The Merge mode can be applied to any inter-frame prediction PU, not just skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector (more precisely, the motion vector difference (MVD) compared to the predicted motion vector value), the corresponding reference picture index for each reference picture list, and the reference picture list are explicitly signaled per PU. Such a mode is named Advanced Motion Vector Prediction (AMVP) in this disclosure.
[0065] When signaling indicates that one of two lists of reference images should be used, a PU is generated from a sample block. This is called "one-way prediction". One-way prediction applies to both P-strips and B-strips.
[0066] When signaling indicates that two reference image lists should be used, a PU is generated from two sample blocks. This is called "bidirectional prediction". Bidirectional prediction is only applicable to B-strips.
[0067] 2.1.1 List of Reference Images
[0068] In HEVC, the term inter-frame prediction is used to describe predictions derived from data elements (e.g., sample values or motion vectors) of reference images other than the currently decoded image. As in H.264 / AVC, images can be predicted from multiple reference images. The reference images used for inter-frame prediction are organized into one or more reference image lists. A reference index identifies which reference image in the list should be used to create the predicted signal.
[0069] A single list of reference images (list 0) is used for the P-strip, and two lists of reference images (list 0 and list 1) are used for the B-strip. It should be noted that the reference images included in lists 0 / 1 can be from past and future images in terms of capture / display order.
[0070] 2.1.2 Merge Mode
[0071] 2.1.2.1 Derivation of Merge Pattern Candidates
[0072] When predicting a PU using the Merge mode, indices pointing to entries in the Merge candidate list are parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized according to the following sequence of steps:
[0073] Step 1: Initial Candidate Derivation
[0074] Step 1.1: Spatial Candidate Derivation
[0075] Step 1.2: Redundancy check of airspace candidates
[0076] Step 1.3: Time-domain candidate derivation
[0077] Step 2: Add candidate insertions
[0078] Step 2.1: Create bidirectional prediction candidates
[0079] Step 2.2: Insert zero-motion candidates
[0080] exist Figure 1 These steps are also illustrated schematically. For spatial merge candidate derivation, up to four merge candidates are selected from candidates located at five different positions. For temporal merge candidate derivation, up to one merge candidate is selected from two candidates. Since it is assumed at the decoder that the number of candidates per PU is constant, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of merge candidates (MaxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, the index of the best merge candidate is encoded using truncated unary (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single merge candidate list, which is the same as the merge candidate list of a 2N×2N prediction unit.
[0081] The operations associated with the foregoing steps are described in detail below.
[0082] 2.1.2.2 Derivation of Airspace Candidates
[0083] In the derivation of the spatial Merge candidate, from the location located Figure 2 Up to four merge candidates are selected from the candidates at the positions depicted. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another strip or slice) or if it is intra-frame encoding / decoding. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only those at positions A1, B1, B0, A0, and B2 are considered. Figure 3The pairs linked by arrows are added to the list only if the candidate used for redundancy checking does not have the same motion information. Another source of duplicate motion information is a "second PU" associated with partitions different from 2N×2N. As an example, Figure 4 The second prediction unit (PU) is depicted for the N×2N and 2N×N cases. When the current PU is segmented into N×2N, the candidate at position A1 is not considered for list construction. In fact, adding this candidate would result in two prediction units having the same motion information, which is redundant for an encoding / decoding unit with only one PU. Similarly, when the current PU is segmented into 2N×N, position B1 is not considered.
[0084] 2.1.2.3 Time-domain candidate derivation
[0085] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the juxtaposed PU in the juxtaposed image. The list of reference images used for the derivation of the juxtaposed PU is displayed in the strip header, indicating that ground signaling will be used. Figure 5 The dashed lines in the diagram show the scaled motion vectors for obtaining the temporal merge candidate. These motion vectors are scaled from the motion vectors of the juxtaposed PU using POC distances tb and td, where tb is defined as the POC difference between the current image and its reference image, and td is defined as the POC difference between the juxtaposed image and its reference image. The reference image index for the temporal merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification. For the B-strip, two motion vectors are obtained, one for reference image list 0 and the other for reference image list 1, and they are combined to form a bidirectional prediction merge candidate.
[0086] 2.1.2.4 Juxtaposed Images and Juxtaposed PUs
[0087] When TMVP is enabled (e.g., slice_temporal_mvp_enabled_flag equals 1), the variable ColPic representing the juxtaposed images is deduced as follows:
[0088] - If the current stripe is a B stripe and the collocated_from_l0_flag of the signaling notification is equal to 0, then ColPic is set to equal to RefPicList1[collocated_ref_idx].
[0089] - Otherwise (slice_type equals B and collocated_from_l0_flag equals 1, or slice_type equals P), ColPic is set to equal RefPicList0[collocated_ref_idx].
[0090] Here, collocated_ref_idx and collocated_from_l0_flag are two syntax elements that can be used for signaling notification in the stripe header.
[0091] like Figure 6 The description describes the selection of a temporal candidate position between candidate C0 and C1 within the juxtaposed PU(Y) belonging to the reference frame. Position C1 is used if the PU at position C0 is unavailable, intra-frame encoded, or outside the current codec tree unit (CTU, aka LCU, maximum codec unit) row. Otherwise, position C0 is used in the derivation of the temporal merge candidate.
[0092] The relevant syntax elements are described below:
[0093] 7.3.6.1 General Striped Segment Header Syntax
[0094]
[0095] 2.1.2.5 Derivation of the MV of TMVP Candidates
[0096] More specifically, the following steps are performed to derive TMVP candidates:
[0097] (1) Set the reference image list X = 0, and the target reference image is the reference image with index 0 in list X (e.g., curr_ref). Call the derivation process of the juxtaposed motion vector to obtain the MV of list X pointing to curr_ref.
[0098] (2) If the current strip is a B strip, set the reference image list X = 1, and the target reference image is the reference image with index 0 in list X (e.g., curr_ref). Call the derivation process of the juxtaposed motion vector to obtain the MV of list X pointing to curr_ref.
[0099] The derivation of the juxtaposed motion vectors is described in the next subsection.
[0100] 2.1.2.5.1 Derivation of the juxtaposed motion vectors
[0101] For a juxtaposed block, it can be intra-frame or inter-frame encoded / decoded using unidirectional or bidirectional prediction. If it is intra-frame encoded / decoded, the TMVP candidate is set to unavailable.
[0102] If it is a one-way prediction from list A, then the motion vector of list A is scaled to the target reference image list X.
[0103] If it is a bidirectional prediction, and the target reference image list is X, then the motion vectors of list A are scaled to the target reference image list X, and A is determined according to the following rules:
[0104] - If no reference image has a larger POC value compared to the current image, then A is set to be equal to X.
[0105] Otherwise, A is set to equal collocated_from_l0_flag.
[0106] Some relevant descriptions are included below:
[0107] 8.5.3.2.9 Derivation of the juxtaposed motion vector
[0108] The input to this process is:
[0109] – The variable currPb specifies the current prediction block.
[0110] – The variable colPb specifies the juxtaposition prediction block within the juxtaposition image specified by ColPic.
[0111] – Luminance position (xColPb, yColPb): Specifies the top left luminance sample of the juxtaposed luminance prediction block specified by colPb relative to the top left luminance sample of the juxtaposed image specified by ColPic.
[0112] – Refer to the index refIdxLX, where X is 0 or 1.
[0113] The output of this process is:
[0114] – Motion vector prediction mvLXCol,
[0115] –Availability flag: availableFlagLXCol.
[0116] The variable currPic specifies the current image.
[0117] The arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] are set to be equal to the PredFlagL0[x][y], MvL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively, and the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] are set to be equal to the PredFlagL1[x][y], MvL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.
[0118] The variables mvLXCol and availableFlagLXCol are derived as follows:
[0119] – If colPb is encoded or decoded in intra-prediction mode, both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0120] Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0121] – If predFlagL0Col[xColPb][yColPb] equals 0, then mvCol, refIdxCol, and listCol are set to equal mvL1Col[xColPb][yColPb], refIdxL1Col[xColPb][yColPb], and L1, respectively.
[0122] Otherwise, if predFlagL0Col[xColPb][yColPb] equals 1 and predFlagL1Col[xColPb][yColPb] equals 0, then mvCol, refIdxCol, and listCol are set to mvL0Col[xColPb][yColPb], refIdxL0Col[xColPb][yColPb], and L0, respectively.
[0123] Otherwise (predFlagL0Col[xColPb][yColPb] equals 1, and predFlagL1Col[xColPb][yColPb] equals 1), perform the following assignment:
[0124] – If NoBackwardPredFlag equals 1, then mvCol, refIdxCol, and listCol are set to mvLXCol[xColPb][yColPb], refIdxLXCol[xColPb][yColPb], and LX, respectively.
[0125] Otherwise, mvCol, refIdxCol, and listCol are set to equal mvLNCol[xColPb][yColPb], refIdxLNCol[xColPb][yColPb], and LN, respectively, where N is the value of collocated_from_l0_flag.
[0126] Furthermore, mvLXCol and availableFlagLXCol are derived as follows:
[0127] – If LongTermRefPic(currPic,currPb,refIdxLX,LX) is not equal to LongTermRefPic(ColPic,colPb,refIdxCol,listCol), then both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0128] Otherwise, the variable availableFlagLXCol is set to 1, refPicListCol[refIdxCol] is set to the reference image listCol containing the stripes of the predicted block colPb in the juxtaposed image specified by ColPic, and the following applies:
[0129] colPocDiff=DiffPicOrderCnt(ColPic,refPicListCol[refIdxCol]) (2-1)
[0130] currPocDiff=DiffPicOrderCnt(currPic,RefPicListX[refIdxLX]) (2-2)
[0131] – If RefPicListX[refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is derived as follows:
[0132] mvLXCol=mvCol (2-3)
[0133] Otherwise, mvLXCol is derived as a scaled version of the motion vector mvCol, as follows:
[0134] tx=(16384+(Abs(td)>>1)) / td (2-4)
[0135] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (2-5)
[0136] mvLXCol=Clip3(-32768,32767,Sign(distScaleFactor*mvCol)*((Abs(distScaleFactor*mvCol)+127)>>8)) (2-6)
[0137] The derivation of td and tb is as follows:
[0138] td=Clip3(-128,127,colPocDiff) (2-7)
[0139] tb=Clip3(-128,127,currPocDiff) (2-8)
[0140] The definition of NoBackwardPredFlag is:
[0141] The variable NoBackwardPredFlag is derived as follows:
[0142] – If the DiffPicOrderCnt(aPic,CurrPic) of each image aPic in the current stripe's RefPicList0 or RefPicList1 is less than or equal to 0, then NoBackwardPredFlag is set to 1.
[0143] Otherwise, NoBackwardPredFlag is set to 0.
[0144] 2.1.2.6 Additional Candidate Insertion
[0145] In addition to the spatiotemporal merge candidate, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated by utilizing the spatiotemporal merge candidate. Combined bidirectional prediction merge candidates are only used for B-strips. Combined bidirectional prediction candidates are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of the other. If the two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. As an example, Figure 7 This describes the case where two candidates from the original list (on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create bidirectional predictive merge candidates that are added to the final list (on the right). There are many rules regarding combinations that are considered to generate these additional merge candidates.
[0146] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the MaxNumMergeCand capacity. These candidates have zero spatial displacements and a reference image index that starts at zero and increments whenever a new zero-motion candidate is added to the list. Finally, no redundancy checks are performed on these candidates.
[0147] 2.1.3 AMVP
[0148] AMVP utilizes the spatiotemporal correlation between motion vectors and neighboring PUs for explicit transmission of motion parameters. For each list of reference images, a motion vector candidate list is constructed by first checking the availability of temporally neighboring PU locations on the left and top, removing redundant candidates, and adding zero vectors to make the candidate list a constant length. The encoder can then select the best prediction from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate uses truncated unary coding. In this case, the maximum value to be encoded is 2 (see...). Figure 8 The following sections provide details of the derivation process for the motion vector prediction candidates.
[0149] 2.1.3.1 Derivation of AMVP Candidates
[0150] Figure 8 The derivation process of motion vector prediction candidates is summarized.
[0151] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. For the derivation of spatial motion vector candidates, based on the location as... Figure 2The motion vectors of each PU at the five different locations are depicted to ultimately derive two motion vector candidates.
[0152] For temporal motion vector candidate derivation, one motion vector candidate is selected from two candidates derived based on two different juxtaposition positions. After generating the first spatiotemporal candidate list, duplicate motion vector candidates in the list are removed. If the number of potential candidates is greater than two, motion vector candidates with reference image indices greater than 1 are removed from the associated reference image list. If the number of spatiotemporal motion vector candidates is less than two, additional zero motion vector candidates are added to the list.
[0153] 2.1.3.2 Candidate Spatial Motion Vectors
[0154] In the derivation of the spatial motion vector candidates, at most two candidates are considered from five potential candidates. These five potential candidates are selected from those located in the spatial domain, such as... Figure 2 The positions depicted are derived from the PUs, which are the same as the positions of the motion merge. The derivation order to the left of the current PU is defined as A0, A1 and scaled A0, scaled A1. The derivation order above the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, for each side, there are four cases that can be used as motion vector candidates, two of which do not require spatial scaling, and two of which do use spatial scaling. These four different cases are summarized as follows:
[0155] *No spatial scaling
[0156] (1) Same list of reference images and same index of reference images (same POC)
[0157] (2) Different lists of reference images but the same reference image (same POC)
[0158] Spatial scaling
[0159] (3) Same list of reference images but different reference images (different POCs)
[0160] (4) Different lists of reference images and different reference images (different POCs)
[0161] First, check for cases without spatial scaling, then check for spatial scaling. Regardless of the reference image list, consider spatial scaling when the Proof of Concept (POC) differs between the reference image of a neighboring PU and the reference image of the current PU. If all candidate PUs on the left are unavailable or intra-frame encoded / decoded, scaling of the upper motion vector is allowed to aid in the parallel derivation of the left and upper motion vectors. Otherwise, spatial scaling of the upper motion vector is not allowed.
[0162] like Figure 9 The described process involves scaling the motion vectors of neighboring PUs in a manner similar to temporal scaling during spatial scaling. The main difference is that a list of reference images and the index of the current PU are given as input; the actual scaling process is the same as that of temporal scaling.
[0163] 2.1.3.3 Candidate Motion Vectors in the Time Domain
[0164] Except for the derivation of the reference image index, all the procedures for deriving the temporal Merge candidate are the same as those for deriving the spatial motion vector candidate (see [link]). Figure 6 The reference image index is signaled to the decoder.
[0165] 2.2 Inter-frame prediction methods in VVC
[0166] Several new codec tools for improving inter-frame prediction have been developed, such as Adaptive Motion Vector Difference Resolution (AMVR) for Signaling Notification MVD, Merge with Motion Vector Difference (MMVD), Triangle Prediction Mode (TPM), Combined Intra-Inter-Frame Prediction (CIIP), Advanced TMVP (ATMVP, also known as SbTMVP), Affine Prediction Mode, Generalized Bidirectional Prediction (GBI), Decoder-Side Motion Vector Refinement (DMVR), and Bidirectional Optical Flow (BIO, also known as BDOF).
[0167] VVC supports three different Merge list construction processes:
[0168] (1) Sub-block Merge Candidate List: This includes ATMVP and Affine Merge candidates. A single Merge list construction process is shared for both Affine and ATMVP modes. Here, ATMVP and Affine Merge candidates can be added sequentially. The size of the sub-block Merge list is signaled in the stripe header and has a maximum value of 5.
[0169] (2) Regular Merge List: For inter-frame codec blocks, a shared Merge List construction process is used. Here, spatial / temporal Merge Candidates, HMVP, paired Merge Candidates, and zero-motion Candidates can be inserted sequentially. The size of the regular Merge List is indicated by signaling in the stripe header and has a maximum value of 6. MMVD, TPM, and CIIP rely on the regular Merge List.
[0170] (3) IBC Merge List: It is done in a manner similar to a regular Merge List.
[0171] Similarly, VVC supports three lists of AMVPs:
[0172] (1) Affine AMVP Candidate List
[0173] (2) Regular AMVP Candidate List
[0174] (3) IBC·AMVP candidate list: The same construction process as the IBC Merge list.
[0175] 2.2.1 Codec Block Structure in VVC
[0176] In VVC, a quadtree / binary tree / ternary tree (QT / BT / TT) structure is used to divide the image into square or rectangular blocks.
[0177] Besides QT / BT / TT, a separate tree (also known as a dual codec tree) is also used for I-frames in VVC. In the case of a separate tree, the codec block structure is signaled separately for the luma and chroma components.
[0178] In addition, except for blocks encoded and decoded using several specific encoding and decoding methods (such as intra-frame sub-segmentation prediction, where PU equals TU but is less than CU, and inter-frame encoded and decoded sub-block transformation, where PU equals CU but TU is less than PU), CU is set to equal PU and TU.
[0179] 2.2.2 Affine Prediction Mode
[0180] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). Although there are many types of motion in the real world, such as zooming in / out, rotation, perspective motion, and other irregular motions, VVC uses simplified affine transformation motion compensation prediction with 4-parameter and 6-parameter affine models. Figure 10 As shown, the affine motion field of the block is described by two control point motion vectors (CPMV) of a 4-parameter affine model and three CPMVs of a 6-parameter affine model.
[0181] The motion vector field (MVF) of the block is described by the following equations using the 4-parameter affine model in equation (1) (where the 4 parameters are defined as variables a, b, e, and f) and the 6-parameter affine model in equation (2) (where the 4 parameters are defined as variables a, b, c, d, e, and f):
[0182]
[0183]
[0184] Where (mv h 0,mv h 0) is the motion vector of the top-left control point, and (mv h1, MV h 1) is the motion vector of the upper right control point, and (mv h 2, MV h 2) is the motion vector of the lower left control point. All three motion vectors are called the control point motion vectors (CPMV). (x, y) represents the coordinates of the point relative to the upper left sample point within the current block, and (mv) h (x, y), mv v (x, y) is the motion vector derived for the sample located at (x, y). The CP motion vector can be signaled (e.g., in affine AMVP mode) or derived on the fly (e.g., in affine Merge mode). w and h are the width and height of the current block. In practice, division is implemented via right shift and rounding. In VTM, the representative point is defined as the center position of the sub-block; for example, when the coordinates of the top-left corner of the sub-block relative to the top-left sample within the current block are (xs, ys), the coordinates of the representative point are defined as (xs+2, ys+2). For each sub-block (e.g., 4×4 in VTM), the motion vector for the entire sub-block is derived using the representative point.
[0185] To further simplify motion compensation prediction, a sub-block-based affine transformation prediction was applied. To derive the motion vector for each M×N (M and N are set to 4 in the current VVC) sub-block, the motion vector of the center sample point of each sub-block (e.g., ...) is... Figure 11 The result (as shown) can be calculated according to equations (1) and (2) and rounded to 1 / 16 fractional precision. Then, a 1 / 16 pixel motion compensation interpolation filter can be applied to generate predictions for each sub-block with derived motion vectors. The affine mode introduces a 1 / 16 pixel interpolation filter.
[0186] After MCP, the high-precision motion vector of each sub-block is rounded and saved with the same precision as the standard motion vector.
[0187] 2.2.3 Merge of the entire block
[0188] 2.2.3.1 Constructing the Merge table from the regular Merge schema
[0189] 2.2.3.1.1 History-Based Motion Vector Prediction (HMVP)
[0190] Unlike the Merge list design, VVC uses a history-based motion vector prediction (HMVP) method.
[0191] The HMVP stores motion information from previous encoding / decoding operations. Motion information from previous encoded / decoded blocks is defined as HMVP candidates. Multiple candidate HMVPs are stored in a table called the HMVP table, which is dynamically maintained during the encoding / decoding process. The HMVP table is cleared when encoding / decoding of a new slice / LCU line / strip begins. Whenever an inter-frame encoded / decoded block and a non-sub-block, non-TPM mode exist, the associated motion information is added to the last entry of the table as a new HMVP candidate. Figure 12 The entire encoding and decoding process is described in the document.
[0192] 2.2.3.1.2 Standard Merge List Construction Process
[0193] The construction of a regular Merge list (for translational motion) can be summarized in the following steps:
[0194] Step 1: Derive the candidate airspace
[0195] Step 2: Insert HMVP candidate
[0196] Step 3: Insert pairwise average candidates
[0197] Step 4: Default motion candidates
[0198] HMVP candidates can be used in the AMVP and Merge candidate list building process. Figure 13 The modified Merge Candidate List Construction Process is described. When the Merge Candidate List is not full after the insertion of TMVP candidates, HMVP candidates stored in the HMVP table can be used to populate the Merge Candidate List. Considering that a block is generally highly correlated with its nearest neighbor in terms of motion information, HMVP candidates in the table are inserted in descending order of their indexes. The last entry in the table is added to the list first, and the first entry is added to the end. Similarly, redundancy removal is also applied to HMVP candidates. The Merge Candidate List Construction Process terminates once the total number of available Merge Candidates reaches the maximum number of Merge Candidates allowed for signaling notification.
[0199] Please note that all spatial / temporal / HMVP candidates should be encoded and decoded in non-IBC mode. Otherwise, they are not allowed to be added to the regular Merge candidate list.
[0200] The HMVP table contains up to five regular sports candidates, and each of them is unique.
[0201] 2.2.3.1.2.1 Pruning process
[0202] If the corresponding candidates used for redundancy checking do not have the same motion information, the candidate is simply added to the list. This comparison process is called the pruning process.
[0203] The pruning process among the airspace candidates depends on the use of TPM for the current block.
[0204] When the current block is not encoded or decoded in TPM mode (e.g., regular Merge, MMVD, CIIP), the HEVC pruning process (e.g., five-pronged pruning) of the spatial merge candidate is utilized.
[0205] 2.2.4 Triangle Prediction Pattern
[0206] In VVC, triangular splitting mode is supported for inter-frame prediction. Triangular splitting mode is only applied to CUs of 8x8 or larger that are encoded and decoded in Merge mode but not in MMVD or CIIP mode. For CUs that meet these conditions, signaling informs the CU level flag to indicate whether triangular splitting mode is applied.
[0207] When using this mode, with either diagonal or anti-diagonal division, the CU is evenly divided into two triangular segments, such as... Figure 14 As depicted, each triangle segment in the CU is used for inter-frame prediction using its own motion; only unidirectional prediction is allowed for each segment, meaning each segment has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, only two motion-compensated predictions are needed for each CU.
[0208] If the CU-level flag indicates that the current CU uses the triangle segmentation mode for encoding and decoding, a flag indicating the triangle segmentation direction (diagonal or anti-diagonal) and two Merge indices (one for each segment) are further signaled. After predicting each triangle segment, a blending process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the predicted signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangle segmentation mode is stored in 4x4 cells.
[0209] The regular Merge candidate list is reused for triangle segmentation Merge prediction without additional motion vector pruning. For each Merge candidate in the regular Merge candidate list, one and only one of its L0 or L1 motion vectors is used for triangle prediction. Furthermore, the order of selection of the L0 and L1 motion vectors is based on the parity of their Merge index. Using this scheme, the regular Merge list can be used directly.
[0210] 2.2.4.1 TPM Merge List Construction Process
[0211] In some embodiments, the standard Merge list construction process may include the following modifications:
[0212] (1) How the pruning process is performed depends on the TPM used for the current block.
[0213] - If the current block is not encoded or decoded using TPM, then invoke HEVC5 pruning applied to the spatial merge candidate.
[0214] Otherwise (if the current block is encoded / decoded using TPM), full pruning will be applied when adding new spatial merge candidates. That is, B1 is compared with A1; B0 is compared with A1 and B1; A0 is compared with A1, B1, and B0; and B2 is compared with A1, B1, A0, and B0.
[0215] (2) Whether to check motion information from B2 depends on the use of TPM for the current block.
[0216] - If the current block is not encoded or decoded using TPM, B2 will only be accessed and checked if there are fewer than 4 space merge candidates before checking B2.
[0217] Otherwise (if the current block is encoded / decoded using TPM), B2 is always accessed and checked before adding B2, regardless of how many available airspace merge candidates there are.
[0218] 2.2.4.2 Adaptive Weighting Process
[0219] After predicting each triangular prediction unit, an adaptive weighting process is applied to the diagonal edges between two triangular prediction units to derive the final prediction for the entire CU. The two sets of weighting factors are defined as follows:
[0220] The first weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luminance and chrominance samples, respectively;
[0221] The second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used for luminance and chrominance samples, respectively.
[0222] The weighting factor set is selected based on a comparison of the motion vectors of the two triangular prediction units. The second weighting factor set is used when any of the following conditions are true:
[0223] - The reference images for the two triangular prediction units are different from each other.
[0224] - The absolute value of the difference between the horizontal values of two motion vectors is greater than 16 pixels.
[0225] - The absolute value of the difference between the perpendicular values of two motion vectors is greater than 16 pixels.
[0226] Otherwise, use the first weighted factor group. Figure 15 An example is shown in the image.
[0227] 2.2.4.3 Motion Vector Storage
[0228] The motion vector of the triangular prediction unit ( Figure 16 Mv1 and Mv2 in the CU are stored in a 4×4 grid. For each 4×4 grid, depending on the grid's position within the CU, either a unidirectional or bidirectional predicted motion vector is stored. Figure 16 As shown, a 4×4 grid located in the unweighted region (i.e., not located at the diagonal edge) stores either a unidirectional predicted motion vector Mv1 or Mv2. Conversely, a 4×4 grid located in the weighted region stores a bidirectional predicted motion vector. The bidirectional predicted motion vector is derived from Mv1 and Mv2 according to the following rules:
[0229] (1) When Mv1 and Mv2 have motion vectors from different directions (L0 or L1), Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.
[0230] (2) When Mv1 and Mv2 both come from the same L0 (or L1) direction,
[0231] If the reference image for Mv2 is the same as an image in the L1 (or L0) reference image list, then Mv2 is scaled to that image. Mv1 and the scaled Mv2 are combined to form a bidirectional predicted motion vector.
[0232] If the reference image for Mv1 is the same as an image in the L1 (or L0) reference image list, then Mv1 is scaled to that image. The scaled Mv1 and Mv2 are combined to form a bidirectional predicted motion vector.
[0233] Otherwise, only Mv1 is stored for the weighted region.
[0234] 2.2.4.4 The syntax, semantics, and decoding process of the Merge pattern
[0235] Added changes are highlighted with underline, bold, and italic. Deleted parts are marked with [[]].
[0236] 7.3.5.1 General Strip Header Syntax
[0237]
[0238]
[0239] 7.3.7.5 Encoding / Decoding Unit Syntax
[0240]
[0241]
[0242] 7.3.7.7 Merge Data Syntax
[0243]
[0244]
[0245] 7.4.6.1 General Strip Header Semantics
[0246] `six_minus_max_num_merge_cand` specifies the maximum number of Merge Motion Vector Prediction (MVP) candidates supported from the strips subtracted from 6. The maximum number of Merge MVP candidates, `MaxNumMergeCand`, is derived as follows:
[0247] MaxNumMergeCand=6-six_minus_max_num_merge_cand (7-57)
[0248] The value of MaxNumMergeCand should be in the range of 1 to 6 (inclusive).
[0249] `five_minus_max_num_subblock_merge_cand` specifies the maximum number of subblock-based Merge Motion Vector Prediction (MVP) candidates supported from the strips subtracted from 5. When `five_minus_max_num_subblock_merge_cand` is not present, it is inferred to be equal to 5 - `sps_sbtmvp_enabled_flag`. The maximum number of subblock-based Merge MVP candidates, `MaxNumSubblockMergeCand`, is derived as follows:
[0250] MaxNumSubblockMergeCand=5-five_minus_max_num_subblock_merge_cand(7-58)
[0251] The value of MaxNumSubblockMergeCand should be in the range of 0 to 5 (inclusive).
[0252] 7.4.8.5 Semantics of Encoding / Decoding Units
[0253] A pred_mode_flag value of 0 indicates that the current codec unit is encoding / decoding in inter-frame prediction mode. A pred_mode_flag value of 1 indicates that the current codec unit is encoding / decoding in intra-frame prediction mode.
[0254] When pred_mode_flag does not exist, it is inferred as follows:
[0255] – If cbWidth equals 4 and cbHeight equals 4, then pred_mode_flag is inferred to be equal to 1.
[0256] Otherwise, when decoding an I stripe, pred_mode_flag is inferred to be equal to 1, and when decoding a P or B stripe, pred_mode_flag is inferred to be equal to 0.
[0257] For x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the variable CuPredMode[x][y] is derived as follows:
[0258] – If pred_mode_flag equals 0, then CuPredMode[x][y] is set to equal MODE_INTER.
[0259] Otherwise (pred_mode_flag equals 1), CuPredMode[x][y] is set to equal MODE_INTRA.
[0260] A pred_mode_ibc_flag value of 1 indicates that the current codec unit is encoding / decoding in IBC prediction mode. A pred_mode_ibc_flag value of 0 indicates that the current codec unit is not encoding / decoding in IBC prediction mode.
[0261] When pred_mode_ibc_flag does not exist, it is inferred as follows:
[0262] – If cu_skip_flag[x0][y0] equals 1, cbWidth equals 4, and cbHeight equals 4, then pred_mode_ibc_flag is inferred to be equal to 1.
[0263] Otherwise, if both cbWidth and cbHeight are equal to 128, then pred_mode_ibc_flag is inferred to be equal to 0.
[0264] Otherwise, when decoding an I-strip, pred_mode_ibc_flag is inferred to be equal to the value of sps_ibc_enabled_flag, and when decoding a P- or B-strip, it is equal to 0.
[0265] When pred_mode_ibc_flag equals 1, for x = x0..x0+cbWidth-1 and y = y0..y0+cbHeight-1, the variable CuPredMode[x][y] is set to equal MODE_IBC.
[0266] `general_merge_flag[x0][y0]` specifies whether the inter-frame prediction parameters of the current codec unit are inferred from neighboring inter-frame prediction segments. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0267] When general_merge_flag[x0][y0] does not exist, it is inferred as follows:
[0268] – If cu_skip_flag[x0][y0] equals 1, then general_merge_flag[x0][y0] is inferred to be equal to 1.
[0269] Otherwise, general_merge_flag[x0][y0] is inferred to be equal to 0.
[0270] mvp_l0_flag[x0][y0] specifies the index of the motion vector prediction value in list 0, where x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0271] When mvp_l0_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0272] mvp_l1_flag[x0][y0] has the same semantics as mvp_l0_flag, where l0 and list 0 are replaced by l1 and list 1, respectively.
[0273] inter_pred_idc[x0][y0] specifies, according to Table 7-10, whether list 0, list 1, or bidirectional prediction is used for the current codec unit. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0274] Table 7-10 – Name Association of Inter-Frame Prediction Modes
[0275]
[0276] When inter_pred_idc[x0][y0] does not exist, it is inferred to be equal to PRED_L0.
[0277] 7.4.8.7 Merge Data Semantics
[0278] The regular_merge_flag[x0][y0] value of 1 specifies the inter-frame prediction parameters used by the regular Merge mode to generate the current codec unit. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0279] When regular_merge_flag[x0][y0] does not exist, it is inferred as follows:
[0280] – If all of the following conditions are true, then regular_merge_flag[x0][y0] is inferred to be equal to 1:
[0281] –sps_mmvd_enabled_flag equals 0.
[0282] –general_merge_flag[x0][y0] equals 1.
[0283] –cbWidth*cbHeight equals 32.
[0284] Otherwise, regular_merge_flag[x0][y0] is inferred to be equal to 0.
[0285] The `mmvd_merge_flag[x0][y0]` value of 1 specifies the Merge mode with motion vector differences used to generate the inter-frame prediction parameters for the current codec unit. The array indices `x0` and `y0` specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0286] When mmvd_merge_flag[x0][y0] does not exist, it is inferred as follows:
[0287] – If all of the following conditions are true, then it is inferred that mmvd_merge_flag[x0][y0] equals 1:
[0288] –sps_mmvd_enabled_flag equals 1.
[0289] –general_merge_flag[x0][y0] equals 1.
[0290] –cbWidth*cbHeight equals 32.
[0291] –regular_merge_flag[x0][y0] equals 0.
[0292] Otherwise, mmvd_merge_flag[x0][y0] is inferred to be equal to 0.
[0293] mmvd_cand_flag[x0][y0] specifies whether the first (0) or second (1) candidate in the Merge candidate list is used with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0294] When mmvd_cand_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0295] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 7-12. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0296] Table 7-12 – Specification of MmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0]
[0297]
[0298] mmvd_direction_idx[x0][y0] specifies the index used to derive MmvdSign[x0][y0] as specified in Table 7-13. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0299] Table 7-13 – Specification of MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]
[0300] mmvd_direction_idx[x0][y0] MmvdSign[x0][y0][0] MmvdSign[x0][y0][1] 0 +1 0 1 -1 0 2 0 +1 3 0 -1
[0301] The two components of Merge, plus the MVD offset MmvdOffset[x0][y0], are derived as follows:
[0302] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0] (7-124)
[0303] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1] (7-125)
[0304] `merge_subblock_flag[x0][y0]` specifies whether the sub-block-based inter-frame prediction parameters of the current codec unit are inferred from neighboring blocks. Array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the considered codec block relative to the top left luminance sample of the image. When `merge_subblock_flag[x0][y0]` does not exist, it is inferred to be equal to 0.
[0305] merge_subblock_idx[x0][y0] specifies the merge candidate index of the merge candidate list based on subblocks, where x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0306] When merge_subblock_idx[x0][y0] does not exist, it is inferred to be equal to 0.
[0307] ciip_flag[x0][y0] specifies whether the combined inter-frame merge and intra-frame prediction are applied to the current codec unit. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0308] When ciip_flag[x0][y0] does not exist, it is inferred to be equal to 0.
[0309] When ciip_flag[x0][y0] equals 1, the variable IntraPredModeY[x][y] (where x = xCb..xCb+cbWidth-1 and y = yCb..yCb+cbHeight-1) is set to equal INTRA_PLANAR.
[0310] The variable MergeTriangleFlag[x0][y0] (which specifies whether motion compensation based on the triangle shape is used to generate prediction samples for the current codec unit when decoding B stripes) is derived as follows:
[0311] – MergeTriangleFlag[x0][y0] is set to 1 if all of the following conditions are true:
[0312] –sps_triangle_enabled_flag equals 1.
[0313] –slice_type equals B.
[0314] –general_merge_flag[x0][y0] equals 1.
[0315] –MaxNumTriangleMergeCand is greater than or equal to 2.
[0316] –cbWidth*cbHeight is greater than or equal to 64.
[0317] –regular_merge_flag[x0][y0] equals 0.
[0318] –mmvd_merge_flag[x0][y0] equals 0.
[0319] –merge_subblock_flag[x0][y0] equals 0.
[0320] –ciip_flag[x0][y0] equals 0.
[0321] Otherwise, MergeTriangleFlag[x0][y0] is set to 0.
[0322] `merge_triangle_split_dir[x0][y0]` specifies the splitting direction of the merge triangle pattern. The array indices x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0323] When merge_triangle_split_dir[x0][y0] does not exist, it is inferred to be equal to 0.
[0324] merge_triangle_idx0[x0][y0] specifies the first merge candidate index of the motion compensation candidate list based on the triangle shape, where x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0325] When merge_triangle_idx0[x0][y0] does not exist, it is inferred to be equal to 0.
[0326] merge_triangle_idx1[x0][y0] specifies the second Merge candidate index of the motion compensation candidate list based on the triangle shape, where x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0327] When merge_triangle_idx1[x0][y0] does not exist, it is inferred to be equal to 0.
[0328] merge_idx[x0][y0] specifies the merge candidate index in the merge candidate list, where x0 and y0 specify the position (x0, y0) of the top left luminance sample of the codec block under consideration relative to the top left luminance sample of the image.
[0329] When merge_idx[x0][y0] does not exist, it is inferred as follows:
[0330] – If mmvd_merge_flag[x0][y0] equals 1, then merge_idx[x0][y0] is inferred to be equal to mmvd_cand_flag[x0][y0].
[0331] Otherwise (mmvd_merge_flag[x0][y0] equals 0), merge_idx[x0][y0] is inferred to be equal to 0.
[0332] 2.2.4.4.1 Decoding process
[0333] In some embodiments, the decoding process is defined as follows:
[0334] 8.5.2.2 Derivation of the Luminance Motion Vector in Merge Mode
[0335] The procedure is invoked only when general_merge_flag[xCb][yCb] equals 1, where (xCb, yCb) specifies the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0336] The input to this process is:
[0337] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0338] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0339] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0340] The output of this process is:
[0341] The brightness motion vectors mvL0[0][0] and mvL1[0][0] with a precision of –1 / 16 fractional sample points.
[0342] –Refer to indices refIdxL0 and refIdxL1,
[0343] – The prediction list uses the flags predFlagL0[0][0] and predFlagL1[0][0].
[0344] – Bidirectional prediction weight index bcwIdx
[0345] –Merge candidate list mergeCandList.
[0346] The bidirectional prediction weight index bcwIdx is set to equal to 0.
[0347] The motion vectors mvL0[0][0] and mvL1[0][0], the reference indices refIdxL0 and refIdxL1, and the prediction utilization flags predFlagL0[0][0] and predFlagL1[0][0] are derived through the following ordered steps:
[0348] 1. Invoke the derivation procedure for spatial merge candidates from neighboring codec units as specified in Clause 8.5.2.4, taking the luma codec block position (xCb, yCb), luma codec block width cbWidth, and luma codec block height cbHeight as input, and outputting availability flags availableFlagA0, availableFlagA1, availableFlagB0, availableFlagB1, and availableFlagB2, and reference indices refIdxLXA0, refIdxL XA1, refIdxLXB0, refIdxLXB1 and refIdxLXB2, the prediction list using flags predFlagLXA0, predFlagLXA1, predFlagLXB0, predFlagLXB1 and predFlagLXB2, motion vectors mvLXA0, mvLXA1, mvLXB0, mvLXB1 and mvLXB2 (where X is 0 or 1), and bidirectional prediction weight indices bcwIdxA0, bcwIdxA1, bcwIdxB0, bcwIdxB1, bcwIdxB2.
[0349] 2. The reference index refIdxLXCol (where X is 0 or 1) of the time-domain Merge candidate Col and the bidirectional prediction weight index bcwIdxCol are set to equal to 0.
[0350] 3. Invoke the derivation process for the temporal luminance motion vector prediction as specified in Clause 8.5.2.11, taking the luminance position (xCb, yCb), luminance codec block width cbWidth, luminance codec block height cbHeight, and variable refIdxL0Col as input, and outputting the availability flag availableFlagL0Col and the temporal motion vector mvL0Col. The variables availableFlagCol, predFlagL0Col, and predFlagL1Col are derived as follows:
[0351] availableFlagCol=availableFlagL0Col (8-263)
[0352] predFlagL0Col=availableFlagL0Col (8-264)
[0353] predFlagL1Col=0 (8-265)
[0354] 4. When slice_type equals B, invoke the derivation process for the temporal lumen motion vector prediction as specified in Clause 8.5.2.11, taking the lumen position (xCb, yCb), lumen codec block width cbWidth, lumen codec block height cbHeight, and variable refIdxL1Col as input, and outputting the availability flag availableFlagL1Col and the temporal motion vector mvL1Col. The variables availableFlagCol and predFlagL1Col are derived as follows:
[0355] availableFlagCol=availableFlagL0Col||availableFlagL1Col (8-266)
[0356] predFlagL1Col=availableFlagL1Col (8-267)
[0357] 5. The Merge Candidate List (mergeCandList) is constructed as follows:
[0358]
[0359] 6. The variables numCurrMergeCand and numOrigMergeCand are set to be equal to the number of Merge candidates in mergeCandList.
[0360] 7. When numCurrMergeCand is less than (MaxNumMergeCand-1) and NumHmvpCand is greater than 0, the following applies:
[0361] – Call the derivation process for history-based Merge candidates as specified in 8.5.2.6, with mergeCandList and numCurrMergeCand as inputs and modified mergeCandList and numCurrMergeCand as outputs.
[0362] –numOrigMergeCand is set to be equal to numCurrMergeCand.
[0363] 8. When numCurrMergeCand is less than MaxNumMergeCand but greater than 1, the following applies:
[0364] – Invoke the derivation procedure for pairwise average merge candidates as specified in Clause 8.5.2.4, where the inputs are mergeCandList, reference indices refIdxL0N and refIdxL1N, prediction list using flags predFlagL0N and predFlagL1N, motion vectors mvL0N and mvL1N for each candidate N in mergeCandList, and numCurrMergeCand, and the outputs are assigned to mergeCandList, numCurrMergeCand, reference indices refIdxL0avgCand and refIdxL1avgCand, prediction list using flags predFlagL0avgCand and predFlagL1avgCand, and motion vectors mvL0avgCand and mvL1avgCand for the candidate avgCands added to mergeCandList. The bidirectional prediction weight index bcwIdx for the candidate avgCands added to mergeCandList is set to 0.
[0365] –numOrigMergeCand is set to be equal to numCurrMergeCand.
[0366] 9. Invoke the derivation procedure for the zero motion vector Merge candidates as specified in Clause 8.5.2.5, wherein the inputs are mergeCandList, reference indices refIdxL0N and refIdxL1N, prediction list using flags predFlagL0N and predFlagL1N, motion vectors mvL0N and mvL1N for each candidate N in mergeCandList, and numCurrMergeCand, and the outputs are assigned to mergeCandList, numCurrMergeCand, and reference index refIdxL0zeroCand. m and refIdxL1zeroCand m The prediction list uses the flags predFlagL0zeroCand m and predFlagL1zeroCand m and each new candidate zeroCand added to mergeCandList m Motion vector mvL0zeroCand m and mvL1zeroCand m Each new zeroCand candidate added to mergeCandList mThe bidirectional prediction weight index bcwIdx is set to 0. The number of candidates added, numZeroMergeCand, is set to (numCurrMergeCand - numOrigMergeCand). When numZeroMergeCand is greater than 0, m ranges from 0 to numZeroMergeCand-1, inclusive.
[0367] 10. Perform the following assignment, where N is the candidate at position merge_idx[xCb][yCb] in the Merge candidate list mergeCandList (N = mergeCandList[merge_idx[xCb][yCb]]), and X is replaced with 0 or 1:
[0368] refIdxLX=refIdxLXN (8-269)
[0369] predFlagLX[0][0]=predFlagLXN (8-270)
[0370] mvLX[0][0][0]=mvLXN[0] (8-271)
[0371] mvLX[0][0][1]=mvLXN[1] (8-272)
[0372] bcwIdx=bcwIdxN (8-273)
[0373] 11. When mmvd_merge_flag[xCb][yCb] equals 1, the following applies:
[0374] – Call the derivation process of the Merge motion vector difference as specified in 8.5.2.7, with the brightness position (xCb, yCb), reference index refIdxL0, refIdxL1, and prediction list using flags predFlagL0[0][0] and predFlagL1[0][0] as input, and the motion vector difference mMvdL0 and mMvdL1 as output.
[0375] – The motion vector difference mMvdLX is added to the merged motion vector mvLX, for X being 0 and 1, as follows:
[0376] mvLX[0][0][0]+=mMvdLX[0] (8-274)
[0377] mvLX[0][0][1]+=mMvdLX[1] (8-275)
[0378] mvLX[0][0][0]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][0]) (8-276)
[0379] mvLX[0][0][1]=Clip3(-2 17 ,2 17 -1,mvLX[0][0][1]) (8-277)
[0380] 2.2.5 MMVD
[0381] In some embodiments, a final motion vector representation (UMVE, also known as MMVD) is presented. UMVE is used for skip or merge modes via a motion vector representation method.
[0382] In VVC, UMVE reuses the same Merge candidates included in the regular Merge candidate list. Among these Merge candidates, a base candidate can be selected and further extended using motion vector representation methods.
[0383] UMVE provides a new method for representing motion vector difference (MVD), in which MVD is represented by the starting point, motion amplitude, and motion direction.
[0384] In some embodiments, the Merge candidate list is used as is. However, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE extensions.
[0385] The basic candidate index defines the starting point. The basic candidate index indicates the best candidate among the candidates in the list, as shown below.
[0386] Table 4. Basic Candidate IDX
[0387] Basic candidate IDX 0 1 2 3 Nth MVP 1st MVP 2nd MVP 3rd MVP 4th MVP
[0388] If the number of basic candidates is equal to 1, then the basic candidate IDX will not be notified by signaling.
[0389] The distance index is motion amplitude information. The distance index indicates a predefined distance from the starting point. The predefined distances are shown below:
[0390] Table 5. Distance from IDX
[0391]
[0392] The direction index represents the direction of MVD relative to the starting point. The direction index can represent the four directions shown below.
[0393] Table 6. Directional IDX
[0394] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –
[0395] Immediately after transmitting the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, the UMVE flag is resolved. If the UMVE flag is equal to 1, the UMVE syntax is resolved. However, if it is not 1, the AFFINE flag is resolved. If the AFFINE flag is equal to 1, it is in AFFINE mode; otherwise, the skip / merge index is resolved to the VTM's skip / merge mode.
[0396] There is no need for an additional line buffer due to UMVE candidates. The software's skip / merge candidates are used directly as the base candidates. The MV supplement is determined immediately before motion compensation using the input UMVE index. There is no need to reserve a long line buffer for this.
[0397] Under the current general testing conditions, the first or second Merge candidate in the Merge candidate list can be selected as the basic candidate.
[0398] UMVE is also known as Merge with MV difference (MMVD).
[0399] 2.2.6 Combined Intra-Inter-Frame Prediction (CIIP)
[0400] In some embodiments, multiple hypothesis prediction is proposed, wherein combining intra-frame and inter-frame prediction is one way to generate multiple hypotheses.
[0401] When multi-hypothesis prediction is applied to improve intra-mode, it combines an intra-mode prediction with a Merge index prediction. In the Merge CU, a flag is communicated for Merge mode signaling when a flag is true, to select an intra-mode from the intra-candidate list. For the luma component, the intra-candidate list is derived from only one intra-prediction mode (e.g., planar mode). The weights applied to the prediction block from intra- and inter-frame predictions are determined by the encoding / decoding modes (intra- or non-intra-frame) of two neighboring blocks (A1 and B1).
[0402] 2.2.7 Merge based on sub-block technology
[0403] It is recommended that, in addition to the regular Merge list of non-sub-block Merge candidates, all sub-block-related motion candidates be placed in a separate Merge list.
[0404] Motion candidates related to sub-blocks are placed in a separate Merge list, which is named "Sub-block Merge Candidate List".
[0405] In one example, the list of sub-block merge candidates includes ATMVP candidates and affine merge candidates.
[0406] The candidate list for sub-block merge is populated in the following order:
[0407] 1. ATMVP candidate (may be available or not);
[0408] 2. Affine Merge List (including candidates for inheriting from affines and candidates for constructing affines)
[0409] 3. Zero-filled MV 4-parameter affine model
[0410] 2.2.7.1 ATMVP (also known as Sub-Block Temporal Motion Vector Prediction, sbTMVP)
[0411] The basic idea of ATMVP is to derive a set of multiple temporal motion vector predictions for a block. Each sub-block is assigned a set of motion information. When generating ATMVP Merge candidates, motion compensation is performed at the 8×8 level rather than at the entire block level.
[0412] In the current design, ATMVP predicts the motion vectors of sub-CUs within a CU in two steps, which will be described in the following two subsections.
[0413] 2.2.7.1.1 Derivation of Initialization of Motion Vectors
[0414] Let tempMv denote the initialization motion vector. When block A1 is available and is not encoded in an intra-frame manner (e.g., encoded in inter-frame or IBC mode), the following is used to derive the initialization motion vector.
[0415] - If all of the following conditions are true, then tempMv is set to be equal to the motion vector of block A1 from list 1, denoted as mvL1A1:
[0416] - The reference image index for list 1 is available (not equal to -1), and it has the same POC value as the juxtaposed image (e.g., DiffPicOrderCnt(ColPic,RefPicList[1][refIdxL1A1]) equals 0).
[0417] - No reference image has a larger Proof of Concept (POC) compared to the current image (e.g., for each image aPic in the reference image list for the current strip, DiffPicOrderCnt(aPic,currPic) is less than or equal to 0).
[0418] -The current stripe is equal to stripe B.
[0419] -collocated_from_l0_flag equals 0.
[0420] - Otherwise, if all of the following conditions are true, then tempMv is set to be equal to the motion vector of block A1 from list 0, denoted as mvL0A1:
[0421] - The reference image index for list 0 is available (not equal to -1).
[0422] - It has the same POC value as the juxtaposed image (e.g., DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A1]) equals 0).
[0423] - Otherwise, the zero motion vector is used to initialize the MV.
[0424] The corresponding block (with the center position of the current block plus the rounded MV, cropped if necessary to be within a specific range) is identified in the juxtaposed picture of the strip header signaling notification along with the initial motion vector.
[0425] If the block is inter-frame encoded or decoded, proceed to step two. Otherwise, the ATMVP candidate is set to unavailable.
[0426] 2.2.7.1.2 Derivation of Sub-CU Motion
[0427] The second step is to divide the current CU into sub-CUs and obtain the motion information of each sub-CU from the blocks corresponding to each sub-CU in the juxtaposed image.
[0428] If the corresponding block of a sub-CU is encoded and decoded in inter-frame mode, the final motion information of the current sub-CU is derived using motion information by invoking a derivation process for the juxtaposed motion vectors, which is no different from the traditional TMVP process. Basically, if the corresponding block is predicted from the target list X used for unidirectional or bidirectional prediction, the motion vector is used; otherwise, if it is predicted from the list Y (Y = 1-X) used for unidirectional or bidirectional prediction, and NoBackwardPredFlag is equal to 1, the MV of list Y is used. Otherwise, no motion candidate can be found.
[0429] If the blocks in the juxtaposed image identified by initializing the MV and the current sub-CU position are intra-frame encoded or IBC encoded, or if motion candidates cannot be found as described above, then the following further applies:
[0430] R will be used to extract the juxtaposed images. col The motion vector of the motion field in the equation is represented by MV. col To minimize the impact of MV scaling, the spatial candidate list is used to derive the MV. col The music video (MV) is selected as follows: if the reference image for a candidate MV is a juxtaposed image, then that MV is selected and used as the MV. col Otherwise, the MV with the closest juxtaposed image is selected to derive the MV using scaling. col .
[0431] The following is an example of the decoding process for the derivation of the juxtaposed motion vectors:
[0432] 8.5.2.12 Derivation of the juxtaposed motion vectors
[0433] The input to this process is:
[0434] – The variable currCb specifies the current encoding / decoding block.
[0435] – The variable colCb specifies the juxtaposition encoding / decoding block within the juxtaposition image specified by ColPic.
[0436] – Luminance position (xColCb, yColCb): Specifies the top left luminance sample of the juxtaposed luminance codec block specified by ColCb relative to the top left luminance sample of the juxtaposed image specified by ColPic.
[0437] – Refer to the index refIdxLX, where X is 0 or 1.
[0438] – Flag, indicating the sub-block time-domain Merge candidate sbFlag.
[0439] The output of this process is:
[0440] Motion vector prediction with a precision of -1 / 16 fractional sample points
[0441] –Availability flag: availableFlagLXCol.
[0442] The variable currPic specifies the current image.
[0443] The arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] are set to be equal to the PredFlagL0[x][y], MvDmvrL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively. The arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] are set to be equal to the PredFlagL1[x][y], MvDmvrL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.
[0444] The variables mvLXCol and availableFlagLXCol are derived as follows:
[0445] – If colCb is encoded or decoded in intra-frame or IBC prediction mode, both components of mvLXCol are set to 0, and availableFlagLXCol is set to equal to 0.
[0446] Otherwise, the motion vector mvCol, the reference index refIdxCol, and the reference list identifier listCol are derived as follows:
[0447] – If sbFlag equals 0, then availableFlagLXCol is set to 1, and the following applies:
[0448] – If predFlagL0Col[xColCb][yColCb] equals 0, then mvCol, refIdxCol, and listCol are set to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.
[0449] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then mvCol, refIdxCol, and listCol are set to mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.
[0450] Otherwise (predFlagL0Col[xColCb][yColCb] equals 1, and predFlagL1Col[xColCb][yColCb] equals 1), perform the following assignment:
[0451] – If NoBackwardPredFlag equals 1, then mvCol, refIdxCol, and listCol are set to mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively.
[0452] Otherwise, mvCol, refIdxCol, and listCol are set to equal mvLNCol[xColCb][yColCb], refIdxLNCol[xColCb][yColCb], and LN, respectively, where N is the value of collocated_from_l0_flag.
[0453] Otherwise (sbFlag equals 1), the following applies:
[0454] – If PredFlagLXCol[xColCb][yColCb] equals 1, then mvCol, refIdxCol, and listCol are set to mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively, and availableFlagLXCol is set to 1.
[0455] – Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following applies:
[0456] – If for each image aPic in the reference image list of the current stripe, DiffPicOrderCnt(aPic,currPic) is less than or equal to 0, and PredFlagLYCol[xColCb][yColCb] is equal to 1, then mvCol, refIdxCol, and listCol are set to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY, respectively, where Y equals ! X, where X is the value of X for which the procedure is called. availableFlagLXCol is set to 1.
[0457] Both components of –mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0458] – When availableFlagLXCol equals TRUE, mvLXCol and availableFlagLXCol are derived as follows:
[0459] – If LongTermRefPic(currPic,currCb,refIdxLX,LX) is not equal to LongTermRefPic(ColPic,colCb,refIdxCol,listCol), then both components of mvLXCol are set to 0, and availableFlagLXCol is set to 0.
[0460] Otherwise, the variable availableFlagLXCol is set to 1, refPicList[listCol][refIdxCol] is set to the list of reference images in listCol containing the stripes of the codec block colCb in the juxtaposed images specified by ColPic, and the following applies:
[0461] colPocDiff=DiffPicOrderCnt(ColPic,refPicList[listCol][refIdxCol])(8-402)
[0462] currPocDiff=DiffPicOrderCnt(currPic,RefPicList[X][refIdxLX]) (8-403)
[0463] – Invoke the temporal motion buffer compression process of the juxtaposed motion vector as specified in Clause 8.5.2.15, with mvCol as input and the modified mvCol as output.
[0464] – If RefPicList[X][refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then mvLXCol is derived as follows:
[0465] mvLXCol=mvCol (8-404)
[0466] Otherwise, mvLXCol is derived as a scaled version of the motion vector mvCol, as follows:
[0467] tx=(16384+(Abs(td)>>1)) / td (8-405)
[0468] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (8-406)
[0469] mvLXCol=Clip3(-131072,131071,(distScaleFactor*mvCol+128-(distScaleFactor*mvCol>=0))>>8)) (8-407)
[0470] The derivation of td and tb is as follows:
[0471] td=Clip3(-128,127,colPocDiff) (8-408)
[0472] tb=Clip3(-128,127,currPocDiff) (8-409)
[0473] 2.2.8 Refinement of motion information
[0474] 2.2.8.1 Decoder-Side Motion Vector Refinement (DMVR)
[0475] In bidirectional prediction, for the prediction of a block region, two prediction blocks formed using motion vectors (MV) from list 0 and MV from list 1 are combined to form a single prediction signal. In the decoder-side motion vector refinement (DMVR) method, the two motion vectors of the bidirectional prediction are further refined.
[0476] For DMVR in VVC, assuming an MVD mirrored between list 0 and list 1, such as... Figure 19 As shown, bilateral matching is performed to refine the motion vectors (MVs), for example, to find the optimal MVD among several MVD candidates. Let MVL0(L0X, L0Y) and MVL1(L1X, L1Y) represent the MVs of two lists of reference images. The MVD of list 0, represented by (MvdX, MvdY), that minimizes the cost function (e.g., SAD) is defined as the optimal MVD. For the SAD function, it is defined as the SAD between the reference block of list 0 and the reference block of list 1, where the reference block of list 0 is derived using the motion vectors (L0X+MvdX, L0Y+MvdY) in the list 0 reference image, and the reference block of list 1 is derived using the motion vectors (L1X-MvdX, L1Y-MvdY) in the list 1 reference image.
[0477] The motion vector thinning process can be iterated twice. In each iteration, up to six MVDs (with integer pixel precision) can be checked in two steps, such as... Figure 20As shown. In the first step, MVD(0,0), (-1,0), (1,0), (0,-1), and (0,1) are checked. In the second step, one of MVD(-1,-1), (-1,1), (1,-1), or (1,1) can be selected and further checked. Assume the function Sad(x,y) returns the SAD value of MVD(x,y). The MVD checked in the second step, represented by (MvdX, MvdY), is determined as follows:
[0478] MvdX = -1;
[0479] MvdY = -1;
[0480] If (Sad(1,0)<Sad(-1,0))
[0481] MvdX = 1;
[0482] If (Sad(0,1)<Sad(0,-1))
[0483] MvdY = 1;
[0484] In the first iteration, the starting point is the MV of the signaling notification, and in the second iteration, the starting point is the MV of the signaling notification plus the best MVD selected in the first iteration. DMVR is only applicable when one reference picture is the preceding picture and the other reference picture is the following picture, and both reference pictures have the same picture order count distance to the current picture.
[0485] To further simplify the DMVR process, in some embodiments, the DMVR design employed has the following key features:
[0486] - Terminate early when the SAD at position (0,0) between list 0 and list 1 is less than the threshold.
[0487] - Terminate early when the SAD between list 0 and list 1 is zero at some position.
[0488] - DMVR block size: W*H>=64&&H>=8, where W and H are the width and height of the block.
[0489] - For DMVRs with CU dimensions > 16x16, divide the CU into multiple 16x16 sub-blocks. If only the width or height of the CU is greater than 16, divide it only in the vertical or horizontal direction.
[0490] - Reference block size (W+7)*(H+7) (for brightness).
[0491] -25-point SAD-based integer pixel search (e.g., (+-)2 refinement of the search range, single stage)
[0492] - DMVR based on bilinear interpolation.
[0493] - Subpixel thinning based on the "parameter error surface equation". This process is only performed if the minimum SAD cost is not equal to zero and the optimal MVD is (0, 0) in the last MV thinning iteration.
[0494] -Luminosity / Chromaticity MC w / Reference Block Fill (if needed).
[0495] - Refined MV for MC and TMVP only.
[0496] 2.2.8.1.1 Use of DMVR
[0497] DMVR can be enabled when all of the following conditions are true:
[0498] - The DMVR enable flag in SPS (e.g., sps_dmvr_enabled_flag) is equal to 1.
[0499] - The TPM flag, inter-frame affine flag, sub-block merge flag (ATMVP or affine merge), and MMVD flag are all equal to 0.
[0500] -Merge flag equals 1
[0501] - The current block is predicted bidirectionally, and the POC distance between the current image and the reference images in list 1 is equal to the POC distance between the reference images in list 0 and the current image.
[0502] - Current CU height is greater than or equal to 8
[0503] - The number of luminance samples (CU width * height) is greater than or equal to 64
[0504] 2.2.8.1.2 Sub-pixel thinning based on "parameter error surface equation"
[0505] The method is summarized as follows:
[0506] 1. Parameter error surface fitting is calculated only if the center position is the optimal cost position in a given iteration.
[0507] 2. The center location cost and the costs at positions (-1, 0), (0, -1), (1, 0), and (0, 1) from the center are used to fit the 2-D parabolic error surface equation of the shape.
[0508] E(x,y)=A(x-x0) 2 +B(y-y0) 2 +C
[0509] Where (x0, y0) corresponds to the position of minimum cost, and C corresponds to the minimum cost value. By solving the five equations with five unknowns, (x0, y0) is calculated as:
[0510] x0=(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0)))
[0511] y0=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0)))
[0512] (x0, y0) can be calculated to any desired sub-pixel precision by adjusting the precision of the division (e.g., how many bits of quotient to calculate). For 1 / 16 th Pixel precision requires only 4 bits of the absolute value of the quotient, making it suitable for implementations based on fast shift subtraction with only 2 divisions per CU.
[0513] 3. Add the calculated (x0, y0) to the integer distance thinning MV to obtain the precise thinning increment MV for each subpixel.
[0514] 2.3 Intra-frame block copying
[0515] Intra-Block Copy (IBC), also known as Current Picture Reference, has been adopted in HEVC Screen Content Codec Extension (HEVC-SCC) and the current VVC test model (VTM-4.0). IBC extends the concept of motion compensation from inter-frame codec to intra-frame codec. For example... Figure 21 As shown, when IBC is applied, the current block is predicted using a reference block in the same image. Samples in the reference block must be reconstructed before the current block is encoded or decoded. While IBC is not very effective for most camera-captured sequences, it shows significant encoding / decoding gains for screen content. This is because there are many repeating patterns, such as icons and text characters in screen content images. IBC can effectively remove redundancy between these repeating patterns. In HEVC-SCC, IBC can be applied if the inter-frame codec unit (CU) selects the current image as its reference image. In this case, the MV is renamed to a block vector (BV), and the BV always has integer pixel precision. For compatibility with the master profile HEVC, the current image is marked as the "long-term" reference image in the decoded image buffer (DPB). It should be noted that, similarly, in multi-view / 3D video codec standards, inter-view reference images are also marked as "long-term" reference images.
[0516] After BV finds its reference block, predictions can be generated by copying the reference block. The residual can be obtained by subtracting the reference pixel from the original signal. Then, transforms and quantization can be applied as in other encoding / decoding modes.
[0517] However, some or all pixel values may be undefined when the reference block is outside the image, overlaps with the current block, is outside the reconstructed region, or is outside a valid region subject to certain constraints. Essentially, there are two solutions to handle such problems. One is to disallow such cases, such as in bitstream consistency. The other is to apply padding to those undefined pixel values. The following subsections describe the solutions in detail.
[0518] 2.3.1 IBC in the VVC Test Model (VTM4.0)
[0519] In the current VVC test model (e.g., VTM-4.0 design), the entire reference block should be together with the current codec tree unit (CTU) and should not overlap with the current block. Therefore, there is no need to pad the reference or prediction block. The IBC flag is encoded as the prediction mode of the current CU. Therefore, for each CU, there are a total of three prediction modes: MODE_INTRA, MODE_INTER, and MODE_IBC.
[0520] 2.3.1.1 IBC Merge Mode
[0521] In IBC Merge mode, indices pointing to entries in the IBC Merge candidate list are parsed from the bitstream. The construction of the IBC Merge list can be summarized in the following steps:
[0522] Step 1: Derive the candidate airspace
[0523] Step 2: Insert HMVP candidate
[0524] Step 3: Insert pairwise average candidates
[0525] In the derivation of the spatial Merge candidate, when located in such a position... Figure 2 Up to four merged candidates are selected from the candidates depicted at positions A1, B1, B0, A0, and B2. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or is not encoded / decoded in IBC mode. After the candidate at position A1 is added, the insertion of the remaining candidates undergoes a redundancy check, which ensures that candidates with the same motion information are excluded from the list to improve encoding / decoding efficiency.
[0526] After inserting an empty domain candidate, if the IBC Merge list size is still smaller than the maximum IBC Merge list size, an IBC candidate from the HMVP table can be inserted. Redundancy checks are performed when inserting an HMVP candidate.
[0527] Finally, the pairwise average candidates are inserted into the IBC Merge list.
[0528] A Merge candidate is considered invalid when the reference block identified by the Merge candidate is outside the image, overlaps with the current block, is outside the reconstructed region, or is outside the valid region subject to some constraints.
[0529] It should be noted that invalid merge candidates can be inserted into the IBC merge list.
[0530] 2.3.1.2 IBC AMVP Mode
[0531] In IBC AMVP mode, the AMVP index pointing to an entry in the IBC AMVP list is parsed from the bitstream. The construction of the IBC AMVP list can be summarized in the following steps:
[0532] Step 1: Derive the candidate airspace
[0533] - Check A0 and A1 until a usable candidate is found.
[0534] - Check B0, B1, B2 until a usable candidate is found.
[0535] Step 2: Insert HMVP candidate
[0536] Step 3: Insert zero candidate
[0537] After inserting the spatial candidate, if the size of the IBC AMVP list is still smaller than the maximum IBC AMVP list size, then the IBC candidate in the HMVP table can be inserted.
[0538] Finally, the zero candidate was inserted into the IBC AMVP list.
[0539] 2.3.1.3 Chromaticity IBC Mode
[0540] In the current VVC, motion compensation in chroma IBC mode is performed at the sub-block level. A chroma block is divided into several sub-blocks. Each sub-block determines whether its corresponding luma block has a block vector, and if so, its validity is determined. The current VTM has encoder constraints; if all sub-blocks in the current chroma CU have valid luma block vectors, the chroma IBC mode is tested. For example, in YUV420 video, the chroma block is N×M, and the co-occurring luma region is 2N×2M. The sub-block size of the chroma block is 2×2. Performing the chroma mv derivation and block copying process requires several steps.
[0541] (1) The chroma block will first be divided into (N>>1)*(M>>1) sub-blocks.
[0542] (2) For each sub-block with upper left sampling coordinates (x, y), obtain the corresponding brightness block with the same upper left sampling coordinates (2x, 2y).
[0543] (3) The encoder verifies the block vector (bv) of the acquired luminance block. If one of the following conditions is met, the bv will be considered invalid.
[0544] a. The bv of the corresponding luminance block does not exist.
[0545] b. The prediction block identified by bv has not yet been reconstructed.
[0546] c. The predicted block identified by bv partially or completely overlaps with the current block.
[0547] (4) The chromaticity motion vector of the sub-block is set to the motion vector of the corresponding luminance sub-block.
[0548] IBC mode is enabled at the encoder when all sub-blocks find a valid bv.
[0549] 2.3.2 Recent Developments in IBC
[0550] 2.3.2.1 Single BV List
[0551] In some embodiments, the BV predictions of the Merge mode and the AMVP mode in IBC share a common prediction list, which consists of the following elements:
[0552] (1) Two adjacent locations in the airspace (e.g.) Figure 2 (A1, B1)
[0553] (2) 5 HMVP entries
[0554] (3) The default value is zero vector
[0555] The number of candidates in the list is controlled by a variable derived from the stripe header. For Merge mode, a maximum of the first 6 entries of the list can be used; for AMVP mode, the first 2 entries can be used. Furthermore, the list must conform to the shared Merge list region requirement (the same list is shared within SMR).
[0556] In addition to the BV prediction candidate list mentioned above, the pruning operation between the HMVP candidate and the existing Merge candidate (A1, B1) can be simplified. In this simplification, there will be a maximum of two pruning operations, as it only compares the first HMVP candidate with (multiple) spatial Merge candidates.
[0557] 2.3.2.2 Size Limitations of IBC
[0558] In some embodiments, the syntax constraints for disabling the 128x128 IBC mode can be explicitly applied over the current bitstream constraints in previous VTM and VVC versions, which makes the presence of the IBC flag dependent on the CU size < 128x128.
[0559] 2.3.2.3 Shared Merge List of IBC
[0560] To reduce decoder complexity and support parallel encoding, in some embodiments, the same Merge candidate list can be shared across all leaf codec units (CUs) of an ancestor node in the CU partitioning tree to enable parallel processing of small skip / Merge codec units. The ancestor node is named the Merge Shared Node. The shared Merge candidate list is generated at the Merge Shared Node, assuming the Merge Shared Node is a leaf CU.
[0561] More specifically, the following may apply:
[0562] - If a block has no more than 32 luminance samples and is divided into two 4x4 sub-blocks, then a shared Merge list between very small blocks (e.g., two adjacent 4x4 blocks) is used.
[0563] - If a block has more than 32 luminance samples, but after partitioning, at least one sub-block is less than the threshold (32), then all sub-blocks of that partition share the same Merge list (e.g., 16x4 or 4x16 partitioned ternary or 8x8 partitioned using quaternary).
[0564] Such restrictions apply only to the IBC Merge mode.
[0565] 3. Problems solved by the embodiments
[0566] A block can be encoded and decoded in IBC mode. However, different sub-regions within a block can have different contents. Further investigation is needed to explore the correlation with previously encoded and decoded blocks within the current frame.
[0567] 4. Examples of Implementation Methods
[0568] In this document, Intra-Block Copy (IBC) may not be limited to current IBC techniques, but can be interpreted as a technique that excludes the use of reference samples within the current strip / piece / brick / picture / other video unit (e.g., CTU line) by traditional intra-frame prediction methods.
[0569] To address the aforementioned issues, a sub-block-based IBC (sbIBC) encoding / decoding method is proposed. In sbIBC, the current IBC-encoded video block (e.g., CU / PU / CB / PB) is divided into multiple sub-blocks. Each sub-block can have a size smaller than the video block size. For each corresponding sub-block among the multiple sub-blocks, the video codec can identify a reference block for the corresponding sub-block within the current picture / strip / piece / tile / piece group. The video codec can then use the motion parameters of the identified reference block for the corresponding sub-block to determine the motion parameters of that sub-block.
[0570] Furthermore, IBC is not limited to application only to unidirectional predictive codec blocks. Bidirectional prediction, where both reference images are the current image, can also be supported. Alternatively, bidirectional prediction, where one reference image is from the current image and the other is from a different image, can also be supported. In yet another example, multiple hypotheses can also be applied.
[0571] The list below should be considered as examples to illustrate general concepts. These techniques should not be interpreted narrowly. Furthermore, these techniques can be combined in any way. Figure 2 The neighboring blocks A0, A1, B0, B1, and B2 are shown in the figure.
[0572] 1. In sbIBC, a block of size M×N can be divided into more than one sub-block.
[0573] a. In one example, the sub-block size is fixed at L×K, for example, L=K=4.
[0574] b. In one example, the sub-block size is fixed as the smallest decoding unit / prediction unit / transform unit / unit for storing motion information.
[0575] c. In one example, a block can be divided into multiple sub-blocks of different or equal sizes.
[0576] d. In one example, signaling can be used to indicate the size of the sub-block.
[0577] e. In one example, the indication of the sub-block size can be changed block by block, for example, based on the block dimension.
[0578] f. In one example, the sub-block size must be of the form (N1×minW)×(N2×minH), where minW×minH represents the smallest decoding unit / prediction unit / transform unit / unit used for storing motion information, and N1 and N2 are positive integers.
[0579] g. In one example, the sub-block dimension may depend on the color format and / or color components.
[0580] i. For example, the size of sub-blocks of different color components can be different.
[0581] 1) Alternatively, the sub-blocks of different color components can be the same size.
[0582] ii. For example, when the color format is 4:2:0, a 2L×2K sub-block of the luminance component can correspond to an L×K sub-block of the chrominance component.
[0583] 1) Alternatively, when the color format is 4:2:0, the four 2L×2K sub-blocks of the luminance component can correspond to the 2L×2K sub-blocks of the chrominance component.
[0584] iii. For example, when the color format is 4:2:2, a 2L×2K sub-block of the luminance component can correspond to a 2L×K sub-block of the chrominance component.
[0585] 1) Alternatively, when the color format is 4:2:2, the two 2L×2K sub-blocks of the luminance component can correspond to the 2L×2K sub-blocks of the chrominance component.
[0586] iv. For example, when the color format is 4:4:4, a 2L×2K sub-block of the luminance component can correspond to a 2L×2K sub-block of the chrominance component.
[0587] h. In one example, the MV of a sub-block of the first color component can be derived from one or more corresponding sub-blocks of the second color component.
[0588] i. For example, the MV of a sub-block of the first color component can be derived as the average MV of multiple corresponding sub-blocks of the second color component.
[0589] ii. Alternatively, the above method can be applied when using a single tree.
[0590] iii. Alternatively, the above method can be applied when dealing with a specific block size (such as a 4×4 chroma block).
[0591] i. In one example, the sub-block size can depend on the encoding / decoding mode, such as IBC Merge / AMVP mode.
[0592] j. In one example, the sub-blocks can be non-rectangular, such as triangles / wedges.
[0593] 2. The motion information of the sub-CU is obtained by using two stages (including identifying the corresponding reference block with the initial motion vector (denoted as initMV) and deriving one or more motion vectors of the sub-CU based on the reference block), wherein at least one reference image is equal to the current image.
[0594] a. In one example, the reference block can be within the current image.
[0595] b. In one example, the reference block can be in the reference image.
[0596] i. For example, it can be in a juxtaposed reference image.
[0597] ii. For example, it can be found in a reference image that is identified by using motion information of juxtaposed blocks or neighboring blocks of juxtaposed blocks.
[0598] Phase 1.a Settings for initMV(vx,vy)
[0599] c. In one example, initMV can be derived from one or more neighboring (adjacent or non-adjacent) blocks of the current block or the current child block.
[0600] i. Neighboring blocks can be blocks within the same image.
[0601] 1) Alternatively, it can be a block in a reference image.
[0602] a. For example, it can be in a juxtaposed reference image.
[0603] b. For example, it can be identified by using motion information of juxtaposed blocks or neighboring blocks of juxtaposed blocks.
[0604] ii. In one example, it can be derived from neighboring block Z.
[0605] 1) For example, initMV can be set to equal the value stored in neighboring block Z.
[0606] In the MV. For example, neighboring block Z can be block A1.
[0607] iii. In one example, it can be derived from multiple blocks checked in sequence.
[0608] 1) In one example, the first identified motion vector associated with the current image, which is a reference image from the inspection block, can be set to initMV.
[0609] d. In one example, initMV can be derived from the list of motion candidates.
[0610] i. In one example, it can be derived from the kth (e.g., the 1st) candidate in the IBC candidate list.
[0611] 1) In one example, the IBC candidate list is the Merge / AMVP candidate list.
[0612] 2) In one example, an IBC candidate list that differs from the existing IBC Merge candidate list construction process can be used, such as using different spatial neighbor blocks.
[0613] ii. In one example, it can be derived from the k-th (e.g., the 1st) candidate in the IBC HMVP table.
[0614] e. In one example, it can be deduced based on the position of the current block.
[0615] f. In one example, it can be derived based on the dimensions of the current block.
[0616] g. In one example, it can be set as the default value.
[0617] h. In one example, the initMV instruction can be communicated via signaling at the video unit level (such as slice / strip / picture / tile / CTU line / CTU / CTB / CU / PU / TU etc.).
[0618] i. The initial MV can be different for two different sub-blocks within the current block.
[0619] j. How to deduce that the initial MV can be changed block by block, slice by slice, strip by strip, etc.
[0620] Phase 1.b: Reference block for identifying sub-CUs using initMV
[0621] k. In one example, the initMV can first be converted to 1-pixel integer precision, and the converted MV can be used to identify the corresponding block of the sub-block. The converted MV is represented by (vx', vy').
[0622] i. In one example, if (vx,vy) is in F pixel integer precision, the converted MV represented by (vx',vy') can be set to (vx*F,vy*F) (e.g., F = 2 or 4).
[0623] ii. Alternatively, (vx',vy') is directly set to equal (vx,vy).
[0624] l. Assume a sub-block has a top left position of (x, y) and a sub-block size of K×L. The corresponding block of the sub-block is set to CU / CB / PU / PB covering the coordinates (x+offsetX+vx', y+offsetY+vy'), where offsetX and offsetY are used to indicate the selected coordinates relative to the current sub-block.
[0625] i. In one example, offsetX and / or offsetY are set to 0.
[0626] ii. In one example, offsetX can be set to (L / 2) or (L / 2+1) or (L / 2–1), where L can be the width of the sub-block.
[0627] iii. In one example, offsetY can be set to (K / 2), (K / 2+1), or (K / 2–1), where K can be the height of the sub-block.
[0628] iv. Alternatively, the horizontal and / or vertical offsets may be further cropped to a range, such as within the boundaries of the image / strip / piece / tile / IBC reference area.
[0629] Phase 2 involves deriving the motion vectors of sub-blocks using the motion information of the identified corresponding reference block (via subMV). (subMVx, subMVy) represents)
[0630] The subMV of a subblock is derived from the motion information of the corresponding block.
[0631] i. In one example, if the corresponding block has a motion vector pointing to the current image, then subMV is set to equal MV.
[0632] ii. In one example, if the corresponding block has a motion vector pointing to the current image, then subMV is set to equal MV plus initMV.
[0633] The derived subMV can be further trimmed to a given range, or trimmed to ensure that it points to the IBC reference region.
[0634] o. In a consistent bitstream, the derived subMV must be a valid MV of the IBC of the subblock.
[0635] 3. It can generate one or more IBC candidates with sub-block motion vectors, which can be represented as sub-block IBC candidates.
[0636] 4. Sub-block IBC candidates can be inserted into sub-block Merge candidates, including ATMVP and affine Merge candidates.
[0637] a. In one example, it can be added before all other sub-block Merge candidates.
[0638] b. In one example, it can be added after the ATMVP candidate.
[0639] c. In one example, it can be added after an inherited affine candidate or a constructed affine candidate.
[0640] d. In one example, it can be added to the IBC Merge / AMVP candidate list.
[0641] i. Alternatively, whether to add it may depend on the current block's schema information.
[0642] For example, if it is an IBC AMVP mode, it may not need to add it.
[0643] e. Which candidate list to add can depend on the partitioning structure, such as a bitree or a single tree.
[0644] f. Alternatively, multiple sub-block IBC candidates can be inserted into the sub-block Merge candidate.
[0645] 5. The list of IBC sub-block motions (e.g., AMVP / Merge) candidates can be constructed using at least one sub-block IBC candidate.
[0646] a. Alternatively, one or more sub-block IBC candidates can be inserted into the IBC sub-block Merge candidate, for example, using different initialization MVs.
[0647] b. Alternatively, the construction of the IBC sub-block motion candidate list or the existing IBC AMVP / Merge candidate list can be signaled via indicators or dynamically (on-the-fly) deduced.
[0648] c. Alternatively, if the current block is encoded or decoded in IBC Merge mode, the index of the IBC sub-block merge candidate list can be signaled.
[0649] d. Alternatively, if the current block is encoded or decoded in IBC AMVP mode, the index of the IBC subblock AMVP candidate list can be signaled.
[0650] i. Alternatively, the signaling notification / derivative MVD of the IBC AMVP mode can be applied to one or more sub-blocks.
[0651] 6. The reference block of a sub-block and the sub-block can belong to the same color component.
[0652] An extension of sbIBC by mixing and using other tools applied to different sub-blocks within the same block.
[0653] 7. A block may be divided into multiple sub-blocks, at least one of which is encoded and decoded using IBC, and at least one of which is encoded and decoded in intra-frame mode.
[0654] a. In one example, motion vectors may not be derived for a sub-block. Instead, one or more intra-prediction modes may be derived for the sub-block.
[0655] b. Alternatively, palette patterns and / or palette tables can be derived.
[0656] c. In one example, an intra-prediction mode can be derived for the entire block.
[0657] 8. A block can be divided into multiple sub-blocks, all of which are encoded and decoded in intra-frame mode.
[0658] 9. A block can be divided into multiple sub-blocks, all of which are encoded and decoded in a palette mode.
[0659] 10. A block can be divided into multiple sub-blocks, wherein at least one sub-block is encoded and decoded in IBC mode, and at least one sub-block is encoded and decoded in palette mode.
[0660] 11. A block may be divided into multiple sub-blocks, wherein at least one sub-block is encoded and decoded in intra-frame mode and at least one sub-block is encoded and decoded in palette mode.
[0661] 12. A block can be divided into multiple sub-blocks, wherein at least one sub-block is encoded and decoded in IBC mode, and at least one sub-block is encoded and decoded in inter-frame mode.
[0662] 13. A block may be divided into multiple sub-blocks, wherein at least one sub-block is encoded and decoded in intra-frame mode and at least one sub-block is encoded and decoded in inter-frame mode.
[0663] Interaction with other tools
[0664] 14. When applying one or more of the above methods, the IBC HMVP table does not need to be updated.
[0665] a. Alternatively, one or more motion vectors from the IBC codec subregions can be used to update the IBCHMVP table.
[0666] 15. When applying one or more of the above methods, it is not necessary to update the non-IBC HMVP table.
[0667] b. Alternatively, one or more motion vectors from inter-frame coding / decoding sub-regions can be used to update non-IBCHMVP tables.
[0668] 16. The loop filtering process (e.g., deblocking process) may depend on the use of the methods described above.
[0669] a. In one example, when one or more of the methods described above are applied, the sub-block boundaries can be filtered.
[0670] a. Alternatively, when one or more of the above methods are applied, the sub-block boundaries can be filtered.
[0671] b. In one example, blocks encoded or decoded using the above method can be processed in a similar manner to traditional IBC encoded or decoded blocks.
[0672] 17. Specific encoding / decoding methods (e.g., sub-block transform, affine motion prediction, multi-reference line intra-prediction, matrix-based intra-prediction, symmetric MVD encoding / decoding, Merge with side motion derivation / refinement using MVD decoder, bidirectional optical flow, simplified quadratic transform, multiple transform sets, etc.) can be disabled for blocks encoded / decoded using one or more of the methods described above.
[0673] 18. Instructions for the use of the above methods and / or sub-block sizes may be signaled or dynamically derived at the sequence / picture / strip / piece group / piece / picture brick / CTU / CTB / CU / PU / TU / other video unit level.
[0674] a. In one example, one or more of the methods described above can be considered as a special IBC mode.
[0675] i. Alternatively, if a block is encoded or decoded in IBC mode, further instructions can be signaled or derived using the conventional whole-block-based IBC approach or sbIBC.
[0676] ii. In one example, subsequent IBC codec blocks can use the motion information of the current sbIBC codec block as MV prediction values.
[0677] 1. Alternatively, subsequent IBC codec blocks may not be allowed to use the motion information of the current sbIBC codec block as the MV prediction value.
[0678] b. In one example, sbIBC can be indicated by the candidate index of the motion candidate list.
[0679] i. In one example, a specific candidate index is assigned to the sbIBC codec block.
[0680] c. In one example, IBC candidates can be categorized into two classes: one for whole-block encoding / decoding and the other for sub-block encoding / decoding. Whether a block is encoded / decoded in sbIBC mode can depend on the category of the IBC candidate.
[0681] Use of tools
[0682] 19. Whether and / or how to apply the above methods may depend on the following information:
[0683] a. In DPS / SPS / VPS / PPS / APS / Image Header / Strip Header / Group Header / Maximum Codec Unit (LCU) / Codec Unit (CU) / LCU Line / LCU Group / TU / PU Block / Video
[0684] Signaling notification messages in the encoding / decoding unit
[0685] b. Location of CU / PU / TU / block / video codec unit
[0686] c. Block dimensions of the current block and / or its neighboring blocks
[0687] d. Block shape of the current block and / or its neighboring blocks
[0688] e. Intra-frame mode of the current block and / or its neighboring blocks
[0689] f. The motion / block vector of its neighboring blocks
[0690] g. Indication of color format (such as 4:2:0, 4:4:4)
[0691] h. Encoding / decoding tree structure
[0692] i. Strip / group type and / or image type
[0693] j. Color components (e.g., can be applied only to the chromaticity component or the luminance component)
[0694] k. Time-domain layer ID
[0695] l. Standard summary / level / tier
[0696] Ideas related to the Merge list construction process and IBC usage
[0697] 20. For blocks in an inter-frame encoded / decoded picture / strip / slice / slice, the IBC mode can be used in conjunction with the inter-frame prediction mode.
[0698] a. In one example, for IBC AMVP mode, signaling can be used to inform syntax elements whether the current block is predicted from the current picture and a reference picture that is different from the current picture (represented as a temporal reference picture).
[0699] i. Alternatively, if the current block is also predicted from a time-domain reference picture, the syntax element can be signaled to indicate which time-domain reference picture and its associated MVP index, MVD, MV precision, etc., to use.
[0700] ii. In one example, for IBC AMVP mode, one list of reference images may include only the current image, and another list of reference images may include only the time-domain reference image.
[0701] b. In one example, for the IBC Merge mode, the motion vector and reference image can be derived from neighboring blocks.
[0702] i. For example, if neighboring blocks are predicted only from the current image, then the motion information derived from neighboring blocks can be referenced only to the current image.
[0703] ii. For example, if neighboring blocks are predicted from the current image and a temporal reference image, the derived motion information can be referenced to the current image and the temporal reference image.
[0704] 1) Alternatively, the derived motion information can be derived by referring only to the current image.
[0705] iii. For example, if neighboring blocks are predicted only from temporal reference images, they can be considered “invalid” or “unavailable” when constructing IBC Merge candidates.
[0706] c. In one example, a fixed weighting factor can be assigned to a reference block from the current image and a reference block from a temporal reference image for bidirectional prediction.
[0707] i. Alternatively, the weighting factor can be signaled.
[0708] 21. The motion candidate list construction process (e.g., regular merge list, IBC merge / AMVP list, sub-block merge list, IBC sub-block candidate list) and / or whether / how to update the HMVP table may depend on block dimensions and / or merge sharing conditions. Let the width and height of the block be represented as W and H, respectively. Condition C may depend on block dimensions and / or encoding / decoding information.
[0709] a. The process of constructing the motion candidate list (e.g., regular merge list, IBC merge / AMVP list, sub-block merge list, IBC sub-block candidate list) and / or whether / how to update the HMVP table may depend on condition C.
[0710] b. In one example, condition C may depend on the encoding and decoding information of the current block and / or its neighboring (adjacent or non-adjacent) blocks.
[0711] c. In one example, condition C may depend on Merge shared conditions.
[0712] d. In one example, condition C may depend on the block dimension of the current block, and / or the block dimensions of neighboring (adjacent or non-adjacent) blocks and / or the encoding / decoding mode of the current and / or neighboring blocks.
[0713] e. In one example, if condition C is satisfied, the derivation of the spatial Merge candidate is skipped.
[0714] f. In one example, if condition C is met, candidates derived from spatially adjacent (adjacent or non-adjacent) blocks are skipped.
[0715] g. In one example, if condition C is met, candidates derived from a specific spatial neighbor (adjacent or non-adjacent) block (e.g., block B2) are skipped.
[0716] h. In one example, if condition C is met, the derivation of the HMVP candidate is skipped.
[0717] i. In one example, if condition C is satisfied, the derivation of the pairwise Merge candidates is skipped.
[0718] j. In one example, if condition C is met, the maximum number of pruning operations is reduced or set to 0.
[0719] i. Alternatively, pruning operations can be reduced or removed from the airspace merge candidates.
[0720] ii. Alternatively, pruning operations can be reduced or removed from HMVP candidates and other Merge candidates.
[0721] k. In one example, if condition C is met, the update of the HMVP candidate is skipped.
[0722] i. In one example, HMVP candidates can be added directly to the sports list without being pruned.
[0723] l. In one example, if condition C is met, no default motion candidate is added (e.g., zero motion candidates in the IBC Merge / AVMP list).
[0724] m. In one example, when condition C is met, different checking orders (e.g., from first to last instead of from last to first) and / or different numbers of HMVP candidates to be checked / added.
[0725] n. In one example, condition C can be satisfied when W*H is greater than or not less than a threshold (e.g., 1024).
[0726] o. In one example, condition C can be satisfied when W and / or H are greater than or not less than a threshold (e.g., 32).
[0727] p. In one example, condition C can be satisfied when W is greater than or not less than a threshold (e.g., 32).
[0728] q. In one example, condition C can be satisfied when H is greater than or not less than a threshold (e.g., 32).
[0729] r. In one example, condition C can be satisfied when W*H is greater than or not less than a threshold (e.g., 1024) and the current block is encoded and decoded in IBC AMVP and / or Merge mode.
[0730] s. In one example, condition C can be satisfied when W*H is less than or no greater than a threshold (e.g., 16, 32, or 64) and the current block is encoded / decoded in IBCAMVP and / or Merge mode.
[0731] i. Alternatively, when condition C is met, the IBC motion list construction process may include candidates from spatially neighboring blocks (e.g., A1, B1) as well as default candidates. That is, the insertion of HMVP candidates is skipped.
[0732] ii. Alternatively, when condition C is met, the IBC motion list construction process may include candidates from the HMVP candidates in the IBCHMVP table, as well as default candidates. That is, the insertion of candidates from spatially neighboring blocks is skipped.
[0733] iii. Alternatively, after decoding a block if condition C is met, the update of the IBC HMVP table is skipped.
[0734] iv. Alternatively, condition C can be satisfied when one / some / all of the following conditions are true:
[0735] 1) When W*H is equal to or not greater than T1 (e.g., 16) and the current block is encoded / decoded in IBC AMVP and / or Merge mode.
[0736] 2) When W equals T2 and H equals T3 (e.g., T2 = 4, T3 = 8), the block above it is available and its size is A x B; and both the current block and the block above it are encoded and decoded in a specific mode.
[0737] a. Alternatively, when W equals T2 and H equals T3 (e.g., T2 = 4, T3 = 8), the block above it is available in the same CTU, and its size is equal to AxB, and the current block and its block above it are encoded and decoded in the same mode.
[0738] b. Alternatively, when W equals T2 and H equals T3 (e.g., T2 = 4, T3 = 8), the block above it is available, and its size is equal to A x B, and the current block and the block above it are encoded and decoded in the same mode.
[0739] c. Alternatively, when W equals T2 and H equals T3 (e.g., T2 = 4, T3 = 8), the block above it is unavailable.
[0740] d. Alternatively, when W equals T2 and H equals T3 (e.g., T2 = 4, T3 = 8), the block above it is unavailable, or the block above it is outside the current CTU.
[0741] 3) When W equals T4 and H equals T5 (e.g., T4 = 8, T5 = 4), the block to its left is available and its size is equal to A x B; the current block and its left block are encoded and decoded in a specific mode.
[0742] a. Alternatively, when W equals T4 and H equals T5 (e.g., T4 = 8, T5 = 4), the left-hand block is unavailable.
[0743] 4) When W*H is not greater than T1 (e.g., 32), the current block is encoded and decoded in IBC AMVP and / or Merge mode; its upper and left neighboring blocks are available and have a size equal to AxB, and are encoded and decoded in a specific mode.
[0744] a. When W*H is not greater than T1 (e.g., 32), the current block is encoded / decoded in a specific mode; its left neighboring block is available, has a size equal to AxB, and is IBC encoded / decoded; and its upper neighboring block is available, within the same CTU, has a size equal to AxB, and is encoded / decoded in the same mode.
[0745] b. When W*H is not greater than T1 (e.g., 32), the current block is encoded and decoded in a specific mode; its left neighboring block is unavailable; and its upper neighboring block is available, within the same CTU, and of size AxB, and encoded and decoded in the same mode.
[0746] c. When W*H is not greater than T1 (e.g., 32), the current block is encoded or decoded in a specific mode; its left neighboring block is unavailable; and its upper neighboring block is unavailable.
[0747] d. When W*H is not greater than T1 (e.g., 32), the current block is encoded and decoded in a specific mode; its left neighboring block is available, has a size equal to AxB, and is encoded and decoded in the same mode; and its upper neighboring block is unavailable.
[0748] e. When W*H is not greater than T1 (e.g., 32), the current block is encoded or decoded in a specific mode; its left neighboring block is unavailable; and its upper neighboring block is unavailable or outside the current CTU.
[0749] f. When W*H is not greater than T1 (e.g., 32), the current block is encoded or decoded in a specific mode; its left neighboring block is available, has a size equal to AxB, and is encoded or decoded in the same mode; and its upper neighboring block is unavailable or outside the current CTU.
[0750] 5) In the above example, "specific mode" is IBC mode.
[0751] 6) In the above example, "specific mode" is inter-frame mode.
[0752] 7) In the example above, “AxB” can be set to 4x4.
[0753] 8) In the above example, "the size of the neighboring block is equal to AxB" can be replaced with "the size of the neighboring block is not greater than or not less than AxB".
[0754] 9) In the example above, the upper and left adjacent blocks are the two blocks visited for the spatial Merge candidate derivation.
[0755] a. In one example, suppose the coordinates of the top left sample point in the current block are (x, y), and the left block is a block that covers (x-1, y+H-1).
[0756] b. In one example, suppose the coordinates of the top left sample point in the current block are (x, y), and the left block is a block that covers (x+W-1, y-1).
[0757] t. The above thresholds can be predefined or signaled.
[0758] i. Alternatively, the threshold may depend on the block's encoding / decoding information, such as the encoding / decoding mode.
[0759] u. In one example, condition C is satisfied when the current block is under a shared node and the current block is encoded / decoded in IBC AMVP and / or Merge mode.
[0760] i. Alternatively, when condition C is met, the IBC motion list construction process may include candidates from spatially neighboring blocks (e.g., A1, B1) as well as default candidates. That is, the insertion of HMVP candidates is skipped.
[0761] ii. Alternatively, when condition C is met, the IBC motion list construction process may include candidates from the HMVP candidates in the IBCHMVP table, as well as default candidates. That is, the insertion of candidates from spatially neighboring blocks is skipped.
[0762] iii. Alternatively, after decoding a block if condition C is met, the update of the IBC HMVP table is skipped.
[0763] v. In one example, condition C can be adaptively changed, such as based on the block's encoding / decoding information.
[0764] i. In one example, condition C can be defined based on the encoding / decoding mode (IBC or non-IBC mode) and the block dimension.
[0765] Whether or not the above method is applied may depend on the block's encoding and decoding information, such as whether it is an IBC encoding / decoding block.
[0766] i. In one example, the above method can be applied when the block is encoded or decoded by IBC.
[0767] IBC Sports List
[0768] 22. Propose that motion candidates in the IBC HMVP table be stored with integer pixel precision instead of 1 / 16 pixel precision.
[0769] a. In one example, all motion candidates are stored with 1-pixel precision.
[0770] b. In one example, when using motion information from spatially adjacent (adjacent or non-adjacent) blocks and / or from the IBC HMVP table, the rounding process for MV is skipped.
[0771] 23. The list of IBC proposals may only contain proposal candidates from one or more IBC HMVP tables.
[0772] a. Alternatively, signaling notifications for candidates in the IBC motion list may depend on the number of available HMVP candidates in the HMVP table.
[0773] b. Alternatively, the signaling notification of candidates in the IBC motion list may depend on the maximum number of HMVP candidates in the HMVP table.
[0774] c. Alternatively, HMVP candidates in the HMVP table are added to the list sequentially without pruning.
[0775] i. In one example, the order is based on the ascending order of the table's entry indexes.
[0776] ii. In one example, the order is based on the descending order of the table's entry indexes.
[0777] iii. In one example, the first N entries in the table can be skipped.
[0778] iv. In one example, the last N entries in the table can be skipped.
[0779] v. In one example, an entry with (multiple) invalid BVs can be skipped.
[0780] vi.
[0781] d. Alternatively, motion candidates derived from HMVP candidates from one or more HMVP tables can be further modified, such as by adding offsets to the horizontal vector and / or the vertical vector.
[0782] i. An HMVP candidate with (multiple) invalid BVs can be modified to provide (multiple) valid BVs.
[0783] e. Alternatively, a default motion candidate can be added after or before one or more HMVP candidates.
[0784] f. How / whether to add HMVP candidates to the IBC motion list can depend on the block dimension.
[0785] i. For example, when the block dimensions (W and H representing width and height) satisfy condition C, the IBC motion list may contain only motion candidates from one or more HMVP tables.
[0786] 1) In one example, the condition C is W <= T1 and H <= T2, for example T1 = T2 = 4.
[0787] 2) In one example, condition C is W <= T1 or H <= T2, for example, T1 = T2 = 4.
[0788] 3) In one example, condition C is W*H <= T, for example, T = 16.
[0789] 5. Examples
[0790] Added changes are highlighted with underline, bold, and italic. Deleted parts are marked with [[]].
[0791] 5.1 Example #1
[0792] The HMVP table is not updated when the current block is under a shared node. A single IBCHMVP table is used only for blocks under a shared node.
[0793] 7.4.8.5 Semantics of Encoding / Decoding Units
[0794] When all of the following conditions are true, update the historical motion vector prediction list for the shared Merge candidate list region by setting NumHmvpSmrIbcCand to equal NumHmvpIbcCand and HmvpSmrIbcCandList[i] to equal HmvpIbcCandList[i] (for i = 0..NumHmvpIbcCand-1):
[0795] –IsInSmr[x0][y0] equals TRUE.
[0796] –SmrX[x0][y0] equals x0.
[0797] –SmrY[x0][y0] equals y0.
[0798] 8.6.2 Derivation of the motion vector components of the IBC block
[0799] 8.6.2.1 General
[0800] The input to this process is:
[0801] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0802] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0803] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0804] The output of this process is:
[0805] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[0806] The brightness motion vector mvL is derived as follows:
[0807] – Call the derivation procedure for IBC luminance motion vector prediction as specified in Clause 8.6.2.2, with luminance position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output as luminance motion vector mvL.
[0808] – When general_merge_flag[xCb][yCb] equals 0, the following applies:
[0809] 1. The variable mvd is derived as follows:
[0810] mvd[0]=MvdL0[xCb][yCb][0] (8-883)
[0811] mvd[1]=MvdL0[xCb][yCb][1] (8-884)
[0812] 2. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal mvL, rightShift set to equal MvShift+2, and leftShift set to equal MvShift+2 as inputs, and the rounded mvL as output.
[0813] 3. The brightness motion vector mvL is modified as follows:
[0814] u[0]=(mvL[0]+mvd[0]+2 18 )%2 18 (8-885)
[0815] mvL[0]=(u[0]>=2 17 )? (u[0]-2 18 ):u[0] (8-886)
[0816] u[1]=(mvL[1]+mvd[1]+2 18 )%2 18 (8-887)
[0817] mvL[1]=(u[1]>=2 17 )? (u[1]-2 18 ):u[1] (8-888)
[0818] Note 1 – The result values of mvL[0] and mvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[0819] When IsInSmr[xCb][yCb] is false The update process for the historical motion vector prediction list, as specified in Clause 8.6.2.6, is invoked using the brightness motion vector mvL.
[0820] The top left position (xRefTL, yRefTL) and bottom right position (xRefBR, yRefBR) inside the reference block are derived as follows:
[0821] (xRefTL,yRefTL)=(xCb+(mvL[0]>>4),yCb+(mvL[1]>>4)) (8-889)
[0822] (xRefBR,yRefBR)=(xRefTL+cbWidth-1,yRefTL+cbHeight-1) (8-890)
[0823] The requirement for bitstream consistency is that the luminance motion vector mvL should comply with the following constraints:
[0824] –...
[0825] 8.6.2.4 Derivation of IBC-based motion vector candidates based on history
[0826] The input to this process is:
[0827] – Motion vector candidate list mvCandList,
[0828] – The number of available motion vector candidates in the list, numCurrCand.
[0829] The output of this process is:
[0830] – Modified motion vector candidate list mvCandList
[0831] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[0832] – The number of modified motion vector candidates in the list, numCurrCand.
[0833] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[0834] The array smrHmvpIbcCandList and the variable smrNumHmvpIbcCand are derived as follows:
[0835] [[smr]]HmvpIbcCandList=[[isInSmr? HmvpSmrIbcCandList:]]HmvpIbcCandList (8-906)
[0836] [[smr]]NumHmvpIbcCand=[[isInSmr? NumHmvpSmrIbcCand:]]NumHmvpIbcCand(8-907)
[0837] For each candidate in smrHmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1..[[smr]]NumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumMergeCand:
[0838] 1. The variable sameMotion is derived as follows:
[0839] – if all of the following conditions are true for any motion vector candidate N (where N is A1 or B1), then sameMotion and isPrunedN are set to TRUE:
[0840] –hMvpIdx is less than or equal to 1.
[0841] – The candidate [[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] is equal to the candidate motion vector N.
[0842] –isPrunedN equals FALSE.
[0843] Otherwise, sameMotion is set to equal FALSE.
[0844] 2. When sameMotion equals FALSE, the candidate [[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] is added to the motion vector candidate list, as shown below:
[0845] mvCandList[numCurrCand++]=[[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] (8-908)
[0846] 5.2 Example #2
[0847] When the block size meets a specific condition (such as width * height < K), the check for spatial merge / AMVP candidates during the IBC motion list construction process is removed. In the following description, the threshold K can be predefined, such as 16.
[0848] 7.4.8.2 Semantics of Encoding / Decoding Tree Units
[0849] CTU is the root node of the codec tree structure.
[0850] The array IsInSmr[x][y], which specifies whether the sample at (x,y) is located within the shared Merge candidate list region, is initialized as follows for x = 0..CtbSizeY-1 and y = 0..CtbSizeY-1:
[0851] IsInSmr[x][y]=FALSE (7-96)]]
[0852] 7.4.8.4 Encoder-decoder tree semantics
[0853] IsInSmr[x][y] is set to TRUE for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, provided all of the following conditions are true:
[0854] –IsInSmr[x0][y0] equals FALSE
[0855] –cbWidth*cbHeight / 4 is less than 32
[0856] –treeType is not equal to DUAL_TREE_CHROMA
[0857] When IsInSmr[x0][y0] equals TRUE, for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the arrays SmrX[x][y], SmrY[x][y], SmrW[x][y], and SmrH[x][y] are derived as follows:
[0858] SmrX[x][y]=x0 (7-98)
[0859] SmrY[x][y]=y0 (7-99)
[0860] SmrW[x][y]=cbWidth (7-100)
[0861] SmrH[x][y]=cbHeight (7-101)
[0862] IsInSmr[x][y] is set to TRUE for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, provided all of the following conditions are true:
[0863] –IsInSmr[x0][y0] equals FALSE
[0864] –One of the following conditions is true:
[0865] –mtt_split_cu_binary_flag equals 1, and cbWidth*cbHeight / 2 is less than 32.
[0866] –mtt_split_cu_binary_flag equals 0, and cbWidth*cbHeight / 4 is less than 32.
[0867] –treeType is not equal to DUAL_TREE_CHROMA
[0868] When IsInSmr[x0][y0] equals TRUE, for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the arrays SmrX[x][y], SmrY[x][y], SmrW[x][y], and SmrH[x][y] are derived as follows:
[0869] SmrX[x][y]=x0 (7-102)
[0870] SmrY[x][y]=y0 (7-103)
[0871] SmrW[x][y]=cbWidth (7-104)
[0872] SmrH[x][y]=cbHeight (7-105)]]
[0873] 7.4.8.5 Semantics of Encoding / Decoding Units
[0874] When all of the following conditions are true, update the historical motion vector prediction list for the shared Merge candidate list region by setting NumHmvpSmrIbcCand to equal NumHmvpIbcCand and HmvpSmrIbcCandList[i] to equal HmvpIbcCandList[i] (for i = 0..NumHmvpIbcCand-1):
[0875] –IsInSmr[x0][y0] equals TRUE.
[0876] –SmrX[x0][y0] equals x0.
[0877] –SmrY[x0][y0] equals y0.
[0878] For x = x0..x0 + cbWidth - 1 and y = y0..y0 + cbHeight - 1, perform the following assignments:
[0879] CbPosX[x][y]=x0 (7-106)
[0880] CbPosY[x][y]=y0 (7-107)
[0881] CbWidth[x][y]=cbWidth (7-108)
[0882] CbHeight[x][y]=cbHeight (7-109)
[0883] 8.6.2 Derivation of the motion vector components of the IBC block
[0884] 8.6.2.1 General
[0885] The input to this process is:
[0886] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0887] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0888] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0889] The output of this process is:
[0890] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[0891] The brightness motion vector mvL is derived as follows:
[0892] – Call the derivation procedure for IBC luminance motion vector prediction as specified in Clause 8.6.2.2, with luminance position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output as luminance motion vector mvL.
[0893] – When general_merge_flag[xCb][yCb] equals 0, the following applies:
[0894] 4. The variable mvd is derived as follows:
[0895] mvd[0]=MvdL0[xCb][yCb][0] (8-883)
[0896] mvd[1]=MvdL0[xCb][yCb][1] (8-884)
[0897] 5. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal mvL, rightShift set to equal MvShift+2, and leftShift set to equal MvShift+2 as inputs, and the rounded mvL as output.
[0898] 6. The brightness motion vector mvL has been modified as follows:
[0899] u[0]=(mvL[0]+mvd[0]+2 18 )%2 18 (8-885)
[0900] mvL[0]=(u[0]>=2 17 )? (u[0]-2 18 ):u[0] (8-886)
[0901] u[1]=(mvL[1]+mvd[1]+2 18 )%2 18 (8-887)
[0902] mvL[1]=(u[1]>=2 17 )? (u[1]-2 18 ):u[1] (8-888)
[0903] Note 1 – The result values of mvL[0] and mvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[0904] When smrWidth*smrHeight is greater than K The update process for the historical motion vector prediction list, as specified in Clause 8.6.2.6, is invoked using the brightness motion vector mvL.
[0905] The top left position (xRefTL, yRefTL) and bottom right position (xRefBR, yRefBR) inside the reference block are derived as follows:
[0906] (xRefTL,yRefTL)=(xCb+(mvL[0]>>4),yCb+(mvL[1]>>4)) (8-889)
[0907] (xRefBR,yRefBR)=(xRefTL+cbWidth-1,yRefTL+cbHeight-1) (8-890)
[0908] The requirement for bitstream consistency is that the luminance motion vector mvL should comply with the following constraints:
[0909] –...
[0910] 8.6.2.2 Derivation of IBC Luminance Motion Vector Prediction
[0911] This procedure is invoked only when CuPredMode[xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top left sample of the current luminance codec block relative to the top left luminance sample of the current image.
[0912] The input to this process is:
[0913] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0914] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0915] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0916] The output of this process is:
[0917] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[0918] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows:
[0919] xSmr=[[IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:]]xCb (8-895)
[0920] ySmr=[[IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:]]yCb (8-896)
[0921] smrWidth=[[IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:]]cbWidth (8-897)
[0922] smrHeight=[[IsInSmr[xCb][yCb]? SmrH[xCb][yCb]:]]cbHeight (8-898)
[0923] smrNumHmvpIbcCand=[[IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand:]]NumHmvpIbcCand (8-899)
[0924] The brightness motion vector mvL is derived through the following ordered steps:
[0925] 1. When smrWidth*smrHeight is greater than K The process calls the derivation procedure for spatial motion vector candidates from neighboring codec units as specified in Clause 8.6.2.3, where the luma codec block position (xCb, yCb) is set to be equal to (xSmr, ySmr), the luma codec block width cbWidth and luma codec block height cbHeight are set to be equal to smrWidth and smrHeight, and the output is the availability flags availableFlagA1 and availableFlagB1, and the motion vectors mvA1 and mvB1.
[0926] 2. When smrWidth*smrHeight is greater than K The candidate list of motion vectors, mvCandList, is constructed as follows:
[0927]
[0928] 3. When smrWidth*smrHeight is greater than K, the variable numCurrCand is set to be equal to the number of Merge candidates in mvCandList.
[0929] 4. When numCurrCand is less than MaxNumMergeCand and smrNumHmvpIbcCand is greater than 0, call the IBC history-based motion vector candidate derivation process as specified in 8.6.2.4, with mvCandList, isInSmr set to be equal to IsInSmr[xCb][yCb] and numCurrCand as input, and the modified mvCandList and numCurrCand as output.
[0930] 5. When numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand equals MaxNumMergeCand:
[0931] 1. mvCandList[numCurrCand][0] is set to equal to 0.
[0932] 2. mvCandList[numCurrCand][1] is set to equal to 0.
[0933] 3.numCurrCand increases by 1.
[0934] 6. The variable mvIdx is derived as follows:
[0935] mvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-901)
[0936] 7. Perform the following assignment:
[0937] mvL[0]=mergeCandList[mvIdx][0] (8-902)
[0938] mvL[1]=mergeCandList[mvIdx][1] (8-903)
[0939] 8.6.2.4 Derivation of IBC-based motion vector candidates based on history
[0940] The input to this process is:
[0941] – Motion vector candidate list mvCandList,
[0942] – The number of available motion vector candidates in the list, numCurrCand.
[0943] The output of this process is:
[0944] – Modified motion vector candidate list mvCandList
[0945] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[0946] – The number of modified motion vector candidates in the list, numCurrCand.
[0947] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[0948] The array smrHmvpIbcCandList and the variable smrNumHmvpIbcCand are derived as follows:
[0949] [[smr]]HmvpIbcCandList=[[isInSmr? HmvpSmrIbcCandList:]]HmvpIbcCandList (8-906)
[0950] [[smr]]NumHmvpIbcCand=[[isInSmr? NumHmvpSmrIbcCand:]]NumHmvpIbcCand(8-907)
[0951] For each candidate in [[smr]]HmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1..smrNumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumMergeCand:
[0952] 1. The variable sameMotion is derived as follows:
[0953] -if smrWidth*smrHeight is greater than K If all of the following conditions are true for any motion vector candidate N (where N is A1 or B1), then sameMotion and isPrunedN are set to TRUE:
[0954] –hMvpIdx is less than or equal to 1.
[0955] – The candidate [[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] is equal to the candidate motion vector N.
[0956] –isPrunedN equals FALSE.
[0957] Otherwise, sameMotion is set to equal FALSE.
[0958] 2. When sameMotion equals FALSE, the candidate [[smr]]HmvpIbcCandList[smrNumHmvpIbcCand-hMvpIdx] is added to the motion vector candidate list, as shown below:
[0959] mvCandList[numCurrCand++]=[[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] (8-908)
[0960] 5.3 Example #3
[0961] When the block size meets certain conditions, such as the current block being under a shared node, the check for spatial merge / AMVP candidates during the IBC motion list construction process is removed, and the HMVP table is not updated.
[0962] 8.6.2.2 Derivation of IBC Luminance Motion Vector Prediction
[0963] This procedure is invoked only when CuPredMode[xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top left sample of the current luminance codec block relative to the top left luminance sample of the current image.
[0964] The input to this process is:
[0965] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[0966] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[0967] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[0968] The output of this process is:
[0969] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[0970] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows:
[0971] The variables xSmr, ySmr, smrWidth, smrHeight, and smrNumHmvpIbcCand are derived as follows:
[0972] xSmr=[[IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:]]xCb (8-895)
[0973] ySmr=[[IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:]]yCb (8-896)
[0974] smrWidth=[[IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:]]cbWidth (8-897)
[0975] smrHeight=[[IsInSmr[xCb][yCb]? SmrH[xCb][yCb]:]]cbHeight (8-898)
[0976] smrNumHmvpIbcCand=[[IsInSmr[xCb][yCb]? NumHmvpSmrIbcCand:]]NumHmvpIbcCand (8-899)
[0977] The brightness motion vector mvL is derived through the following ordered steps:
[0978] 1. When IsInSmr[xCb][yCb] is false The process calls the derivation procedure for spatial motion vector candidates from neighboring codec units as specified in Clause 8.6.2.3, where the luma codec block position (xCb, yCb) is set to be equal to (xSmr, ySmr), the luma codec block width cbWidth and luma codec block height cbHeight are set to be equal to smrWidth and smrHeight, and the output is the availability flags availableFlagA1 and availableFlagB1, and the motion vectors mvA1 and mvB1.
[0979] 2. When IsInSmr[xCb][yCb] is false The candidate list of motion vectors, mvCandList, is constructed as follows:
[0980]
[0981] 3. When IsInSmr[xCb][yCb] is false The variable numCurrCand is set to be equal to the number of Merge candidates in mvCandList.
[0982] 4. When numCurrCand is less than MaxNumMergeCand and smrNumHmvpIbcCand is greater than 0, call the IBC history-based motion vector candidate derivation process as specified in 8.6.2.4, with mvCandList, isInSmr set to be equal to IsInSmr[xCb][yCb] and numCurrCand as input, and the modified mvCandList and numCurrCand as output.
[0983] 5. When numCurrCand is less than MaxNumMergeCand, the following applies until numCurrCand equals MaxNumMergeCand:
[0984] 1. mvCandList[numCurrCand][0] is set to equal to 0.
[0985] 2. mvCandList[numCurrCand][1] is set to equal to 0.
[0986] 3.numCurrCand increases by 1.
[0987] 6. The variable mvIdx is derived as follows:
[0988] mvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-901)
[0989] 7. Perform the following assignment:
[0990] mvL[0]=mergeCandList[mvIdx][0] (8-902)
[0991] mvL[1]=mergeCandList[mvIdx][1] (8-903)
[0992] 8.6.2.4 Derivation of IBC-based motion vector candidates based on history
[0993] The input to this process is:
[0994] – Motion vector candidate list mvCandList,
[0995] – The number of available motion vector candidates in the list, numCurrCand.
[0996] The output of this process is:
[0997] – Modified motion vector candidate list mvCandList
[0998] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[0999] – The number of modified motion vector candidates in the list, numCurrCand.
[1000] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[1001] The array smrHmvpIbcCandList and the variable smrNumHmvpIbcCand are derived as follows:
[1002] [[smr]]HmvpIbcCandList=[[isInSmr? HmvpSmrIbcCandList:]]HmvpIbcCandList (8-906)
[1003] [[smr]]NumHmvpIbcCand=[[isInSmr? NumHmvpSmrIbcCand:]]NumHmvpIbcCand(8-907)
[1004] For each candidate in [[smr]]HmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1..smrNumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumMergeCand:
[1005] 3. The variable sameMotion is derived as follows:
[1006] -if isInSmr is false, and If all of the following conditions are true for any motion vector candidate N (where N is A1 or B1), then sameMotion and isPrunedN are set to TRUE:
[1007] –hMvpIdx is less than or equal to 1.
[1008] – The candidate [[smr]]HmvpIbcCandList[[smr]]NumHmvpIbcCand-hMvpIdx] is equal to the candidate motion vector N.
[1009] –isPrunedN equals FALSE.
[1010] Otherwise, sameMotion is set to equal FALSE.
[1011] 4. When sameMotion equals FALSE, the candidate [[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] is added to the motion vector candidate list, as shown below:
[1012] mvCandList[numCurrCand++]=[[smr]]HmvpIbcCandList[[[smr]]NumHmvpIbcCand-hMvpIdx] (8-908)
[1013] 8.6.2 Derivation of the motion vector components of the IBC block
[1014] 8.6.2.1 General
[1015] The input to this process is:
[1016] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1017] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1018] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1019] The output of this process is:
[1020] The brightness motion vector mvL with a precision of -1 / 16 fractional sample points.
[1021] The brightness motion vector mvL is derived as follows:
[1022] – Call the derivation procedure for IBC luminance motion vector prediction as specified in Clause 8.6.2.2, with luminance position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output as luminance motion vector mvL.
[1023] – When general_merge_flag[xCb][yCb] equals 0, the following applies:
[1024] 7. The variable mvd is derived as follows:
[1025] mvd[0]=MvdL0[xCb][yCb][0] (8-883)
[1026] mvd[1]=MvdL0[xCb][yCb][1] (8-884)
[1027] 8. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal mvL, rightShift set to equal MvShift+2, and leftShift set to equal MvShift+2 as inputs, and the rounded mvL as output.
[1028] 9. The brightness motion vector mvL is modified as follows:
[1029] u[0]=(mvL[0]+mvd[0]+2 18 )%2 18 (8-885)
[1030] mvL[0]=(u[0]>=2 17 )? (u[0]-2 18 ):u[0] (8-886)
[1031] u[1]=(mvL[1]+mvd[1]+2 18 )%2 18 (8-887)
[1032] mvL[1]=(u[1]>=2 17 )? (u[1]-2 18 ):u[1] (8-888)
[1033] Note 1 – The result values of mvL[0] and mvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[1034] When IsInSmr[xCb][yCb] is false The update process for the historical motion vector prediction list, as specified in Clause 8.6.2.6, is invoked using the brightness motion vector mvL.
[1035] The top left position (xRefTL, yRefTL) and bottom right position (xRefBR, yRefBR) inside the reference block are derived as follows:
[1036] (xRefTL,yRefTL)=(xCb+(mvL[0]>>4),yCb+(mvL[1]>>4)) (8-889)
[1037] (xRefBR,yRefBR)=(xRefTL+cbWidth-1,yRefTL+cbHeight-1) (8-890)
[1038] The requirement for bitstream consistency is that the luminance motion vector mvL should comply with the following constraints:
[1039] –...
[1040] 5.4 Example #4
[1041] When the block size meets specific conditions (such as width * height <= K, or width = N, height = 4 and the left neighboring block is 4x4 and encoded / decoded in IBC mode, or width = 4, height = N and the upper neighboring block is 4x4 and encoded / decoded in IBC mode), the spatial merge / AMVP candidate check is removed during the IBC motion list construction process, and the HMVP table is not updated. In the following description, the threshold K can be predefined, such as 16, and N can be predefined, such as 8.
[1042] 7.4.9.2 Semantics of Encoding / Decoding Tree Units
[1043] CTU is the root node of the codec tree structure.
[1044] The array IsAvailable[cIdx][x][y], which specifies whether the sample at (x,y) is available for the derivation of the availability of the neighboring block as specified in Clause 6.4.4, is initialized as follows for cIdx = 0..2, x = 0..CtbSizeY-1, and y = 0..CtbSizeY-1:
[1045] IsAvailable[cIdx][x][y]=FALSE (7-123)
[1046] The array IsInSmr[x][y], which specifies whether the sample at (x,y) is located within the shared Merge candidate list region, is initialized as follows for x = 0..CtbSizeY-1 and y = 0..CtbSizeY-1:
[1047] IsInSmr[x][y]=FALSE (7-124)]]
[1048] 7.4.9.4 Encoder-decoder tree semantics
[1049] IsInSmr[x][y] is set to TRUE for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, provided all of the following conditions are true:
[1050] –IsInSmr[x0][y0] equals FALSE
[1051] –cbWidth*cbHeight / 4 is less than 32
[1052] –treeType is not equal to DUAL_TREE_CHROMA
[1053] When IsInSmr[x0][y0] equals TRUE, for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the arrays SmrX[x][y], SmrY[x][y], SmrW[x][y], and SmrH[x][y] are derived as follows:
[1054] SmrX[x][y]=x0 (7-126)
[1055] SmrY[x][y]=y0 (7-127)
[1056] SmrW[x][y]=cbWidth (7-128)
[1057] SmrH[x][y]=cbHeight (7-129)]]
[1058] 8.6.2 Derivation of the block vector components of the IBC block
[1059] 8.6.2.1 General
[1060] The input to this process is:
[1061] - The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1062] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1063] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1064] The output of this process is:
[1065] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1066] The luminance block vector mvL is derived as follows:
[1067] - Call the derivation procedure for IBC luma block vector prediction as specified in Clause 8.6.2.2, with luma position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output luma block vector bvL.
[1068] - When general_merge_flag[xCb][yCb] equals 0, the following applies:
[1069] 1. The variable bvd is derived as follows:
[1070] bvd[0]=MvdL0[xCb][yCb][0] (8-900)
[1071] bvd[1]=MvdL0[xCb][yCb][1] (8-901)
[1072] 2. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal bvL, rightShift set to equal AmvrShift, and leftShift set to equal AmvrShift as inputs, and the rounded bvL as output.
[1073] 3. The luma block vector bvL is modified as follows:
[1074] u[0]=(bvL[0]+bvd[0]+2 18 )%2 18 (8-902)
[1075] bvL[0]=(u[0]>=217)? (u[0]-218):u[0] (8-903)
[1076] u[1]=(bvL[1]+bvd[1]+218)%218 (8-904)
[1077] bvL[1]=(u[1]>=217)? (u[1]218:u[1] (8-905)
[1078] Note 1 - The result values of bvL[0] and bvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[1079] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1080] If IsLgrBlk is true, then when Cb Width equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1081] If IsLgrBlk is true, then when Cb Width equals 4 and CbHeight equals N, and the upper neighboring block is 4x4. When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1082] (Or, alternatively:)
[1083] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1084] If IsLgrBlk is true, then when Cb Width equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1085] If IsLgrBlk is true, then when Cb Width equals 4 and CbHeight equals N, and the adjacent block above is available, Furthermore, when it is 4x4 and encoded / decoded in IBC mode, IsLgrBlk is set to false. )
[1086] when IsLgrBlk When the value is true [[IsInSmr[xCb][yCb] equals false]], the update procedure for the list of historical block vector predictions is invoked using the luminance block vector bvL, as specified in Clause 8.6.2.6.
[1087] Bitstream consistency requires that the luminance block vector bvL must comply with the following constraints:
[1088] -CtbSizeY is greater than or equal to ((yCb+(bvL[1]>>4))&(CtbSizeY-1))+cbHeight.
[1089] -For x=xCb..xCb+cbWidth-1 and y=yCb..yCb+cbHeight-1, IbcVirBuf[0][(x+(bvL[0]>>4))&(IbcVirBufWidth-1)][(y+(bvL[1]>>4))&(CtbSizeY-1)] should not be equal to -1.
[1090] 8.6.2.2 Derivation of IBC Luminance Block Vector Prediction
[1091] This procedure is invoked only when CuPredMode[0][xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top left sample of the current luminance codec block relative to the top left luminance sample of the current image.
[1092] The input to this process is:
[1093] - The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1094] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1095] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1096] The output of this process is:
[1097] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1098] The variables xSmr, ySmr, smrWidth, and smrHeight are derived as follows:
[1099] xSmr=IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:xCb (8-906)
[1100] ySmr=IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:yCb (8-907)
[1101] smrWidth=IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:cbWidth (8-908)
[1102] smrHeight=IsInSmr[xCb][yCb]? SmrH[xCb][yCb]:cbHeight (8-909)]]
[1103] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1104] If IsLgrBlk is true, then when Cb Width equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1105] If IsLgrBlk is true, then when Cb Width equals 4 and CbHeight equals N, and the upper neighboring block is 4x4. When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1106] (Alternatively:)
[1107] The variable IslgrBlk is set to (cb Width × cbHeight greater than K? True: False).
[1108] If IsLgrBlk is true, then when Cb Width equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1109] If IsLgrBlk is true, then when CbWidth equals 4 and CbHeight equals N, and the adjacent block above is available, Furthermore, when it is 4x4 and encoded / decoded in IBC mode, IsLgrBlk is set to false. )
[1110] The luminance block vector bvL is derived through the following ordered steps:
[1111] 1. When IslgrBlk is true Invoke the derivation procedure for spatial block vector candidates from neighboring codec units as specified in Clause 8.6.2.3, where is set to equal to ( xCb, yCb The luminance codec block position (xCb, yCb) of [[xSmr, ySmr]] is set to equal to [[smr]]. Cb Width and [[smr]] Cb The luminance codec block width cbWidth and luminance codec block height cbHeight are taken as inputs, and the outputs are availability flags availableFlagA1 and availableFlagB1, as well as block vectors bvA1 and bvB1.
[1112] 2. When IslgrBlk is true The block vector candidate list bvCandList is constructed as follows:
[1113]
[1114] 3. When IslgrBlk is true The variable numCurrCand is set to be equal to the number of Merge candidates in bvCandList.
[1115] 4. When numCurrCand is less than MaxNumIbcMergeCand and NumHmvpIbcCand is greater than 0, invoke the IBC-based derivation process for historical block vector candidates as specified in 8.6.2.4, where bvCandList, and IsLgrBlk,It takes numCurrCand as input and outputs modified bvCandList and numCurrCand.
[1116] 5. When numCurrCand is less than MaxNumIbcMergeCand, the following applies until numCurrCand equals MaxNumIbcMergeCand:
[1117] 1. bvCandList[numCurrCand][0] is set to equal to 0.
[1118] 2. bvCandList[numCurrCand][1] is set to equal to 0.
[1119] 3.numCurrCand increases by 1.
[1120] 6. The variable bvIdx is derived as follows:
[1121] bvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-911)
[1122] 7. Perform the following assignment:
[1123] bvL[0]=bvCandList[mvIdx][0] (8-912)
[1124] bvL[1]=bvCandList[mvIdx][1] (8-913)
[1125] 8.6.2.3 Derivation of IBC Spatial Block Vector Candidates
[1126] The input to this process is:
[1127] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1128] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1129] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1130] The output of this process is as follows:
[1131] –Availability flags for neighboring codec units, availableFlagA1 and availableFlagB1,
[1132] – Block vectors bvA1 and bvB1 with 1 / 16 fractional sample precision from neighboring codec units
[1133] The following applies to the derivation of availableFlagA1 and mvA1:
[1134] – The luminance position (xNbA1, yNbA1) within the adjacent luminance codec block is set to equal to (xCb-1, yCb+cbHeight-1).
[1135] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbA1, yNbA1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableA1.
[1136] – The variables availableFlagA1 and bvA1 are derived as follows:
[1137] – If availableA1 equals FALSE, then availableFlagA1 is set to 0, and both components of bvA1 are set to 0.
[1138] Otherwise, availableFlagA1 is set to 1 and the following assignment is made:
[1139] bvA1=MvL0[xNbA1][yNbA1] (8-914)
[1140] The following applies to the derivation of availableFlagB1 and bvB1:
[1141] – The luminance position (xNbB1, yNbB1) within the adjacent luminance codec block is set to equal to (xCb+cbWidth-1, yCb-1).
[1142] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbB1, yNbB1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableB1.
[1143] – The variables availableFlagB1 and bvB1 are derived as follows:
[1144] – If one or more of the following conditions are true, availableFlagB1 is set to 0, and both components of bvB1 are set to 0:
[1145] –availableB1 equals FALSE.
[1146] –availableA1 equals TRUE, and the luminance locations (xNbA1,yNbA1) and (xNbB1,yNbB1) have the same block vector.
[1147] Otherwise, availableFlagB1 is set to 1 and the following assignment is made:
[1148] bvB1=MvL0[xNbB1][yNbB1] (8-915)
[1149] 8.6.2.4 Derivation of IBC-based Block Vector Candidates
[1150] The input to this process is:
[1151] –Block vector candidate list bvCandList
[1152] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[1153] – The variable IslgrBlk, which indicates a non-small block,
[1154] – The number of available block vector candidates in the list, numCurrCand.
[1155] The output of this process is:
[1156] – Modified block vector candidate list bvCandList,
[1157] – The number of modified motion vector candidates in the list, numCurrCand.
[1158] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[1159] For each candidate in [[smr]]HmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1.[[smr]]NumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumIbcMergeCand:
[1160] 1. The variable sameMotion is derived as follows:
[1161] -if IsLgrBlk is true, and If all of the following conditions are true for any candidate block vector N (where N is A1 or B1), then sameMotion and isPrunedN are set to TRUE:
[1162] –hMvpIdx is less than or equal to 1.
[1163] – Candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is equal to candidate block vector N.
[1164] –isPrunedN equals FALSE.
[1165] Otherwise, sameMotion is set to equal FALSE.
[1166] 2. When sameMotion equals FALSE, candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is added to the block vector candidate list, as shown below:
[1167] bvCandList[numCurrCand++]=HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx](8-916)
[1168] 5.5 Example #5
[1169] When the block size meets specific conditions (such as width = N, height = 4 and the left neighboring block is 4x4 and encoded / decoded in IBC mode, or width = 4, height = N and the upper neighboring block is 4x4 and encoded / decoded in IBC mode), the check for spatial merge / AMVP candidates during the IBC motion list construction process is removed, and the update of the HMVP table is removed. In the following description, N can be predefined, such as 4 or 8.
[1170] 7.4.9.2 Semantics of Encoding / Decoding Tree Units
[1171] CTU is the root node of the codec tree structure.
[1172] The array IsAvailable[cIdx][x][y], which specifies whether the sample at (x,y) is available for the derivation of the availability of the neighboring block as specified in Clause 6.4.4, is initialized as follows for cIdx = 0..2, x = 0..CtbSizeY-1, and y = 0..CtbSizeY-1:
[1173] IsAvailable[cIdx][x][y]=FALSE (7-123)
[1174] The array IsInSmr[x][y], which specifies whether the sample at (x,y) is located within the shared Merge candidate list region, is initialized as follows for x = 0..CtbSizeY-1 and y = 0..CtbSizeY-1:
[1175] IsInSmr[x][y]=FALSE (7-124)]]
[1176] 7.4.9.4 Encoder-decoder tree semantics
[1177] IsInSmr[x][y] is set to TRUE for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, provided all of the following conditions are true:
[1178] –IsInSmr[x0][y0] equals FALSE
[1179] –cbWidth*cbHeight / 4 is less than 32
[1180] –treeType is not equal to DUAL_TREE_CHROMA
[1181] When IsInSmr[x0][y0] equals TRUE, for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the arrays SmrX[x][y], SmrY[x][y], SmrW[x][y], and SmrH[x][y] are derived as follows:
[1182] SmrX[x][y]=x0 (7-126)
[1183] SmrY[x][y]=y0 (7-127)
[1184] SmrW[x][y]=cbWidth (7-128)
[1185] SmrH[x][y]=cbHeight (7-129)]]
[1186] 8.6.2 Derivation of the block vector components of the IBC block
[1187] 8.6.2.1 General
[1188] The input to this process is:
[1189] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1190] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1191] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1192] The output of this process is:
[1193] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1194] The luminance block vector mvL is derived as follows:
[1195] – Call the derivation procedure for IBC luma block vector prediction as specified in Clause 8.6.2.2, with luma position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output luma block vector bvL.
[1196] – When general_merge_flag[xCb][yCb] equals 0, the following applies:
[1197] 1. The variable bvd is derived as follows:
[1198] bvd[0]=MvdL0[xCb][yCb][0] (8-900)
[1199] bvd[1]=MvdL0[xCb][yCb][1] (8-901)
[1200] 2. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal bvL, rightShift set to equal AmvrShift, and leftShift set to equal AmvrShift as inputs, and the rounded bvL as output.
[1201] 3. The luma block vector bvL is modified as follows:
[1202] u[0]=(bvL[0]+bvd[0]+2 18 )%2 18 (8-902)
[1203] bvL[0]=(u[0]>=217)? (u[0]-218):u[0] (8-903)
[1204] u[1]=(bvL[1]+bvd[1]+218)%218 (8-904)
[1205] bvL[1]=(u[1]>=217)? (u[1]218:u[1] (8-905)
[1206] Note 1 - The result values of bvL[0] and bvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[1207] The variable IslgrBlk is set to true when one of the following conditions is true.
[1208] - When CbWidth equals N and CbHeight equals 4, and the left neighboring block is 4x4 and encoding is performed in IBC mode. During decoding.
[1209] - When CbWidth equals 4 and CbHeight equals N, and the upper neighboring block is 4x4 and encoding is performed in IBC mode. During decoding.
[1210] Otherwise, IsLgrBlk is set to false.
[1211] (Or, alternatively:)
[1212] The variable IslgrBlk is set to (CbWidth*CbHeight>16?true:false), and further checks are performed. The following content:
[1213] If IsLgrBlk is true, then when CbWidth equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1214] If IsLgrBlk is true, then when CbWidth equals 4 and CbHeight equals N, and the upper neighboring block is 4x4. When encoding and decoding in IBC mode, IsLgrBlk is set to false. )
[1215] when IsLgrBlk is true When [[IsInSmr[xCb][yCb] equals false]], the update procedure for the list of historical block vector predictions is invoked using the luminance block vector bvL, as specified in Clause 8.6.2.6.
[1216] Bitstream consistency requires that the luminance block vector bvL must comply with the following constraints:
[1217] -CtbSizeY is greater than or equal to ((yCb+(bvL[1]>>4))&(CtbSizeY-1))+cbHeight.
[1218] -For x=xCb..xCb+cbWidth-1 and y=yCb..yCb+cbHeight-1, IbcVirBuf[0][(x+(bvL[0]>>4))&(IbcVirBufWidth-1)][(y+(bvL[1]>>4))&(CtbSizeY-1)] should not be equal to -1.
[1219] 8.6.2.2 Derivation of IBC Luminance Block Vector Prediction
[1220] This procedure is invoked only when CuPredMode[0][xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top left sample of the current luminance codec block relative to the top left luminance sample of the current image.
[1221] The input to this process is:
[1222] - The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1223] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1224] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1225] The output of this process is:
[1226] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1227] The variables xSmr, ySmr, smrWidth, and smrHeight are derived as follows:
[1228] xSmr=IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:xCb (8-906)
[1229] ySmr=IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:yCb (8-907)
[1230] smrWidth=IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:cbWidth (8-908)
[1231] smrHeight=IslnSmr[xCb][yCb]? SmrH[xCb][yCb]:cbHeight (8-909)]]
[1232] The variable IslgrBlk is set to true when one of the following conditions is true.
[1233] - When CbWidth equals N and CbHeight equals 4, and the left neighboring block is 4x4 and encoding is performed in IBC mode. During decoding.
[1234] - When CbWidth equals 4 and CbHeight equals N, and the upper neighboring block is 4x4 and encoding is performed in IBC mode. During decoding.
[1235] Otherwise, IsLgrBlk is set to false.
[1236] (Or, alternatively:)
[1237] The variable IslgrBlk is set to (CbWidth*CbHeight>16?true:false), and further checks are performed. The following content:
[1238] If IsLgrBlk is true, then when CbWidth equals N and CbHeight equals 4, and the left neighboring block is 4x4... When encoding and decoding in IBC mode, IsLgrBlk is set to false.
[1239] If IsLgrBlk is true, then when CbWidth equals 4 and CbHeight equals N, and the upper neighboring block is 4x4. (And when encoding and decoding in IBC mode, IsLgrBlk is set to false.)
[1240] The luminance block vector bvL is derived through the following ordered steps:
[1241] 1. When IslgrBlk is true Invoke the derivation procedure for spatial block vector candidates from neighboring codec units as specified in Clause 8.6.2.3, where is set to equal to ( xCb, yCb The luminance codec block position (xCb, yCb) of [[xSmr, ySmr]] is set to equal to [[smr]]. Cb Width and [[smr]] Cb The luminance codec block width cbWidth and luminance codec block height cbHeight are taken as inputs, and the outputs are availability flags availableFlagA1 and availableFlagB1, as well as block vectors bvA1 and bvB1.
[1242] 2. When IslgrBlk is true The block vector candidate list bvCandList is constructed as follows:
[1243]
[1244] 3. When IslgrBlk is true The variable numCurrCand is set to be equal to the number of Merge candidates in bvCandList.
[1245] 4. When numCurrCand is less than MaxNumIbcMergeCand and NumHmvpIbcCand is greater than 0, invoke the IBC-based derivation process for historical block vector candidates as specified in 8.6.2.4, where bvCandList, and IsLgrBlk, It takes numCurrCand as input and outputs modified bvCandList and numCurrCand.
[1246] 5. When numCurrCand is less than MaxNumIbcMergeCand, the following applies until numCurrCand equals MaxNumIbcMergeCand:
[1247] 1. bvCandList[numCurrCand][0] is set to equal to 0.
[1248] 2. bvCandList[numCurrCand][1] is set to equal to 0.
[1249] 3.numCurrCand increases by 1.
[1250] 6. The variable bvIdx is derived as follows:
[1251] bvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-911)
[1252] 7. Perform the following assignment:
[1253] bvL[0]=bvCandList[mvIdx][0] (8-912)
[1254] bvL[1]=bvCandList[mvIdx][1] (8-913)
[1255] 8.6.2.3 Derivation of IBC Spatial Block Vector Candidates
[1256] The input to this process is:
[1257] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1258] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1259] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1260] The output of this process is as follows:
[1261] –Availability flags for neighboring codec units, availableFlagA1 and availableFlagB1,
[1262] – Block vectors bvA1 and bvB1 with 1 / 16 fractional sample precision from neighboring codec units
[1263] The following applies to the derivation of availableFlagA1 and mvA1:
[1264] – The luminance position (xNbA1, yNbA1) within the adjacent luminance codec block is set to equal to (xCb-1, yCb+cbHeight-1).
[1265] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbA1, yNbA1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableA1.
[1266] – The variables availableFlagA1 and bvA1 are derived as follows:
[1267] – If availableA1 equals FALSE, then availableFlagA1 is set to 0, and both components of bvA1 are set to 0.
[1268] Otherwise, availableFlagA1 is set to 1 and the following assignment is made:
[1269] bvA1=MvL0[xNbA1][yNbA1] (8-914)
[1270] The following applies to the derivation of availableFlagB1 and bvB1:
[1271] – The luminance position (xNbB1, yNbB1) within the adjacent luminance codec block is set to equal to (xCb+cbWidth-1, yCb-1).
[1272] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbB1, yNbB1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableB1.
[1273] – The variables availableFlagB1 and bvB1 are derived as follows:
[1274] – If one or more of the following conditions are true, availableFlagB1 is set to 0, and both components of bvB1 are set to 0:
[1275] –availableB1 equals FALSE.
[1276] –availableA1 equals TRUE, and the luminance locations (xNbA1,yNbA1) and (xNbB1,yNbB1) have the same block vector.
[1277] Otherwise, availableFlagB1 is set to 1 and the following assignment is made:
[1278] bvB1=MvL0[xNbB1][yNbB1] (8-915)
[1279] 8.6.2.4 Derivation of IBC-based Block Vector Candidates
[1280] The input to this process is:
[1281] –Block vector candidate list bvCandList
[1282] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[1283] – The variable IslgrBlk, which indicates a non-small block,
[1284] – The number of available block vector candidates in the list, numCurrCand.
[1285] The output of this process is:
[1286] – Modified block vector candidate list bvCandList,
[1287] – The number of modified motion vector candidates in the list, numCurrCand.
[1288] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[1289] For each candidate in [[smr]]HmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1.[[smr]]NumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumIbcMergeCand:
[1290] 1. The variable sameMotion is derived as follows:
[1291] -if IsLgrBlk is true, and If all of the following conditions are true for any candidate block vector N (where N is A1 or B1), then sameMotion and isPrunedN are set to TRUE:
[1292] –hMvpIdx is less than or equal to 1.
[1293] – Candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is equal to candidate block vector N.
[1294] –isPrunedN equals FALSE.
[1295] Otherwise, sameMotion is set to equal FALSE.
[1296] 2. When sameMotion equals FALSE, candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is added to the block vector candidate list, as shown below:
[1297] bvCandList[numCurrCand++]=HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx]–(8-916)
[1298] 5.6 Example #6
[1299] When the block size meets specific conditions (such as width * height <= K, and the left or top neighboring block is 4x4 and being encoded / decoded in IBC mode), the check for spatial merge / AMVP candidates during the IBC motion list construction process is removed, and the update of the HMVP table is removed. In the following description, the threshold K can be predefined, such as 16 or 32.
[1300] 7.4.9.2 Semantics of Encoding / Decoding Tree Units
[1301] CTU is the root node of the codec tree structure.
[1302] The array IsAvailable[cIdx][x][y], which specifies whether the sample at (x,y) is available for the derivation of the availability of the neighboring block as specified in Clause 6.4.4, is initialized as follows for cIdx = 0..2, x = 0..CtbSizeY-1, and y = 0..CtbSizeY-1:
[1303] IsAvailable[cIdx][x][y]=FALSE (7-123)
[1304] The array IsInSmr[x][y], which specifies whether the sample at (x,y) is located within the shared Merge candidate list region, is initialized as follows for x = 0..CtbSizeY-1 and y = 0..CtbSizeY-1:
[1305] IsInSmr[x][y]=FALSE (7-124)]]
[1306] 7.4.9.4 Encoder-decoder tree semantics
[1307] IsInSmr[x][y] is set to TRUE for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, provided all of the following conditions are true:
[1308] –IsInSmr[x0][y0] equals FALSE
[1309] –cbWidth*cbHeight / 4 is less than 32
[1310] –treeType is not equal to DUAL_TREE_CHROMA
[1311] When IsInSmr[x0][y0] equals TRUE, for x = x0..x0 + cbWidth-1 and y = y0..y0 + cbHeight-1, the arrays SmrX[x][y], SmrY[x][y], SmrW[x][y], and SmrH[x][y] are derived as follows:
[1312] SmrX[x][y]=x0 (7-126)
[1313] SmrY[x][y]=y0 (7-127)
[1314] SmrW[x][y]=cbWidth (7-128)
[1315] SmrH[x][y]=cbHeight (7-129)]]
[1316] 8.6.2 Derivation of the block vector components of the IBC block
[1317] 8.6.2.1 General
[1318] The input to this process is:
[1319] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1320] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1321] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1322] The output of this process is:
[1323] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1324] The luminance block vector mvL is derived as follows:
[1325] - Call the derivation procedure for IBC luma block vector prediction as specified in Clause 8.6.2.2, with luma position (xCb, yCb), variables cbWidth and cbHeight as inputs, and output luma block vector bvL.
[1326] - When general_merge_flag[xCb][yCb] equals 0, the following applies:
[1327] 1. The variable bvd is derived as follows:
[1328] bvd[0]=MvdL0[xCb][yCb][0] (8-900)
[1329] bvd[1]=MvdL0[xCb][yCb][1] (8-901)
[1330] 2. Call the rounding procedure for the motion vector as specified in Clause 8.5.2.14, with mvX set to equal bvL, rightShift set to equal AmvrShift, and leftShift set to equal AmvrShift as inputs, and the rounded bvL as output.
[1331] 3. The luma block vector bvL is modified as follows:
[1332] u[0]=(bvL[0]+bvd[0]+2 18 )%2 18 (8-902)
[1333] bvL[0]=(u[0]>=217)? (u[0]-218):u[0] (8-903)
[1334] u[1]=(bvL[1]+bvd[1]+218)%218 (8-904)
[1335] bvL[1]=(u[1]>=217)? (u[1]218:u[1] (8-905)
[1336] Note 1 - The result values of bvL[0] and bvL[1] specified above will always be in the range of -2. 17 Up to 2 17 The range is within -1 (inclusive).
[1337] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1338] If IsLgrBlk is true, then when the left neighboring block is 4x4 and encoding / decoding is performed in IBC mode, IsLgrBlk is set to false.
[1339] If IsLgrBlk is true, then when the upper neighboring block is 4x4 and encoding / decoding is performed in IBC mode, IsLgrBlk is set to false.
[1340] (Or, alternatively,
[1341] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1342] If IsLgrBlk is true, then when the left neighboring block is 4x4 and encoding / decoding is performed in IBC mode, IsLgrBlk is set to false.
[1343] If IsLgrBlk is true, then the block above it is in the same CTU as the current block, and is 4x4 and in the IBC. When encoding and decoding in this mode, IsLgrBlk is set to false. )
[1344] when IsLgrBlk is true When [[IsInSmr[xCb][yCb] equals false]], the update procedure for the list of historical block vector predictions is invoked using the luminance block vector bvL, as specified in Clause 8.6.2.6.
[1345] Bitstream consistency requires that the luminance block vector bvL must comply with the following constraints:
[1346] -CtbSizeY is greater than or equal to ((yCb+(bvL[1]>>4))&(CtbSizeY-1))+cbHeight.
[1347] -For x=xCb..xCb+cbWidth-1 and y=yCb..yCb+cbHeight-1, IbcVirBuf[0][(x+(bvL[0]>>4))&(IbcVirBufWidth-1)][(y+(bvL[1]>>4))&(CtbSizeY-1)] should not be equal to -1.
[1348] 8.6.2.2 Derivation of IBC Luminance Block Vector Prediction
[1349] This procedure is invoked only when CuPredMode[0][xCb][yCb] equals MODE_IBC, where (xCb, yCb) specifies the top left sample of the current luminance codec block relative to the top left luminance sample of the current image.
[1350] The input to this process is:
[1351] - The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1352] - The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1353] - The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1354] The output of this process is:
[1355] Luminance block vector bvL with a precision of -1 / 16 fractional sample points.
[1356] The variables xSmr, ySmr, smrWidth, and smrHeight are derived as follows:
[1357] xSmr=IsInSmr[xCb][yCb]? SmrX[xCb][yCb]:xCb (8-906)
[1358] ySmr=IsInSmr[xCb][yCb]? SmrY[xCb][yCb]:yCb (8-907)
[1359] smrWidth=IsInSmr[xCb][yCb]? SmrW[xCb][yCb]:cbWidth (8-908)
[1360] smrHeight=IsInSmr[xCb][yCb]? SmrH[xCb][yCb]:cbHeight (8-909)]]
[1361] The variable IslgrBlk is set to (cb Width × cbHeight greater than K? True: False).
[1362] If IsLgrBlk is true, then when the left neighboring block is 4x4 and encoding / decoding is performed in IBC mode, IsLgrBlk is set to false.
[1363] If IsLgrBlk is true, then when the upper neighboring block is 4x4 and encoding / decoding is performed in IBC mode, IsLgrBlk is set to false.
[1364] (Or, alternatively,
[1365] The variable IslgrBlk is set to (cbWidth × cbHeight greater than K? True: False).
[1366] If IsLgrBlk is true, then IsLgrBlk is set to false when the left neighboring block is 4x4 and is being encoded / decoded in IBC mode.
[1367] If IsLgrBlk is true, then IsLgrBlk is set to false when the preceding neighboring block is in the same CTU as the current block, is 4x4, and is being encoded / decoded in IBC mode.
[1368] The luminance block vector bvL is derived through the following ordered steps:
[1369] 1. When IslgrBlk is true Invoke the derivation procedure for spatial block vector candidates from neighboring codec units as specified in Clause 8.6.2.3, where is set to equal to ( xCb,yCb The luminance codec block position (xCb, yCb) of [[xSmr, ySmr]]) is set to equal to [[smr]]. Cb Width and [[smr]] Cb The luminance codec block width cbWidth and luminance codec block height cbHeight are taken as inputs, and the outputs are availability flags availableFlagA1 and availableFlagB1, as well as block vectors bvA1 and bvB1.
[1370] 2. When IslgrBlk is true The block vector candidate list bvCandList is constructed as follows:
[1371]
[1372] 3. When IslgrBlk is true The variable numCurrCand is set to be equal to the number of Merge candidates in bvCandList.
[1373] 4. When numCurrCand is less than MaxNumIbcMergeCand and NumHmvpIbcCand is greater than 0, invoke the IBC-based derivation process for historical block vector candidates as specified in 8.6.2.4, where bvCandList, and IsLgrBlk, It takes numCurrCand as input and outputs modified bvCandList and numCurrCand.
[1374] 5. When numCurrCand is less than MaxNumIbcMergeCand, the following applies until numCurrCand equals MaxNumIbcMergeCand:
[1375] 1. bvCandList[numCurrCand][0] is set to equal to 0.
[1376] 2. bvCandList[numCurrCand][1] is set to equal to 0.
[1377] 3.numCurrCand increases by 1.
[1378] 6. The variable bvIdx is derived as follows:
[1379] bvIdx=general_merge_flag[xCb][yCb]? merge_idx[xCb][yCb]:mvp_l0_flag[xCb][yCb] (8-911)
[1380] 7. Perform the following assignment:
[1381] bvL[0]=bvCandList[mvIdx][0] (8-912)
[1382] bvL[1]=bvCandList[mvIdx][1] (8-913)
[1383] 8.6.2.3 Derivation of IBC Spatial Block Vector Candidates
[1384] The input to this process is:
[1385] – The brightness position (xCb, yCb) of the top left sample of the current luminance block relative to the top left luminance sample of the current image.
[1386] – The variable cbWidth specifies the width of the current codec block in the luminance sample.
[1387] – The variable cbHeight specifies the height of the current codec block in the luminance sample.
[1388] The output of this process is as follows:
[1389] –Availability flags for neighboring codec units, availableFlagA1 and availableFlagB1,
[1390] – Block vectors bvA1 and bvB1 with 1 / 16 fractional sample precision from neighboring codec units
[1391] The following applies to the derivation of availableFlagA1 and mvA1:
[1392] – The luminance position (xNbA1, yNbA1) within the adjacent luminance codec block is set to equal to (xCb-1, yCb+cbHeight-1).
[1393] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbA1, yNbA1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableA1.
[1394] – The variables availableFlagA1 and bvA1 are derived as follows:
[1395] – If availableA1 equals FALSE, then availableFlagA1 is set to 0, and both components of bvA1 are set to 0.
[1396] Otherwise, availableFlagA1 is set to 1 and the following assignment is made:
[1397] bvA1=MvL0[xNbA1][yNbA1] (8-914)
[1398] The following applies to the derivation of availableFlagB1 and bvB1:
[1399] – The luminance position (xNbB1, yNbB1) within the adjacent luminance codec block is set to equal to (xCb+cbWidth-1, yCb-1).
[1400] – Call the derivation procedure for the availability of neighboring blocks as specified in Clause 6.4.4, with the current luminance position (xCurr, yCurr) set to equal to (xCb, yCb), the neighboring luminance position (xNbB1, yNbB1), checkPredModeY set to equal to TRUE, and cIdx set to equal to 0 as inputs, and assign the output to the block availability flag availableB1.
[1401] – The variables availableFlagB1 and bvB1 are derived as follows:
[1402] – If one or more of the following conditions are true, availableFlagB1 is set to 0, and both components of bvB1 are set to 0:
[1403] –availableB1 equals FALSE.
[1404] –availableA1 equals TRUE, and the luminance locations (xNbA1,yNbA1) and (xNbB1,yNbB1) have the same block vector.
[1405] Otherwise, availableFlagB1 is set to 1 and the following assignment is made:
[1406] bvB1=MvL0[xNbB1][yNbB1] (8-915)
[1407] 8.6.2.4 Derivation of IBC-based Block Vector Candidates
[1408] The input to this process is:
[1409] –Block vector candidate list bvCandList
[1410] –[[The variable isInSmr specifies whether the current codec unit is inside the shared Merge candidate region,]]
[1411] – The variable IslgrBlk, which indicates a non-small block,
[1412] – The number of available block vector candidates in the list, numCurrCand.
[1413] The output of this process is:
[1414] – Modified block vector candidate list bvCandList,
[1415] – The number of modified motion vector candidates in the list, numCurrCand.
[1416] Both variables isPrunedA1 and isPrunedB1 are set to FALSE.
[1417] For each candidate in [[smr]]HmvpIbcCandList[hMvpIdx] (where index hMvpIdx = 1.[[smr]]NumHmvpIbcCand), repeat the following ordered steps until numCurrCand equals MaxNumIbcMergeCand:
[1418] 1. The variable sameMotion is derived as follows:
[1419] -if IsLgrBlk is true, and All of the following conditions apply to any candidate block vector N
[1420] If both (where N is A1 or B1) are true, then both sameMotion and isPrunedN are set to TRUE:
[1421] –hMvpIdx is less than or equal to 1.
[1422] – Candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is equal to candidate block vector N.
[1423] –isPrunedN equals FALSE.
[1424] Otherwise, sameMotion is set to equal FALSE.
[1425] 2. When sameMotion equals FALSE, candidate HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx] is added to the block vector candidate list, as shown below:
[1426] bvCandList[numCurrCand++]=HmvpIbcCandList[NumHmvpIbcCand-hMvpIdx]–(8-916)
[1427] Figure 22This is a block diagram of a video processing apparatus 2200. Apparatus 2200 can be used to implement one or more of the methods described herein. Apparatus 2200 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 2200 may include one or more processors 2202, one or more memories 2204, and video processing hardware 2206. The processors(multiple) 2202 can be configured to implement one or more methods described in this document. The memories(multiple) 2204 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 2206 can be used to implement some of the techniques described in this document in hardware circuitry. The video processing hardware 2206 may be partially or wholly included within the processors(multiple) 2202 in the form of dedicated hardware or a graphics processing unit (GPU) or a dedicated signal processing block.
[1428] Figure 23 This is a flowchart illustrating an example of video processing method 2300. Method 2300 includes, in operation 2302, determining the sub-block intra-block copy (sbIBC) codec mode to use. Method 2300 also includes, in operation 2304, performing a conversion using the sbIBC codec mode.
[1429] Some embodiments can be described using the following terms-based description.
[1430] The following clauses illustrate example embodiments of the techniques discussed in Section 1 of the preceding chapter.
[1431] 1. A video processing method comprising: determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and a bitstream representation of the current video block, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region, wherein the size of the sub-blocks is based on a partitioning rule; and using the sbIBC encoding / decoding mode for the multiple sub-blocks to perform the conversion.
[1432] 2. The method according to Clause 1, wherein the current video block is an MxN block, where M and N are integers, and wherein the partitioning rule specifies that each sub-block has the same size.
[1433] 3. The method according to Clause 1, wherein the bitstream representation includes syntax elements indicating partitioning rules or the size of sub-blocks.
[1434] 4. The method according to any one of Clauses 1-3, wherein the partitioning rules specify the size of the sub-blocks based on the color components of the current video block.
[1435] 5. The method according to any one of clauses 1-4, wherein the sub-blocks of the first color component can derive their motion vector information from the sub-blocks of the second color component corresponding to the current video block.
[1436] The following clauses illustrate example embodiments of the techniques discussed in items 2 and 6 of the preceding section.
[1437] 1. A video processing method, comprising: determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and a bitstream representation of the current video block, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region; and performing the conversion using the sbIBC encoding / decoding mode for the multiple sub-blocks, wherein the conversion includes: determining an initial motion vector (initMV) for a given sub-block; identifying a reference block from the initMV; and deriving MV information for the given sub-block using motion vector (MV) information of the reference block.
[1438] 2. The method according to Clause 1, wherein determining initMV includes determining initMV from one or more neighboring blocks of a given sub-block.
[1439] 3. The method according to Clause 2, wherein one or more neighboring blocks are checked in sequence.
[1440] 4. The method according to Clause 1, wherein determining initMV includes deriving initMV from a list of motion candidates.
[1441] 5. The method according to any one of clauses 1-4, wherein identifying the reference block includes converting the initMV to a pixel precision and identifying the reference block based on the converted initMV.
[1442] 6. The method according to any one of clauses 1-4, wherein identifying the reference block includes applying initMV to an offset position within a given block, wherein the offset position is represented as an offset (offsetX, offsetY) from a predetermined position of a given sub-block.
[1443] 7. The method according to any one of clauses 1 to 6, wherein deriving the MV of a given sub-block includes MV information of a trimmed reference block.
[1444] 8. The method according to any one of Clauses 1 to 7, wherein the reference block is in a color component different from the color component of the current video block.
[1445] The following clauses illustrate example embodiments of the techniques discussed in items 3, 4, and 5 of the preceding section.
[1446] 1. A video processing method comprising: determining a sub-block intra-block copy (sbIBC) encoding / decoding mode for a conversion between a current video block in a video region and a bitstream representation of the current video block, wherein the current video block is divided into multiple sub-blocks and each sub-block is encoded / decoded based on reference samples from the video region; and performing the conversion using the sbIBC encoding / decoding mode for the multiple sub-blocks, wherein the conversion includes generating sub-block IBC candidates.
[1447] 2. The method according to Clause 1, wherein the sub-block IBC candidate is added to a candidate list that includes optional motion vector prediction value candidates.
[1448] 3. The method according to Clause 1, wherein the sub-block IBC candidate is added to the list including affine Merge candidates.
[1449] The following clauses illustrate exemplary embodiments of the techniques discussed in items 7, 8, 9, 10, 11, 12, and 13 of the preceding section.
[1450] 1. A video processing method, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into a plurality of sub-blocks, wherein the conversion includes processing a first sub-block among the plurality of sub-blocks using a sub-block intra-block codec (sbIBC) mode and processing a second sub-block among the plurality of sub-blocks using an intra-frame codec mode.
[1451] 2. A video processing method, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing all sub-blocks of the multiple sub-blocks using an intra-frame encoding / decoding mode.
[1452] 3. A video processing method, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into multiple sub-blocks, wherein the conversion includes processing all sub-blocks in the multiple sub-blocks using a palette encoding / decoding mode that uses a palette of representative pixel values for encoding / decoding each sub-block.
[1453] 4. A video processing method, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into a plurality of sub-blocks, wherein the conversion includes processing the first sub-block using a palette mode for encoding and decoding a first sub-block among the plurality of sub-blocks using a palette of representative pixel values, and processing the second sub-block among the plurality of sub-blocks using an intra-block copy encoding and decoding mode.
[1454] 5. A video processing method, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into a plurality of sub-blocks, wherein the conversion includes processing the first sub-block using a palette mode for encoding and decoding a first sub-block among the plurality of sub-blocks using a palette of representative pixel values, and processing a second sub-block among the plurality of sub-blocks using an intra-frame encoding and decoding mode.
[1455] 6. A method of video processing, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into a plurality of sub-blocks, wherein the conversion includes processing a first sub-block among the plurality of sub-blocks using a sub-block intra-block coding / decoding (sbIBC) mode and processing a second sub-block among the plurality of sub-blocks using an inter-frame coding / decoding mode.
[1456] 7. A method of video processing, comprising: performing a conversion between a bitstream representation of a current video block and a current video block divided into a plurality of sub-blocks, wherein the conversion includes processing a first sub-block among the plurality of sub-blocks using a sub-block intra-codec mode and processing a second sub-block among the plurality of sub-blocks using an inter-codec mode.
[1457] The following clauses illustrate example embodiments of the techniques discussed in item 14 of the previous section.
[1458] 8. The method according to any one of clauses 1-7, wherein the method further includes avoiding updating the IBC-based motion vector prediction table after the current video block has been transformed.
[1459] The following clauses illustrate example embodiments of the techniques discussed in item 15 of the previous chapter.
[1460] 9. The method described in any one or more of Clauses 1-7 further includes avoiding updating the non-IBC history-based motion vector prediction table after the current video block has been converted.
[1461] The following clauses illustrate example embodiments of the techniques discussed in item 16 of the previous section.
[1462] 10. The method according to any one of clauses 1-7, wherein the conversion includes the selective use of a loop filter based on the processing.
[1463] The following clauses illustrate example embodiments of the techniques discussed in Section 1 of the preceding chapter.
[1464] 17. The method according to any one of Clauses 1-7, wherein performing the transformation includes performing the transformation by disabling a specific codec mode for the current video block due to the use of the method, wherein the specific codec mode includes one or more of sub-block transform, affine motion prediction, multi-reference line intra-frame prediction, matrix-based intra-frame prediction, symmetric motion vector difference (MVD) codec, Merge with motion derivation / refinement on the MVD decoder side, bidirectional optical flow, simplified quadratic transform, or multiple transform sets.
[1465] The following clauses illustrate example embodiments of the techniques discussed in item 18 of the previous section.
[1466] 1. The method according to any one of the preceding clauses, wherein the indicator in the bitstream representation includes information about how to apply the method to the current video block.
[1467] The following clauses illustrate example embodiments of the techniques discussed in item 19 of the previous section.
[1468] 1. A video encoding method, comprising: deciding to encode a current video block into a bitstream representation using the method described in any one of the preceding clauses; and including information indicating the decision in the bitstream representation at the decoder parameter set level or sequence parameter set level or video parameter set level or picture parameter set level or picture header level or strip header level or slice header level or maximum codec unit level or codec unit level or maximum codec unit line level or LCU group level or transform unit level or prediction unit level or video codec unit level.
[1469] 2. A video encoding method, comprising: deciding to encode a current video block into a bitstream representation based on encoding conditions using any one of the preceding clauses; and performing encoding using any one of the preceding clauses, wherein the conditions are based on one or more of the following: encoding / decoding units, prediction units, transform units, the current video block, or the position of the video encoding / decoding unit of the current video block.
[1470] The block dimensions of the current video block and / or its neighboring blocks.
[1471] The block shape of the current video block and / or its neighboring blocks.
[1472] Intra-frame mode of the current video block and / or its neighboring blocks,
[1473] Motion / block vectors of neighboring blocks of the current video block;
[1474] The color format of the current video block;
[1475] Encoder / decoder tree structure;
[1476] The current video block's strip, segment, or image type;
[1477] The color components of the current video block;
[1478] The temporal layer ID of the current video block;
[1479] A table, level, or standard used for bitstream representation.
[1480] The following clauses illustrate example embodiments of the techniques discussed in item 20 of the previous section.
[1481] 1. A video processing method, comprising: determining a conversion between blocks in a video region and a bitstream representation of the video region using intra-block copy mode and inter-frame prediction mode; and performing the conversion using intra-block copy mode and inter-frame prediction mode for blocks in the video region.
[1482] 2. The method according to Clause 1, wherein the video area includes video images, video strips, video clips, or video segments.
[1483] 3. The method according to any one of Clauses 1-2, wherein the inter-frame prediction mode uses the Optional Motion Vector Prediction Value (AMVP) encoding / decoding mode.
[1484] 4. The method according to any one of clauses 1-3, wherein performing the conversion includes deriving a Merge candidate for the intra-block copy mode from neighboring blocks.
[1485] The following clauses illustrate example embodiments of the techniques discussed in item 21 of the previous section.
[1486] 1. A video processing method, comprising: during a conversion between a current video block and a bitstream representation of the current video block, performing a motion candidate list construction process and / or a table update process for updating a historical motion vector prediction table based on encoding / decoding conditions, and performing a conversion based on the motion candidate list construction process and / or the table update process.
[1487] 2. As described in Clause 1, different processes may be applied when the encoding / decoding conditions are met or not.
[1488] 3. As described in Clause 1, when the encoding / decoding conditions are met, the historical motion vector prediction table should not be updated.
[1489] 4. According to the method described in Clause 1, when the encoding / decoding conditions are met, the derivation of candidates from spatially adjacent (adjacent or non-adjacent) blocks is skipped.
[1490] 5. As described in Clause 1, the derivation of HMVP candidates is skipped when the encoding / decoding conditions are met.
[1491] 6. The encoding / decoding conditions described in Clause 1 include that the block width multiplied by the height is no greater than 16, 32, or 64.
[1492] 7. The encoding / decoding conditions described in Clause 1 include that the block is encoded / decoded in IBC mode.
[1493] 8. The method according to Clause 1, wherein the encoding and decoding conditions are as described in Section 21.bsiv of the preceding chapter.
[1494] The method according to any one of the foregoing clauses, wherein the transformation includes generating a bitstream representation from the current video block.
[1495] The method according to any one of the foregoing clauses, wherein the transformation includes generating samples of the current video block from the bitstream representation.
[1496] A video processing apparatus includes a processor configured to implement the method described in any one or more of the foregoing provisions.
[1497] A computer-readable medium having code stored thereon, which, when executed, causes a processor to perform the methods described in accordance with any one or more of the foregoing provisions.
[1498] Figure 24 This is a block diagram illustrating an example video processing system 2400 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 2400. System 2400 may include an input 2402 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or it may be in a compressed or encoded format. Input 2402 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.
[1499] System 2400 may include an encoding / decoding component 2404 capable of implementing the various encoding / decoding or coding methods described in this document. Encoding / decoding component 2404 can reduce the average bit rate of the video from input 2402 to its output to produce an encoded / decoded representation of the video. Encoding / decoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding / decoding component 2404 may be stored or transmitted via a communication connection, as represented by component 2406. The bitstream (or encoded / decoded) representation of the video received at input 2402, whether stored or communicated, can be used by component 2408 to generate pixel values or transmit as displayable video to display interface 2410. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although specific video processing operations are referred to as “encoding / decoding” operations or tools, it will be understood that encoding / decoding tools or operations are used at the encoder, and the corresponding decoding tools or operations that inversely represent the encoding / decoding results will be performed by the decoder.
[1500] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[1501] Figure 25 This is a flowchart representation of a method 2500 for video processing according to the present technology. Method 2500 includes, in operation 2510, determining that the current block is divided into multiple sub-blocks for a conversion between a current block of video and a bitstream representation of the video. At least one of the multiple blocks is encoded using a modified intra-block copy (IBC) encoding / decoding technique, wherein the technique uses reference samples from one or more video regions of the current block's current frame. Method 2500 includes, in operation 2520, performing a conversion based on this determination.
[1502] In some embodiments, the video region includes the current image, strip, slice, tile, or group of slices. In some embodiments, when the current block has a dimension of M×N, the current block is divided into multiple sub-blocks, where M and N are integers. In some embodiments, the multiple sub-blocks have the same size of L×K, where L and K are integers. In some embodiments, L = K. In some embodiments, L = 4 or K = 4.
[1503] In some embodiments, the multiple sub-blocks have different sizes. In some embodiments, the multiple sub-blocks have non-rectangular shapes. In some embodiments, the multiple sub-blocks have triangular or wedge shapes.
[1504] In some embodiments, the size of at least one sub-block among a plurality of sub-blocks is determined based on the size of the smallest decoding unit, smallest prediction unit, smallest transform unit, or smallest unit for motion information storage. In some embodiments, the size of at least one sub-block is expressed as (N1 × minW) × (N2 × minH), where minW × minH represents the size of the smallest decoding unit, prediction unit, transform unit, or unit for motion information storage, and where N1 and N2 are positive integers. In some embodiments, the size of at least one sub-block among a plurality of sub-blocks is based on the encoding / decoding mode in which the current block is encoded / decoded in the bitstream representation. In some embodiments, the encoding / decoding mode includes at least an intra-block copy (IBC) merge mode or a sub-block temporal motion vector prediction mode. In some embodiments, the size of at least one sub-block among a plurality of sub-blocks is signaled in the bitstream representation.
[1505] In some embodiments, the method includes determining that subsequent blocks of video are divided into a plurality of sub-blocks for the transformation, wherein a first sub-block in the current block has a different size than a second sub-block in the subsequent block. In some embodiments, the size of the first sub-block differs from the size of the second sub-block based on the dimensions of the current block and the dimensions of the subsequent blocks. In some embodiments, the size of at least one of the plurality of sub-blocks is based on the color format or color components of the video. In some embodiments, the first sub-block is associated with a first color component of the video, and the second sub-block is associated with a second color component of the video, the first and second sub-blocks having different dimensions. In some embodiments, when the color format of the video is 4:2:0, the first sub-block associated with the luma component has a 2L×2K dimension, and the second sub-block associated with the chroma component has an L×K dimension. In some embodiments, when the color format of the video is 4:2:2, the first sub-block associated with the luma component has a 2L×2K dimension, and the second sub-block associated with the chroma component has a 2L×K dimension. In some embodiments, the first sub-block is associated with a first color component of the video, and the second sub-block is associated with a second color component of the video, the first and second sub-blocks having the same dimensions. In some embodiments, when the color format of the video is 4:2:0 or 4:4:4, the first sub-block associated with the luminance component has a dimension of 2L×2K, and the second sub-block associated with the chrominance component has a dimension of 2L×2K.
[1506] In some embodiments, the motion vector of a first sub-block associated with a first color component of the video is determined based on one or more sub-blocks associated with a second color component of the video. In some embodiments, the motion vector of the first sub-block is the average of the motion vectors of one or more sub-blocks associated with the second color component. In some embodiments, the current block is divided into multiple sub-blocks based on a single-tree segmentation structure. In some embodiments, the current block is associated with the chroma components of the video. In some embodiments, the current block has a 4×4 size.
[1507] In some embodiments, motion information of at least one sub-block among a plurality of sub-blocks is determined based on identifying a reference block based on an initial motion vector and determining the motion information of the sub-block based on the reference block. In some embodiments, the reference block is located within the current image. In some embodiments, the reference block is located within a reference image among one or more reference images. In some embodiments, at least one of the one or more reference images is the current image. In some embodiments, the reference block is located within a juxtaposed reference image juxtaposed with one or more reference images according to temporal information. In some embodiments, the reference image is determined based on motion information of the juxtaposed block or its neighboring blocks. The juxtaposed block is juxtaposed with the current block according to temporal information. In some embodiments, the initial motion vector of the reference block is determined based on one or more neighboring blocks of the current block or one or more neighboring blocks of a sub-block. In some embodiments, the one or more neighboring blocks include adjacent blocks and / or non-adjacent blocks of the current block or sub-block. In some embodiments, at least one neighboring block among one or more neighboring blocks and the reference block are located within the same image. In some embodiments, at least one neighboring block among one or more neighboring blocks is located within a reference image among one or more reference images. In some embodiments, at least one neighboring block among one or more neighboring blocks is located within a juxtaposed reference image juxtaposed with one or more reference images according to temporal information. In some embodiments, the reference image is determined based on motion information of the juxtaposed block or neighboring blocks of the juxtaposed block, wherein the juxtaposed block is juxtaposed with the current block according to temporal information.
[1508] In some embodiments, the initial motion vector is equal to the motion vector stored in one or more neighboring blocks. In some embodiments, the initial motion vector is determined based on the order in which one or more neighboring blocks are checked for the transformation. In some embodiments, the initial motion vector is a first identified motion vector associated with the current image. In some embodiments, the initial motion vector of a reference block is determined based on a motion candidate list. In some embodiments, the motion candidate list includes an intra-block copy (IBC) candidate list, a merge candidate list, a sub-block temporal motion vector prediction candidate list, or a history-based motion candidate list determined based on past motion prediction results. In some embodiments, the initial motion vector is determined based on a selected candidate in the list. In some embodiments, the selected candidate is the first candidate in the list. In some embodiments, the candidate list is constructed based on a process using spatial neighboring blocks that differ from those used in conventional construction processes.
[1509] In some embodiments, the initial motion vector of a reference block is determined based on the position of the current block in the current image. In some embodiments, the initial motion vector of a reference block is determined based on the dimension of the current block. In some embodiments, the initial motion vector is set to a default value. In some embodiments, the initial motion vector is indicated in the bitstream representation at the video unit level. In some embodiments, a video unit includes a slice, strip, picture, tile, codec tree unit (CTU) row, CTU, codec tree block (CTB), codec unit (CU), prediction unit (PU), or transform unit (TU). In some embodiments, the initial motion vector of a sub-block differs from that of another initial motion block of a second sub-block of the current block. In some embodiments, the initial motion vectors of multiple sub-blocks of the current block are determined differently based on the video unit. In some embodiments, a video unit includes a block, slice, or strip.
[1510] In some embodiments, before identifying the reference block, the initial motion vector is converted to a pixel-integer precision of F, where F is a positive integer greater than or equal to 1. In some embodiments, F is 1, 2, or 4. In some embodiments, the initial motion vector is represented as (vx, vy), and the converted motion vector (vx', vy') is represented as (vx×F, vy×F). In some embodiments, the top left position of the sub-block is represented as (x, y), and the sub-block has a size of L×K, where L is the width of the current sub-block, and K is the height of the sub-block. The reference block is identified as the region covering (x+offsetX+vx', y+offsetY+vy'), where offsetX and offset are non-negative values. In some embodiments, offsetX is 0 and / or offsetY is 0. In some embodiments, offsetX is equal to L / 2, L / 2+1, or L / 2-1. In some embodiments, offsetY is equal to K / 2, K / 2+1, or K / 2-1. In some embodiments, offsetX and / or offsetY are cropped within a range that includes a reference area for copying images, strips, slices, tiles, or intra-frame blocks.
[1511] In some embodiments, the motion vector of the sub-block is further determined based on the motion information of the reference block. In some embodiments, when the motion vector of the reference block points to the current image, the motion vector of the sub-block is the same as the motion vector of the reference block. In some embodiments, when the motion vector of the reference block points to the current image, the motion vector of the sub-block is determined based on the motion vector of the reference block to which the initial motion vector is added. In some embodiments, the motion vector of the sub-block is clipped within a range such that the motion vector of the sub-block points to the intra-block copy reference region. In some embodiments, the motion vector of the sub-block is a valid motion vector of the sub-block's intra-block copy candidate.
[1512] In some embodiments, one or more intra-block copy candidates for a sub-block are determined to determine the motion information of the sub-block. In some embodiments, one or more intra-block copy candidates are added to a motion candidate list, wherein the motion candidate list includes one of the following: a merge candidate for the sub-block, a sub-block temporal motion vector prediction candidate for the sub-block, or an affine merge candidate for the sub-block. In some embodiments, one or more intra-block copy candidates precede any merge candidate for the sub-block in the list. In some embodiments, one or more intra-block copy candidates follow any sub-block temporal motion vector prediction candidate for the sub-block in the list. In some embodiments, one or more intra-block copy candidates follow any inherited or constructed affine candidate for the sub-block in the list. In some embodiments, whether one or more intra-block copy candidates are added to the motion candidate list is based on the encoding / decoding mode of the current block. In some embodiments, when the current block is encoded / decoded using the intra-block copy (IBC) sub-block temporal motion vector prediction mode, one or more intra-block copy candidates are excluded from the motion candidate list.
[1513] In some embodiments, whether one or more intra-block copy candidates are added to the motion candidate list is based on the segmentation structure of the current block. In some embodiments, one or more intra-block copy candidates are added as merge candidates for sub-blocks in the list. In some embodiments, one or more intra-block copy candidates are added to the motion candidate list based on different initial motion vectors. In some embodiments, whether one or more intra-block copy candidates are added to the motion candidate list is indicated in the bitstream representation. In some embodiments, whether the index of the motion candidate list is signaled in the bitstream representation is based on the encoding / decoding mode of the current block. In some embodiments, when encoding / decoding the current block using the intra-block copy merge mode, the index of the motion candidate list including intra-block copy merge candidates is signaled in the bitstream representation. In some embodiments, when encoding / decoding the current block using the intra-block copy sub-block temporal motion vector prediction mode, the index of the motion candidate list including intra-block copy sub-block temporal motion vector prediction candidates is signaled in the bitstream representation. In some embodiments, the motion vector difference of the intra-block copy sub-block temporal motion vector prediction mode is applied to multiple sub-blocks.
[1514] In some embodiments, reference blocks and sub-blocks are associated with the same color components of the video. In some embodiments, whether the current block is divided into multiple sub-blocks is based on codec characteristics associated with the current block. In some embodiments, codec characteristics include decoder parameter sets, sequence parameter sets, video parameter sets, picture parameter sets, APS, picture headers, strip headers, slice headers, maximum codec unit (LCU), codec unit (CU), LCU rows, LCU groups, transform units, prediction units, prediction unit blocks, or syntax flags in the bitstream representation of video codec units. In some embodiments, codec characteristics include the position of codec units, prediction units, transform units, blocks, or video codec units. In some embodiments, codec characteristics include the dimension of the current block or neighboring blocks of the current block. In some embodiments, codec characteristics include the shape of the current block or neighboring blocks of the current block. In some embodiments, codec characteristics include the intra-frame codec mode of the current block or neighboring blocks of the current block. In some embodiments, codec characteristics include the motion vectors of neighboring blocks of the current block. In some embodiments, codec characteristics include an indication of the color format of the video. In some embodiments, codec characteristics include the codec tree structure of the current block. In some embodiments, the codec features include a stripe type, slice type, or picture type associated with the current block. In some embodiments, the codec features include a color component associated with the current block. In some embodiments, the codec features include a temporal layer identifier associated with the current block. In some embodiments, the codec features include a summary, level, or hierarchy of the bitstream representation standard.
[1515] Figure 26 This is a flowchart representation of a method 2600 for video processing according to the present technology. Method 2600 includes, in operation 2610, determining that the current block is divided into multiple sub-blocks for a conversion between a current block of video and a bitstream representation of the video. Depending on the mode, each of the multiple sub-blocks is encoded or decoded in a codec representation using a corresponding codec technique. The method further includes, in operation 2620, performing a conversion based on this determination.
[1516] In some embodiments, the mode specifies encoding and decoding a first sub-block of a plurality of sub-blocks using a modified intra-block copy (IBC) codec technique in which reference samples from video regions are used. In some embodiments, the mode specifies encoding and decoding a second sub-block of a plurality of sub-blocks using an intra-prediction codec technique in which samples from the same sub-block are used. In some embodiments, the mode specifies encoding and decoding a second sub-block of a plurality of sub-blocks using a palette codec technique in which a palette of representative pixel values is used. In some embodiments, the mode specifies encoding and decoding a second sub-block of a plurality of sub-blocks using an inter-frame codec technique in which temporal information is used.
[1517] In some embodiments, the mode specifies that a first sub-block among a plurality of sub-blocks is encoded and decoded using an intra-prediction coding / decoding technique in which samples from the same sub-block are used. In some embodiments, the mode specifies that a second sub-block among a plurality of sub-blocks is encoded and decoded using a palette coding / decoding technique in which a palette of representative pixel values is used. In some embodiments, the mode specifies that a second sub-block among a plurality of sub-blocks is encoded and decoded using an inter-frame coding / decoding technique in which temporal information is used.
[1518] In some embodiments, the mode specifies the use of a single encoding / decoding technique to encode and decode all sub-blocks across a plurality of sub-blocks. In some embodiments, the single encoding / decoding technique includes an intra-prediction encoding / decoding technique that uses samples from the same sub-block to encode and decode the plurality of sub-blocks. In some embodiments, the single encoding / decoding technique includes a palette encoding / decoding technique that uses a palette of representative pixel values to encode and decode the plurality of sub-blocks.
[1519] In some embodiments, where the mode of one or more encoding / decoding techniques is applicable to the current block, the history-based table of motion candidates for the sub-block temporal motion vector prediction mode remains the same for the transitions. The history-based table of motion candidates is determined based on motion information from past transitions. In some embodiments, the history-based table is used for IBC encoding / decoding techniques or non-IBC encoding / decoding techniques.
[1520] In some embodiments, when the mode specifies that at least one sub-block among a plurality of sub-blocks is encoded and decoded using IBC encoding / decoding technology, one or more motion vectors of at least one sub-block are used to update a history-based table of motion candidates for the IBC sub-block temporal motion vector prediction mode. The history-based table of motion candidates is determined based on motion information from past transitions. In some embodiments, when the mode specifies that at least one sub-block among a plurality of sub-blocks is encoded and decoded using inter-frame encoding / decoding technology, one or more motion vectors of at least one sub-block are used to update a history-based table of motion candidates for the non-IBC sub-block temporal motion vector prediction mode. The history-based table of motion candidates is determined based on motion information from past transitions.
[1521] In some embodiments, the use of a filtering process to filter the boundaries of multiple sub-blocks is based on the use of at least one encoding / decoding technique according to a mode. In some embodiments, when at least one encoding / decoding technique is applied, the filtering process to filter the boundaries of multiple sub-blocks is applied. In some embodiments, when at least one encoding / decoding technique is applied, the filtering process to filter the boundaries of multiple sub-blocks is omitted.
[1522] In some embodiments, depending on the mode, the second coding / decoding technique is disabled for the current block for this transformation. In some embodiments, the second coding / decoding technique includes at least one of the following: sub-block transform coding / decoding technique, affine motion prediction coding / decoding technique, multi-reference line intra-prediction coding / decoding technique, matrix-based intra-prediction coding / decoding technique, symmetric motion vector difference (MVD) coding / decoding technique, Merge using MVD decoder-side motion derivation or refinement coding / decoding technique, bidirectional optical flow coding / decoding technique, quadratic transform coding / decoding technique with dimensionality reduction based on the current block dimension, or multi-transform set coding / decoding technique.
[1523] In some embodiments, the use of at least one codec technique according to the mode is signaled in the bitstream representation. In some embodiments, this use is signaled at the sequence level, picture level, stripe level, slice group level, slice level, tile level, codec tree unit (CTU) level, codec tree block (CTB) level, codec unit (CU) level, prediction unit (PU) level, transform unit (TU) level, or another video unit level. In some embodiments, at least one codec technique includes a modified IBC codec technique, and the modified IBC codec technique is indicated in the bitstream representation based on the index value of a candidate in the motion candidate list. In some embodiments, a predefined value is assigned to the current block encoded using the modified IBC codec technique.
[1524] In some embodiments, the use of at least one encoding / decoding technique according to the mode is determined during conversion. In some embodiments, the use of an intra-block copy (IBC) encoding / decoding technique for encoding / decoding the current block, which is used for reference samples from the current block, is signaled in the bitstream. In some embodiments, the use of an intra-block copy (IBC) encoding / decoding technique for encoding / decoding the current block, which is used for reference samples from the current block, is determined during conversion.
[1525] In some embodiments, motion information from multiple sub-blocks of the current block is used as motion vector prediction values for the transformation between subsequent blocks of the video and the bitstream representation. In some embodiments, motion information from multiple sub-blocks of the current block is not permitted for the transformation between subsequent blocks of the video and the bitstream representation. In some embodiments, determining that the current block is divided into multiple sub-blocks for encoding and decoding using at least one encoding and decoding technique is based on whether motion candidates are candidates for blocks or sub-blocks within a block suitable for the video.
[1526] Figure 27This is a flowchart representation of a method 2700 for video processing according to the present technology. Method 2700 includes, in operation 2710, determining an operation associated with a motion candidate list based on conditions related to the characteristics of the current block for a conversion between a current block of video and a bitstream representation of the video. The motion candidate list is constructed for encoding / decoding techniques or based on information from previously processed blocks of video. Method 2700 also includes, in operation 2720, performing the conversion based on this determination.
[1527] In some embodiments, the encoding and decoding techniques include Merge encoding and decoding techniques, intra-block copy (IBC) sub-block temporal motion vector prediction encoding and decoding techniques, sub-block Merge encoding and decoding techniques, IBC encoding and decoding techniques, or IBC encoding and decoding techniques that use reference samples from video regions of the current block for modifications to encode and decode at least one sub-block of the current block.
[1528] In some embodiments, the current block has a dimension of W×H, where W and H are positive integers. The condition relates to the dimension of the current block. In some embodiments, the condition relates to the encoding / decoding information of the current block or the encoding / decoding information of neighboring blocks. In some embodiments, the condition relates to a Merge sharing condition used to share a motion candidate list between the current block and another block.
[1529] In some embodiments, the operation includes deriving spatial merge candidates for the motion candidate list using merge encoding / decoding techniques. In some embodiments, the operation includes deriving motion candidates for the motion candidate list based on spatially neighboring blocks of the current block. In some embodiments, spatially neighboring blocks include adjacent blocks or non-adjacent blocks of the current block.
[1530] In some embodiments, the operation includes deriving motion candidates from a motion candidate list constructed based on information from previously processed blocks of video. In some embodiments, the operation includes deriving pairwise merge candidates from the motion candidate list. In some embodiments, the operation includes one or more pruning operations to remove redundant entries from the motion candidate list. In some embodiments, one or more pruning operations are used for spatial merge candidates in the motion candidate list.
[1531] In some embodiments, the operation includes updating a motion candidate list constructed based on information from previously processed blocks of the video after the transformation. In some embodiments, the update includes adding derived candidates to the motion candidate list without removing redundant trimming operations from the motion candidate list. In some embodiments, the operation includes adding default motion candidates to the motion candidate list. In some embodiments, the default motion candidates include zero motion candidates using IBC sub-block temporal motion vector prediction encoding and decoding techniques. In some embodiments, the operation is skipped if certain conditions are met.
[1532] In some embodiments, the operation includes checking motion candidates in a motion candidate list in a predefined order. In some embodiments, the operation includes checking a predefined number of motion candidates in the motion candidate list. In some embodiments, a condition is met if W×H is greater than or equal to a threshold. In some embodiments, a condition is met if W×H is greater than or equal to a threshold and the current block is encoded / decoded using IBC sub-block temporal motion vector prediction encoding / decoding technology or Merge encoding / decoding technology. In some embodiments, the threshold is 1024.
[1533] In some embodiments, the condition is met when W and / or H are greater than or equal to a threshold. In some embodiments, the threshold is 32. In some embodiments, the condition is met when W × H is less than or equal to a threshold and the current block is encoded using IBC sub-block temporal motion vector prediction encoding / decoding technology or Merge encoding / decoding technology. In some embodiments, the threshold is 16. In some embodiments, the threshold is 32 or 64. In some embodiments, when the condition is met, the operation of inserting candidate motion candidates determined based on spatially neighboring blocks is skipped.
[1534] In some embodiments, the condition is satisfied when W equals T2, H equals T3, and the neighboring block above the current block is available and encoded using the same encoding / decoding technique as the current block, where T2 and T3 are positive integers. In some embodiments, the condition is satisfied when the neighboring block and the current block are in the same encoding / decoding tree unit.
[1535] In some embodiments, when W equals T2, H equals T3, and the neighboring block above the current block is unavailable or outside the current codec tree unit where the current block resides, the condition that T2 and T3 are positive integers is satisfied. In some embodiments, T2 is 4 and T3 is 8. In some embodiments, when W equals T4, H equals T5, and the neighboring block to the left of the current block is available and encoded using the same codec technique as the current block, the condition that T4 and T5 are positive integers is satisfied. In some embodiments, when W equals T4, H equals T5, and the neighboring block to the left of the current block is unavailable, the condition that T4 and T5 are positive integers is satisfied. In some embodiments, T4 is 8 and T5 is 4.
[1536] In some embodiments, the condition is met when W×H is less than or equal to a threshold, the current block is encoded and decoded using IBC sub-block temporal motion vector prediction coding and decoding or Merge coding and decoding, and the first neighboring block above the current block and the second neighboring block to the left of the current block are encoded and decoded using the same coding and decoding technique. In some embodiments, the first neighboring block and the second neighboring block are available and encoded and decoded using IBC coding and decoding, wherein the second neighboring block is within the same coding and decoding tree unit as the current block. In some embodiments, the first neighboring block is not available, wherein the second neighboring block is available and is within the same coding and decoding tree unit as the current block. In some embodiments, the first neighboring block and the second neighboring block are not available. In some embodiments, the first neighboring block is available, and the second neighboring block is not available. In some embodiments, the first neighboring block is not available, and the second neighboring block is outside the coding and decoding tree unit where the current block is located. In some embodiments, the first neighboring block is available, and the second neighboring block is outside the coding and decoding tree unit where the current block is located. In some embodiments, the threshold is 32. In some embodiments, the first neighboring block and the second neighboring block are used to derive spatial domain Merge candidates. In some embodiments, the top left sample of the current block is located at (x, y), and the second neighboring block covers the sample located at (x-1, y+H-1). In some embodiments, the top left sample of the current block is located at (x, y), and the second neighboring block covers the sample located at (x+W-1, y-1).
[1537] In some embodiments, the same encoding / decoding technology includes IBC encoding / decoding technology. In some embodiments, the same encoding / decoding technology includes inter-frame encoding / decoding technology. In some embodiments, the neighboring blocks of the current block have a dimension equal to A×B. In some embodiments, the neighboring blocks of the current block have a dimension greater than A×B. In some embodiments, the neighboring blocks of the current block have a dimension less than A×B. In some embodiments, A×B equals 4×4. In some embodiments, the threshold is predefined. In some embodiments, the threshold is signaled in the bitstream representation. In some embodiments, the threshold is based on the encoding / decoding characteristics of the current block, including the encoding / decoding mode in which the current block is encoded / decoded.
[1538] In some embodiments, the condition is satisfied when the current block has a parent node that shares a list of motion candidates and is encoded / decoded using IBC sub-block temporal motion vector prediction encoding / decoding or Merge encoding / decoding. In some embodiments, the condition is adaptively changed based on the encoding / decoding characteristics of the current block.
[1539] Figure 28This is a flowchart representation of a method 2800 for video processing according to the present technology. Method 2800 includes, in operation 2810, determining that the current block, encoded using a temporally information-based inter-frame encoding / decoding technique, is divided into multiple sub-blocks for a conversion between the current block of video and a bitstream representation of the video. At least one of the multiple blocks is encoded using a modified intra-block copy (IBC) encoding / decoding technique, wherein the technique uses reference samples from one or more video regions including the current block. Method 2800 includes, in operation 2820, performing a conversion based on this determination.
[1540] In some embodiments, the video region includes the current picture, strip, slice, tile, or group of slices. In some embodiments, the inter-frame coding technique includes sub-block temporal motion vector coding, and one or more syntax elements indicating whether the current block is encoded based on the current picture and a reference picture different from the current picture are included in the bitstream representation. In some embodiments, when the current block is encoded based on the current picture and a reference picture, one or more syntax elements indicate the reference picture used for encoding the current block. In some embodiments, one or more syntax elements also indicate motion information associated with the reference picture, the motion information including at least a motion vector prediction index, motion vector difference, or motion vector precision. In some embodiments, the first reference picture list includes only the current picture, and the second reference picture list includes only reference pictures. In some embodiments, the inter-frame coding technique includes temporal merge coding, and the motion information is determined based on the neighboring blocks of the current block, the motion information including at least motion vectors or reference pictures. In some embodiments, when neighboring blocks are determined only based on the current picture, the motion information applies only to the current picture. In some embodiments, when neighboring blocks are determined based on the current picture and a reference picture, the motion information applies to both the current picture and the reference picture. In some embodiments, motion information is applied to the current image only if neighboring blocks are determined based on the current image and a reference image. In some embodiments, if neighboring blocks are determined only based on a reference image, neighboring blocks are discarded when determining merge candidates.
[1541] In some embodiments, a fixed weighting factor is assigned to a reference block from the current image and a reference block from a reference image. In some embodiments, the fixed weighting factor is signaled in the bitstream representation.
[1542] In some embodiments, performing the conversion includes generating a bitstream representation based on video blocks. In some embodiments, performing the conversion includes generating video blocks from the bitstream representation.
[1543] It should be understood that techniques for video encoding or video decoding are disclosed. Video encoders or decoders can employ these techniques together with intra-frame block copying and sub-block-based video processing to achieve greater encoding / decoding efficiency and performance.
[1544] Some embodiments of the disclosed technology involve making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement that tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on that decision or determination, the conversion from video blocks to a bitstream representation of the video will use the video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will process the bitstream knowing that it has been modified based on the video processing tool or mode. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on that decision or determination.
[1545] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder will not use the tool or mode in the conversion from video block to video bitstream representation. In another example, when a video processing tool or mode is disabled, the decoder will process the bitstream knowing that it has not been modified using the video processing tool or mode enabled based on that decision or determination.
[1546] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this document and their equivalents), or in a combination of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, such as one or more modules of computer program instructions encoded on a computer-readable medium for performing or controlling the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a combination of substances affecting machine-readable propagation signals, or a combination of one or more of them. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, generated to encode information for transmission to a suitable receiver device.
[1547] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., a file storing one or more modules, subroutines, or code portions). Computer programs can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected through a communications network.
[1548] The processes and logic described in this document can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processes and logic can also be executed by dedicated logic circuits, and the devices can be implemented as dedicated logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).
[1549] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or operatively coupled to receive data from, transfer data to, or receive data from and transfer data to such mass storage devices. However, a computer does not require such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[1550] While this patent document contains numerous details, these details should not be construed as limiting any subject matter or potentially claimed scope, but rather as descriptions of features specific to particular embodiments of a particular technology. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be excluded from the combination, and the claimed combination may be for sub-combinations or variations thereof.
[1551] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential manner, or as performing all shown operations to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.
[1552] Only some implementation methods and examples are described, and other implementation methods, enhancements and variations can be made based on the content described and shown in this patent document.
Claims
1. A video processing method, comprising: For the conversion between the current block of the video and the bitstream of the video, it is determined that the current block is divided into multiple sub-blocks, wherein each of the multiple sub-blocks is encoded and decoded in the bitstream using a corresponding encoding and decoding technology according to the mode; as well as The conversion is performed based on the determination; The first sub-block among the plurality of sub-blocks is encoded and decoded using sub-block-based IBC (sbIBC) encoding and decoding technology. The second sub-block among the plurality of sub-blocks is encoded and decoded using one of the following technologies: intra-frame prediction encoding and decoding, palette encoding and decoding, and inter-frame encoding and decoding. In the sub-block-based IBC (sbIBC) encoding and decoding technology, a corresponding reference block is identified based on the initial motion vector of the sub-block, and the motion parameters of the identified reference block are used to determine the motion parameters of the corresponding sub-block. The sub-block-based IBC (sbIBC) encoding and decoding technology uses reference samples from video regions. Specifically, motion information of the first sub-block of the current block is not permitted for use in the conversion between subsequent blocks of the video and the bitstream.
2. The method according to claim 1, wherein, When one or more encoding / decoding techniques are applicable to the current block, the history-based table of motion candidates for the sub-block temporal motion vector prediction mode remains the same for the transitions, and the history-based table of motion candidates is determined based on motion information from past transitions.
3. The method according to claim 2, wherein, The history-based table is used for IBC encoding / decoding technology or non-IBC encoding / decoding technology.
4. The method according to any one of claims 1 to 3, wherein, When the mode specifies that at least one of the plurality of sub-blocks is encoded and decoded using IBC encoding and decoding technology, one or more motion vectors of at least one sub-block are used to update the history-based table of motion candidates for the IBC sub-block temporal motion vector prediction mode, the history-based table of motion candidates being determined based on motion information from past transformations.
5. The method according to any one of claims 1 to 3, wherein, When the mode specifies the use of inter-frame encoding / decoding techniques to encode or decode at least one of the plurality of sub-blocks, one or more motion vectors from at least one sub-block are used to update a history-based table of motion candidates for the non-IBC sub-block temporal motion vector prediction mode, the history-based table of motion candidates being determined based on motion information from past transitions.
6. The method according to claim 1, wherein, When at least one encoding / decoding technique is used, a filtering process is applied to filter the boundaries of the plurality of sub-blocks.
7. The method according to claim 1, wherein, When at least one encoding / decoding technique is applied, the filtering process for filtering the boundaries of the plurality of sub-blocks is omitted.
8. The method according to any one of claims 1 to 3, wherein, According to the mode, the second encoding / decoding technology is disabled for the current block for the transformation.
9. The method according to claim 8, wherein, The second encoding / decoding technology includes at least one of the following: sub-block transform encoding / decoding technology, affine motion prediction encoding / decoding technology, multi-reference line intra-prediction encoding / decoding technology, matrix-based intra-prediction encoding / decoding technology, symmetric motion vector difference (MVD) encoding / decoding technology, Merge using MVD decoder-side motion derivation or refinement encoding / decoding technology, bidirectional optical flow encoding / decoding technology, quadratic transform encoding / decoding technology with dimension reduction based on the current block dimension, or multi-transform set encoding / decoding technology.
10. The method according to any one of claims 1 to 3, wherein, The use of at least one encoding / decoding technique according to the mode is signaled in the bitstream.
11. The method according to claim 10, wherein, The signaling is used at the sequence level, picture level, strip level, slice group level, slice level, tile level, codec tree unit (CTU) level, codec tree block (CTB) level, codec unit (CU) level, prediction unit (PU) level, transform unit (TU) level, or at another video unit level.
12. The method according to claim 10, wherein, The at least one encoding / decoding technique includes a modified IBC encoding / decoding technique, and the modified IBC encoding / decoding technique is indicated in the bitstream based on the index value of a candidate in the indicated motion candidate list.
13. The method according to claim 12, wherein, Predefined values are assigned to the current block that is encoded or decoded using the modified IBC encoding / decoding technology.
14. The method according to any one of claims 1 to 3, wherein, The use of at least one encoding / decoding technique according to the mode is determined during conversion.
15. The method according to any one of claims 1 to 3, wherein, The use of the intra-block copy (IBC) encoding / decoding technique used to encode and decode the current block, as indicated by the reference samples from the current block, is signaled in the bitstream.
16. The method according to any one of claims 1 to 3, wherein, The use of the intra-block copy (IBC) encoding / decoding technique for encoding and decoding the current block, which is used by the reference samples from the current block, is determined during the conversion.
17. The method according to any one of claims 1 to 3, wherein, The motion information of the multiple sub-blocks of the current block is used as motion vector prediction values for the conversion between subsequent blocks of the video and the bitstream.
18. The method according to any one of claims 1 to 3, wherein, The determination that the current block is divided into the plurality of sub-blocks encoded and decoded using at least one encoding and decoding technique is based on whether the motion candidate is a candidate for a block or a sub-block within a block that is suitable for video.
19. The method according to any one of claims 1 to 3, wherein, Performing the conversion includes generating the bitstream from the current block of the video.
20. The method according to any one of claims 1 to 3, wherein, Performing the conversion includes generating the current block of the video from the bitstream.
21. A video processing apparatus, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: For the conversion between the current block of the video and the bitstream of the video, it is determined that the current block is divided into multiple sub-blocks, wherein each of the multiple sub-blocks is encoded and decoded in the bitstream using a corresponding encoding and decoding technology according to the mode; as well as The conversion is performed based on the determination; The first sub-block among the plurality of sub-blocks is encoded and decoded using sub-block-based IBC (sbIBC) encoding and decoding technology. The second sub-block among the plurality of sub-blocks is encoded and decoded using one of the following technologies: intra-frame prediction encoding and decoding, palette encoding and decoding, and inter-frame encoding and decoding. In the sub-block-based IBC (sbIBC) encoding and decoding technology, a corresponding reference block is identified based on the initial motion vector of the sub-block, and the motion parameters of the identified reference block are used to determine the motion parameters of the corresponding sub-block. The sub-block-based IBC (sbIBC) encoding and decoding technology uses reference samples from video regions. Specifically, motion information of the first sub-block of the current block is not permitted for use in the conversion between subsequent blocks of the video and the bitstream.
22. A computer-readable medium having code stored thereon, said code causing a processor, when executed, to: The current block of the video is divided into multiple sub-blocks for the conversion between the current block and the bitstream of the video, wherein each sub-block is encoded and decoded in the bitstream using a corresponding encoding / decoding technique according to a mode; and The conversion is performed based on the determination; in, The first sub-block of the plurality of sub-blocks is encoded and decoded using sub-block-based IBC (sbIBC) encoding and decoding technology. The second sub-block of the plurality of sub-blocks is encoded and decoded using one of intra-frame prediction encoding and decoding technology, palette encoding and decoding technology, and inter-frame encoding and decoding technology. In the sub-block-based IBC (sbIBC) encoding and decoding technology, a corresponding reference block is identified based on the initial motion vector of the sub-block, and the motion parameters of the identified reference block are used to determine the motion parameters of the corresponding sub-block. The sub-block-based IBC (sbIBC) encoding and decoding technology uses reference samples from video regions. Specifically, motion information of the first sub-block of the current block is not permitted for use in the conversion between subsequent blocks of the video and the bitstream.
23. A method for storing a video bitstream, comprising: For the current block of the video, it is determined that the current block is divided into multiple sub-blocks, wherein each of the multiple sub-blocks is encoded and decoded in the bitstream using the corresponding encoding and decoding technology according to the mode; The bit stream is generated based on the determination; as well as A computer program / instructions and the bit stream are stored in a non-transitory computer-readable storage medium, wherein the computer program / instructions, when executed by a processor, implement a video processing method to generate the bit stream; The first sub-block among the plurality of sub-blocks is encoded and decoded using sub-block-based IBC (sbIBC) encoding and decoding technology. The second sub-block among the plurality of sub-blocks is encoded and decoded using one of the following technologies: intra-frame prediction encoding and decoding, palette encoding and decoding, and inter-frame encoding and decoding. In the sub-block-based IBC (sbIBC) encoding and decoding technology, a corresponding reference block is identified based on the initial motion vector of the sub-block, and the motion parameters of the identified reference block are used to determine the motion parameters of the corresponding sub-block. The sub-block-based IBC (sbIBC) encoding and decoding technology uses reference samples from video regions. In this context, motion information from the first sub-block of the current block is not permitted to be used to generate the bitstream of subsequent blocks of the video.