Recursive partitioning of video coding blocks

By recursively dividing and transforming video blocks using an intra-frame block copying tool and a low-frequency indivisible transform model, the problem of low video encoding and decoding efficiency in existing technologies is solved, achieving more efficient video compression and bandwidth utilization, and is applicable to HEVC and future video encoding and decoding standards.

CN119743597BActive Publication Date: 2026-04-28DOUYIN VISION CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2020-09-09
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies suffer from low encoding and decoding efficiency and high bandwidth usage when processing video blocks, especially in the Internet and digital communication networks, where the bandwidth demand for video continues to grow as the number of connected user devices increases.

Method used

By employing an intra-block copying tool and a low-frequency inseparable transform encoding/decoding model, the encoding/decoding process is optimized through recursive partitioning and transformation of video blocks. This includes the use of the intra-block copying model, the application of low-frequency inseparable transform, the segmentation of motion vectors, and the determination of encoding/decoding conditions, thereby improving encoding/decoding efficiency.

Benefits of technology

It improves the compression performance of video encoding and decoding, reduces bandwidth requirements, and increases encoding and decoding efficiency. It is applicable to existing HEVC and future video encoding and decoding standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743597B_ABST
    Figure CN119743597B_ABST
Patent Text Reader

Abstract

A method of video processing includes, for a conversion between a current block in a video region of a video and a coded representation of the video, determining that a partitioning of the current block into a plurality of transform units is allowed based on a characteristic of the current block. Signaling of the partitioning is omitted in the coded representation. The method further includes performing the conversion based on the determining.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is a divisional application of the invention patent application filed on September 9, 2020, with application number 202080063234.4 and invention title "Recursive Partition of Video Codec Blocks". Technical Field

[0003] This patent document relates to the encoding and decoding of video and images. Background Technology

[0004] Despite advancements in video compression, digital video still accounts for the largest share of bandwidth usage on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth required for digital video usage is expected to continue to grow. Summary of the Invention

[0005] Apparatus, systems and methods relating to digital video encoding and decoding, particularly to the encoding and decoding of video and images, wherein intra-frame block copying tools are used for encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes: for a transformation between a current block in a video region of a video and a codec representation of the video, determining whether the current block can be divided into multiple transform units based on characteristics of the current block. Signaling notification of the division is omitted in the codec representation. The method also includes performing the transformation based on the determination.

[0007] In another example, a video processing method is disclosed. The method includes: for a conversion between a current block in a video region of a video and the codec representation of the video, determining predicted samples of the current block using an intra-block copy (IBC) model based on filtered reconstructed reference samples in the video region. The method also includes performing the conversion based on the determination.

[0008] In another example, a video processing method is disclosed. The method includes: for a conversion between a current block of a video and its codec representation, determining, according to rules, whether a syntax element indicating a skip mode using an Intra-Block Copy (IBC) codec model is included in the codec representation. The rules specify that signaling notification for the syntax element is based on the dimension of the current block and / or on the maximum allowed dimension of a block encoded using the IBC codec model. The method also includes performing the conversion based on the determination.

[0009] In another example aspect, a video processing method is disclosed. The method includes: for the conversion between a current block of a video and the codec representation of the video, determining at least one context for encoding and decoding an index associated with a low-frequency non-separable transform (LFNST) codec model. The LFNST codec model includes: during encoding, applying a positive quadratic transform between the positive main transform and the quantization step; or during decoding, applying an inverse quadratic transform between the dequantization step and the inverse main transform. The size of the positive quadratic transform and the inverse quadratic transform is smaller than the size of the current block. Determining at least one context based on the segmentation type of the current block, without considering the positive main transform or the inverse main transform. The method further includes performing the conversion according to the determination.

[0010] In another example aspect, a video processing method is disclosed. The method includes: for the conversion between a current block of a video region of a video and the codec representation of the video, determining whether to enable an intra block copy (IBC) codec model based on the maximum transform unit size applicable to the video region. The method further includes performing the conversion according to the determination.

[0011] In another example aspect, a video processing method is disclosed. The method includes: for the conversion between a current block of a video and the codec representation of the video, determining that the absolute value of a component of the motion vector of the current block is divided into two parts, where the motion vector is represented as (Vx, Vy), the component is represented as Vi, and Vi is Vx or Vy. The first part of the two parts is equal to |Vi| - ((|Vi| >> N) << N), the second part of the two parts is equal to |Vi| >> N, and N is a positive integer. The two parts are encoded and decoded separately in the codec representation. The method further includes performing the conversion according to the determination.

[0012] In another example aspect, a video processing method is disclosed. The method includes: for the conversion between a current block of a video and the codec representation of the video, determining information about the maximum dimension of the current block that allows sub-block transformation in the current block based on the maximum allowed dimension of the transform block. The method further includes performing the conversion according to the determination.

[0013] In another example aspect, a video processing method is disclosed. The method includes: for the conversion between the codec representation of a video block in a video region and the video block based on the intra block copy tool, determining that the video block is allowed to be divided into multiple transform units, where the determination is based on the codec conditions of the video block, and the codec representation omits signaling of the division; and performing the conversion based on the division.

[0014] In another example, a video processing method is disclosed. The method includes: converting between the codec representation of a video block and the video block for a video region; determining whether an intra-block copying (IBC) tool is enabled for the conversion of the video block based on the maximum transform size of the video region; and performing the conversion based on the determination.

[0015] In another example, a video processing method is disclosed. The method includes: for a conversion between a codec representation of a video block of a video region and the video block itself, determining whether signaling notification from an intra-block copy (IBC) tool used for the conversion is included in the codec representation; and performing the conversion based on the determination, wherein the determination is based on the width and / or height of the video block and the maximum permissible IBC block size of the video region.

[0016] In another example, a video processing method is disclosed. The method includes: determining, for a conversion between a codec representation of a video block in a video region and the video block itself, performed using an intra-block copy (IBC) tool, that the video block can be divided into multiple transform units (TUs) for conversion; and performing the conversion based on the determination, wherein the conversion includes using individual motion information on the multiple TUs.

[0017] In another example, a video processing method is disclosed. The method includes: determining, for a conversion between a codec representation of a video block of a video region and the video block itself, that an intra-block copying tool is enabled for the conversion; and performing the conversion using the intra-block copying tool, wherein video block prediction is performed using filtered reconstructed samples of the video region.

[0018] In another example, a video processing method is disclosed. The method includes: performing a conversion between a video comprising multiple video blocks and a video codec representation, wherein at least some of the video blocks are encoded using motion vector information, and wherein the motion vector information is represented in the codec representation as a first part based on a first least significant bit of the absolute value of the motion vector information, and the motion vector information is represented in the codec representation as a second part based on the remaining more significant bits, which are more effective than the first least significant bit.

[0019] In another example, a video processing method is disclosed. The method includes: for a conversion between a codec representation of a video block of a video region and the video block, determining whether a sub-block transformation tool is enabled for the conversion; and performing the conversion based on the determination, wherein the determination is based on the maximum permissible transform block size of the video region, and wherein, based on the maximum permissible transform block size, signaling notifications are conditionally included in the codec representation.

[0020] In another example, a video processing method is disclosed. The method includes: for a conversion between a codec representation of a video block of a video region and the video block, determining whether a Low Frequency Inseparable Transform (LFNST) is used during the conversion; and performing the conversion based on the determination, wherein the determination is based on codec conditions applied to the video block, and wherein two contexts are used to encode and decode the LFNST and the matrix index of the LFNST in the codec representation.

[0021] In another representative aspect, the above method is implemented as processor-executable code and stored in a computer-readable program medium.

[0022] In another representative aspect, an apparatus configured or operable to perform the above-described methods is disclosed. This apparatus may include a processor programmed to implement the methods.

[0023] In another representative aspect, video decoder devices can implement the methods described herein.

[0024] The above and other aspects and features of the present disclosure are described in more detail in the accompanying drawings, specification and claims. Attached Figure Description

[0025] Figure 1 An example derivation of the Merge candidate list construction process is shown.

[0026] Figure 2 Example locations of spatial merge candidates are shown.

[0027] Figure 3 An example candidate pair is shown that considers redundancy checks for spatial merge candidates.

[0028] Figure 4 Example locations of the second PU are shown, divided into N×2N and 2N×N segments.

[0029] Figure 5 This is a diagram illustrating motion vector scaling for temporal Merge candidates.

[0030] Figure 6 Examples of candidate positions C0 and C1 for time-domain Merge candidates are shown.

[0031] Figure 7 An example of combined bidirectional prediction of Merge candidates is shown.

[0032] Figure 8 An example derivation of motion vector prediction candidates is shown.

[0033] Figure 9 This is a diagram illustrating motion vector scaling for spatial motion vector candidates.

[0034] Figure 10 An example of candidate positions for the affine Merge pattern is shown.

[0035] Figure 11 The modified Merge list construction process is shown.

[0036] Figure 12 An example of inter-frame prediction based on triangulation is shown.

[0037] Figure 13 An example of applying the first weighting factor group to CU is shown.

[0038] Figure 14 An example of motion vector storage is shown.

[0039] Figure 15 This is an example of UMVE search processing.

[0040] Figure 16 An example of a UMVE search point is shown.

[0041] Figure 17 This is a diagram illustrating the operation of the intra-frame block copying tool.

[0042] Figure 18 An example of a quadratic transformation in JEM is shown.

[0043] Figure 19 An example of the Reduced Secondary Transform (RST) is shown.

[0044] Figure 20 This is a block diagram of an example video processing system that can implement the disclosed technology.

[0045] Figure 21 This is a block diagram of an example video processing device.

[0046] Figure 22 This is a flowchart of an example method for video processing.

[0047] Figure 23 This is a flowchart representation of the video processing method based on this technology.

[0048] Figure 24 This is a flowchart representation of another video processing method based on this technology.

[0049] Figure 25 This is a flowchart representation of another video processing method based on this technology.

[0050] Figure 26 This is a flowchart representation of another video processing method based on this technology.

[0051] Figure 27 This is a flowchart representation of another video processing method based on this technology.

[0052] Figure 28 This is a flowchart representation of another video processing method based on this technology.

[0053] Figure 29 This is a flowchart representation of another video processing method based on this technology. Detailed Implementation

[0054] Embodiments of the disclosed techniques can be applied to existing video codec standards (e.g., HEVC, H.265) and future standards to improve compression performance. Section headings are used in this document to improve readability and do not in any way limit the discussion or embodiments (and / or implementations) to the relevant sections only.

[0055] 1. Overview

[0056] This application relates to video codec technology. Specifically, it relates to intra-block copying (IBC, also known as current picture reference (CPR)) codec. It can be applied to existing video codec standards such as HEVC, or to pending standards (general video codecs). It may also be applicable to future video codec standards or video codecs.

[0057] 2. Background

[0058] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) was established between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) to work on the VVC standard, which aims to reduce the bit rate by 50% compared to HEVC.

[0059] 2.1 Inter-frame prediction in HEVC / H.265

[0060] For inter-frame encoding / decoding, a codec unit (CU) can be encoded / decoded using one prediction unit (PU) or two PUs, depending on the segmentation mode. Each inter-frame prediction PU has motion parameters for one or two reference picture lists. The motion parameters include motion vectors and reference picture indices. The use of one of the two reference picture lists can also be notified via inter_pred_idc signaling. The motion vectors can be explicitly encoded / decoded as increments relative to the predictor.

[0061] When a CU uses skip mode encoding / decoding, a PU is associated with the CU and has no significant residual coefficients, no encoded motion vector increments, or a reference picture index. A Merge mode is specified, through which the motion parameters of the current PU can be obtained from neighboring PUs (including spatial and temporal candidates). The Merge mode can be applied to any PU for inter-frame prediction, not just skip mode. Another option for the Merge mode is explicit transmission of motion parameters, where the motion vectors (more precisely, the motion vector difference (MVD) compared to the motion vector predictor), the corresponding reference picture index for each reference picture list, and the use of the reference picture list are explicitly signaled according to each PU. In this disclosure, such a mode is named Advanced Motion Vector Prediction (AMVP).

[0062] When signaling indicates that one of two reference image lists should be used, a PU is generated from a block of samples. This is called "one-way prediction". One-way prediction is available for both P-slices and B-slices.

[0063] When signaling indicates that two lists of reference images should be used, a PU is generated from two blocks of samples. This is called "bidirectional prediction". Bidirectional prediction is only available for B-strips.

[0064] The following section provides details regarding inter-frame prediction modes as defined in HEVC. The description will begin with Merge mode.

[0065] 2.1.1 List of Reference Images

[0066] In HEVC, the term inter-frame prediction is used to refer to predictions derived from data elements (e.g., sample values ​​or motion vectors) of reference images other than the currently decoded image. As in H.264 / AVC, images can be predicted from multiple reference images. The reference images used for inter-frame prediction are organized into one or more reference image lists. A reference index identifies which reference images in the list are used to create the predicted signal.

[0067] A single list of reference images (list 0) is used for the P-strip, and two lists of reference images (list 0 and list 1) are used for the B-strip. It is important to note that the reference images contained in lists 0 and 1 can be from past and future images, in terms of capture / display order.

[0068] 2.1.2 Merge Mode

[0069] 2.1.2.1 Derivation of Merge Pattern Candidates

[0070] When using the Merge pattern to predict the PU, the indices pointing to entries in the Merge candidate list are parsed from the bitstream and used to retrieve motion information. The construction of this list is specified in the HEVC standard and can be summarized in the following steps:

[0071] Step 1: Initial Candidate Derivation

[0072] Step 1.1: Spatial Candidate Derivation

[0073] Step 1.2: Redundancy check of spatial candidates

[0074] Step 1.3: Temporal Candidate Derivation

[0075] Step 2: Adding candidate insertions

[0076] Step 2.1: Create bidirectional prediction candidates

[0077] Step 2.2: Insert zero-motion candidates

[0078] exist Figure 1 These steps are also schematically illustrated, showing an example derivation process for Merge candidate list reconstruction. For spatial Merge candidate derivation, at most four Merge candidates are selected from candidates located at five different positions. For temporal Merge candidate derivation, at most one Merge candidate is selected from two candidates. Since the number of candidates per PU is assumed to be constant at the decoder, additional candidates are generated when the number of candidates obtained from step 1 does not reach the maximum number of Merge candidates (maxNumMergeCand) signaled in the stripe header. Because the number of candidates is constant, the index of the best Merge candidate is encoded using truncated unigram binarization (TU). If the size of the CU is equal to 8, all PUs of the current CU share a single Merge candidate list, which is the same as the Merge candidate list of a 2N×2N prediction unit.

[0079] The operations associated with the foregoing steps are described in detail below.

[0080] 2.1.2.2 Derivation of Airspace Candidates

[0081] In the derivation of the spatial Merge candidate, located in Figure 2 At most four merge candidates are selected from the candidates at the indicated positions. The derivation order is A1, B1, B0, A0, and B2. Position B2 is considered only if any PU at positions A1, B1, B0, or A0 is unavailable (e.g., because it belongs to another stripe or slice) or if it is intra-frame encoding / decoding. After adding candidates at position A1, a redundancy check is performed on the remaining candidates, ensuring that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only those with... Figure 3 Only pairs of arrow links in the code are considered, and a candidate is added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information. Another source of duplicated motion information is the "second PU" associated with different 2Nx2N partitions. For example, Figure 4 The second prediction unit (PU) is described for N×2N and 2N×N cases respectively. When the current PU is segmented into N×2N, candidates for position A1 are not considered during list construction. In fact, adding this candidate could result in two prediction units with the same motion information, which is redundant for an encoding / decoding unit with only one PU. Similarly, when the current PU is segmented into 2N×N, position B1 is not considered.

[0082] 2.1.2.3 Time-domain candidate derivation

[0083] In this step, only one candidate is added to the list. Specifically, in the derivation of the current-domain Merge candidate, the scaling motion vector is derived based on the co-located PU in the co-located image. Figure 5 This is a diagram illustrating motion vector scaling for temporal merge candidates. (See diagram for example.) Figure 5 The dashed line in the middle shows the scaled motion vector for obtaining the temporal Merge candidate, which is scaled from the motion vector of the co-located PU using the POC distances tb and td, where tb is defined as the POC difference between the reference image of the current image and the current image, and td is defined as the POC difference between the reference image of the co-located image and the co-located image. The reference image index of the temporal Merge candidate is set to zero. The actual implementation of the scaling process is described in the HEVC specification [1]. For the B strip, two motion vectors are obtained (one for reference image list 0 and the other for reference image list 1) and combined to form a bidirectional prediction Merge candidate.

[0084] 2.1.2.4 Co-located images and co-located PUs

[0085] When TMVP is enabled (e.g., slice_temporal_mvp_enabled_flag equals 1), the derivation of the variable ColPic representing the juxtaposed images is as follows:

[0086] – If the current stripe is a B stripe and the collocated_from_l0_flag of the signaling notification is equal to 0, then set ColPic to be equal to RefPicList1[collocated_ref_idx].

[0087] Otherwise (slice_type equals B and collocated_from_l0_flag equals 1, or slice_type equals P), set ColPic to equal RefPicList0[collocated_ref_idx].

[0088] Here, collocated_ref_idx and collocated_from_l0_flag are two syntax elements that can be signaled in the stripe header.

[0089] In the co-occurrence PU(Y) of the reference frame, the position of the temporal candidate is selected between candidate C0 and C1, such as... Figure 6 As shown. If the PU at position C0 is unavailable, is intra-coded, or is outside the current codec tree unit (CTU, also known as LCU, maximum codec unit) row, then position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate.

[0090] The relevant syntax elements are described below:

[0091] 7.3.6.1 General Striped Segment Header Syntax

[0092]

[0093] 2.1.2.5 Derivation of the MV of TMVP Candidates

[0094] In some embodiments, the following steps are performed to derive TMVP candidates:

[0095] 1) Set the reference image list X = 0 and the target reference image to the reference image with index 0 in list X (e.g., curr_ref). Call the derivation process of the juxtaposed motion vector to obtain the MV of list X pointing to curr_ref.

[0096] 2) If the current stripe is stripe B, then set the reference image list X = 1 and the target reference image to the reference image with index 0 in list X (e.g., curr_ref). Call the derivation process of the juxtaposed motion vectors to obtain the MV of list X pointing to curr_ref.

[0097] The derivation of the juxtaposed motion vector is described in the next subsection 2.1.2.5.1.

[0098] 2.1.2.5.1 Derivation of the juxtaposed motion vectors

[0099] For co-occurring blocks, intra-frame or inter-frame encoding / decoding can be performed using unidirectional or bidirectional prediction. If intra-frame encoding / decoding is used, the TMVP candidate is set to unavailable.

[0100] If it is a one-way prediction from list A, then the motion vector of list A is scaled to the target reference image list X.

[0101] If it is a bidirectional prediction, and the list of target reference images is X, then the motion vectors of list A are scaled to the list of target reference images X, and A is determined according to the following rules:

[0102] – If no reference image has a larger POC value than the current image, then set A to equal X.

[0103] Otherwise, set A to equal collocated_from_l0_flag.

[0104] 2.1.2.6 Additional Candidate Insertion

[0105] In addition to spatial and temporal merge candidates, there are two additional types of merge candidates: combined bidirectional prediction merge candidates and zero merge candidates. Combined bidirectional prediction merge candidates are generated using spatial and temporal merge candidates. Combined bidirectional prediction merge candidates are only used for B-strips. They are generated by combining the motion parameters of the first reference image list of the initial candidate with the motion parameters of the second reference image list of another candidate. If these two tuples provide different motion hypotheses, they will form a new bidirectional prediction candidate. Figure 7 An example of combined bidirectional prediction of Merge candidates is shown. As an example, Figure 7 The following scenario illustrates a situation where two candidates from the original list (on the left) with mvL0 and refIdxL0 or mvL1 and refIdxL1 are used to create combined bidirectional prediction merge candidates added to the final list (on the right). There are many rules governing the combinations considered to be used to generate these additional merge candidates.

[0106] Zero-motion candidates are inserted to populate the remaining entries in the Merge candidate list, thus reaching the capacity of MaxNumMergeCand. These candidates have zero spatial displacement and a reference image index that starts from zero and increases each time a new zero-motion candidate is added to the list. Finally, no redundancy check is performed on these candidates.

[0107] 2.1.3 Advanced Motion Vector Prediction (AMVP)

[0108] AMVP utilizes the spatiotemporal correlation between motion vectors and neighboring PUs for explicit transmission of motion parameters. For each list of reference images, a motion vector candidate list is first constructed by checking the availability of temporally neighboring PU locations in the upper left corner, removing redundant candidates, and adding zero vectors to keep the candidate list length constant. The encoder can then select the best predictor from the candidate list and send the corresponding index indicating the selected candidate. Similar to Merge index signaling, the index of the best motion vector candidate is encoded using truncated unary values. In this case, the maximum value to be encoded is 2 (see [link to documentation]). Figure 8 The following sections provide details of the derivation process for motion vector prediction candidates.

[0109] 2.1.3.1 Derivation of AMVP Candidates

[0110] Figure 8 The derivation process of motion vector prediction candidates is summarized.

[0111] In motion vector prediction, two types of motion vector candidates are considered: spatial motion vector candidates and temporal motion vector candidates. The derivation of the spatial motion vector candidate is based on the fact that... Figure 2 The motion vectors of each PU at the five different locations shown ultimately yield two motion vector candidates.

[0112] For the derivation of temporal motion vector candidates, one motion vector candidate is selected from two candidates derived based on two different co-located positions. After creating the first space-time candidate list, duplicate motion vector candidates are removed from the list. If the number of potential candidates is greater than two, motion vector candidates with reference image indices greater than 1 in the associated reference image list are removed from the list. If the number of space-time motion vector candidates is less than two, additional zero motion vector candidates are added to the list.

[0113] 2.1.3.2 Candidate Spatial Motion Vectors

[0114] When deriving the spatial motion vector candidates, at most two candidates are considered from the five potential candidates. These five potential candidates are derived from the five vectors located in the region of ... Figure 2The positions of the PUs shown are the same as those of the motion merge. The derivation order to the left of the current PU is defined as A0, A1, and scaled A0, scaled A1. The derivation order above the current PU is defined as B0, B1, B2, scaled B0, scaled B1, scaled B2. Therefore, there are four cases on each side that can be used as motion vector candidates, two of which do not require spatial scaling, and two of which do. The four different cases are summarized as follows:

[0115] • No spatial scaling

[0116] –(1) Same list of reference images, and same index of reference images (same POC)

[0117] –(2) Different lists of reference images, but the same reference image (same POC)

[0118] • Spatial scaling

[0119] –(3) Same list of reference images, but different reference images (different POCs)

[0120] –(4) A list of different reference images, and different reference images (different POCs)

[0121] First, we check the case where spatial scaling is not allowed, then we check the case where spatial scaling is allowed. Spatial scaling is considered when the POC differs between the reference image of the neighboring PU and the reference image of the current PU, regardless of the list of reference images. If all PUs in the left-hand candidate pool are unavailable or are intra-frame encoded / decoded, scaling of the above motion vectors is allowed to aid in the parallel derivation of the MV candidates on the left and above. Otherwise, spatial scaling of the above motion vectors is not allowed.

[0122] During spatial scaling, the motion vectors of neighboring PUs are scaled in a manner similar to temporal scaling, such as... Figure 9 As shown. The main difference is that a list of reference images and their indices for the current PU are given as input; the actual scaling process is the same as the temporal scaling process.

[0123] 2.1.3.3 Candidate Motion Vectors in the Time Domain

[0124] Except for the derivation of the reference image index, all derivation processes for the temporal merge candidate are the same as those for the spatial motion vector candidate (see [link]). Figure 6 The reference image index signal is sent to the decoder.

[0125] 2.2 Inter-frame prediction methods in VVC

[0126] Several new codec tools for improving inter-frame prediction exist, such as Adaptive Motion Vector Differential Resolution (AMVR) for Signaling Notification MVD, Motion Vector Differential Merge (MMVD), Triangular Prediction Mode (TPM), Combined Intra-Inter-Frame Prediction (CIIP), Advanced TMVP (ATMVP, also known as SbTMVP), Affine Prediction Mode, General Bidirectional Prediction (GBI), Decoder-Side Motion Vector Refinement (DMVR), and Bidirectional Optical Flow (BIO, also known as BDOF).

[0127] VVC supports two different Merge list construction processes:

[0128] 1) Sub-block Merge Candidate List: This includes ATMVP and affine Merge candidates. Both affine and ATMVP modes share a single Merge list construction process. ATMVP and affine Merge candidates can be added sequentially. The sub-block Merge list size is signaled in the strip header, and its maximum value is 5.

[0129] 2) Regular Merge List: For inter-frame codec blocks, a shared Merge list construction process is used. Spatial / temporal Merge Candidates, HMVP, Paired Merge Candidates, and Zero Motion Candidates can be inserted sequentially. The signaling in the stripe header informs the size of the regular Merge list, with a maximum value of 6. MMVD, TPM, and CIIP depend on the regular Merge list.

[0130] Similarly, VVC supports three lists of AMVPs:

[0131] 1) Affine AMVP Candidate List

[0132] 2) Regular AMVP Candidate List

[0133] 2.2.1 Codec Block Structure in VVC

[0134] In VVC, a quadtree / binary tree / ternary tree (QT / BT / TT) structure is used to divide the image into square or rectangular blocks.

[0135] In addition to QT / BT / TT, VVC also employs a separate tree (also known as a dual codec tree) for I-frames. The separate tree is used to signal the codec block structure separately for the luminance and chrominance components.

[0136] In addition, except for blocks encoded using several specific encoding / decoding methods (such as intra-frame sub-segmentation prediction where PU equals TU but is less than CU, and sub-block transformation of inter-frame encoded / decoded blocks where PU equals CU but TU is less than PU), CU is set to equal PU and TU.

[0137] 2.2.2 Merge the entire block

[0138] 2.2.2.1 Constructing the Merge List from the Regular Merge Pattern

[0139] 2.2.2.1.1 History-Based Motion Vector Prediction (HMVP)

[0140] Unlike the Merge list design, VVC uses a history-based motion vector prediction (HMVP) method.

[0141] The HMVP stores motion information from previously encoded / decoded blocks. Motion information from previously encoded / decoded blocks is defined as HMVP candidates. Multiple HMVP candidates are stored in a table called the HMVP table, and this table is maintained during on-the-fly encoding / decoding. When encoding / decoding a new slice / LCU line / strip begins, the HMVP table is cleared. Whenever an inter-frame encoded / decoded block or non-sub-block, non-TPM mode exists, the associated motion information is added as a new HMVP candidate to the last entry in the table. Figure 10 The overall encoding and decoding process is described in the document.

[0142] 2.2.2.1.2 Standard Merge List Construction Process

[0143] The construction of a regular Merge list (for translational motion) can be summarized in the following steps:

[0144] Step 1: Derive airspace candidates

[0145] Step 2: Insert HMVP candidate

[0146] Step 3: Insert pairwise average candidates

[0147] Step 4: Default motion candidates

[0148] HMVP candidates can be used in both the AMVP and Merge candidate list construction processes. Figure 11 The modified Merge Candidate List construction process is depicted (highlighted in blue). When the Merge Candidate List is not full after inserting TMVP candidates, it can be populated using HMVP candidates stored in the HMVP table. Considering that a block is generally highly correlated with its nearest neighbor in terms of motion information, HMVP candidates are inserted into the table in descending index order. The last entry in the table is added to the list first, and the first entry is added last. Similarly, redundancy removal is applied to HMVP candidates. Once the total number of available Merge Candidates reaches the maximum number of Merge Candidates allowed for signaling notification, the Merge Candidate List construction process terminates.

[0149] It is important to note that all spatial / temporal / HMVP candidates should be encoded and decoded in non-IBC mode. Otherwise, they are not allowed to be added to the regular merge candidate list.

[0150] The HMVP table contains up to 5 regular motion candidates, and each candidate is unique.

[0151] 2.2.2.1.2.1 Pruning

[0152] Candidates are added to the list only if the corresponding candidates used for redundancy checks do not have the same motion information. This comparison process is called pruning.

[0153] The pruning process between candidates in the airspace depends on the TPM used in the current block.

[0154] When the current block is not encoded and decoded using TPM mode (e.g., regular Merge, MMVD, CIIP), HEVC pruning processing (e.g., five prunings) for spatial merge candidates is utilized.

[0155] 2.2.3 Triangular Prediction Model (TPM)

[0156] In VVC, inter-frame prediction in triangular segmentation mode is supported. Triangular segmentation mode is only applicable to CUs of 8x8 or larger that are encoded and decoded in Merge mode rather than MMVD or CIIP mode. For CUs that meet these conditions, signaling informs the CU level flag to indicate whether triangular segmentation mode is applied.

[0157] like Figure 11 As depicted, when this mode is used, the CU is uniformly divided into two triangular segments using either diagonal or anti-diagonal partitioning. Each triangular segment in the CU is used for inter-frame prediction using its own motion; each segment allows only unidirectional prediction, i.e., each segment has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion-compensated predictions are required for each CU.

[0158] Figure 12 An example of inter-frame prediction based on triangulation is shown.

[0159] If the CU level flag indicates that the current CU is encoded and decoded using the triangulation mode, then signaling is also provided indicating the triangulation direction (diagonal or anti-diagonal) and two merge indices (one for each segment). After predicting each triangulation, a mixing process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal edges. This is the predicted signal for the entire CU, and like other prediction modes, transform and quantization processing is applied to the entire CU. Finally, the CU motion field predicted using the triangulation mode is stored in 4x4 cells.

[0160] The regular Merge candidate list is reused for triangulation Merge prediction without additional motion vector pruning. For each Merge candidate in the regular Merge candidate list, exactly one of its L0 or L1 motion vectors is used for triangulation prediction. Furthermore, the order of the L0 and L1 motion vectors is selected based on the parity of their Merge index. Using this scheme, the regular Merge list can be used directly.

[0161] 2.2.3.1 TPM Merge List Construction Process

[0162] Basically, some modifications can be added to the regular Merge list building process.

[0163] Specifically, the following were applied:

[0164] 1) How to perform pruning depends on the TPM used in the current block.

[0165] - If TPM encoding / decoding of the current block is not utilized, then HEVC 5 pruning applied to the spatial merge candidate is invoked.

[0166] Otherwise (if the current block is encoded / decoded using TPM), full pruning is applied when a new spatial merge candidate is added. That is, B1 is compared with A1; B0 is compared with A1 and B1; A0 is compared with A1, B1, and B0; and B2 is compared with A1, B1, A0, and B0.

[0167] 2) Whether to check motion information from B2 depends on the current block's TPM usage.

[0168] – If TPM encoding and decoding of the current block is not utilized, B2 is accessed and inspected only if there are fewer than 4 space merge candidates before inspecting B2.

[0169] Otherwise (if the current block is encoded or decoded using TPM), B2 is always accessed and checked before adding B2, regardless of how many available space merge candidates there are.

[0170] 2.2.3.2 Adaptive Weighted Processing

[0171] After predicting each triangular prediction unit, an adaptive weighting process is applied to the diagonal edges between two triangular prediction units to derive the final prediction for the entire CU. Two sets of weighting factors are defined below:

[0172] • The first weighting factor group: {7 / 8, 6 / 8, 4 / 8, 2 / 8, 1 / 8} and {7 / 8, 4 / 8, 1 / 8} are used for luminance and chrominance samples, respectively;

[0173] • The second weighting factor group: {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8} are used for luminance and chrominance samples, respectively.

[0174] The weighting factor set is selected based on a comparison of the motion vectors of two triangular prediction units. The second weighting factor set is used when any of the following conditions are true:

[0175] – The reference images for the two triangular prediction units are different.

[0176] - The absolute value of the difference between the horizontal values ​​of two motion vectors is greater than 16 pixels.

[0177] – The absolute value of the difference between the perpendicular values ​​of two motion vectors is greater than 16 pixels.

[0178] Otherwise, use the first weighted factor group. Figure 13 An example is shown.

[0179] 2.2.3.3 Motion Vector Storage

[0180] The motion vector of the triangular prediction unit ( Figure 14 Mv1 and Mv2 in the CU are stored in a 4×4 grid. For each 4×4 grid, depending on the position of the 4×4 grid in the CU, either unidirectional or bidirectional predicted motion vectors are stored. Figure 14 As shown, the unidirectional predicted motion vector (Mv1 or Mv2) is stored in a 4×4 grid located in the unweighted region (i.e., not at the diagonal edge). On the other hand, the bidirectional predicted motion vector is stored in a 4×4 grid located in the weighted region. The bidirectional predicted motion vector is derived from Mv1 and Mv2 according to the following rules:

[0181] 1) In the case where Mv1 and Mv2 have motion vectors from different directions (L0 or L1), Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.

[0182] 2) When Mv1 and Mv2 both originate from the same L0 (or L1) direction,

[0183] – If the reference image for Mv2 is the same as an image in the L1 (or L0) reference image list, then Mv2 is scaled to that image. Mv1 and the scaled Mv2 are combined to form a bidirectional predicted motion vector.

[0184] – If the reference image for Mv1 is the same as an image in the L1 (or L0) reference image list, then scale Mv1 to that image. Combine the scaled Mv1 and Mv2 to form a bidirectional predicted motion vector.

[0185] Otherwise, store Mv1 only in the weighted region.

[0186] 2.2.4 Merge (MMVD) with Motion Vector Difference

[0187] In some embodiments, the final motion vector representation (UMVE, also known as MMVD) and the proposed motion vector representation method are used for skip or merge modes.

[0188] UMVE reuses the same Merge candidates as those included in the regular Merge candidate list in VVC. Within the Merge candidate list, a base candidate can be selected, and this base candidate can be further extended using the proposed motion vector representation method.

[0189] UMVE provides a new representation of motion vector difference (MVD), which uses the origin, amplitude, and direction of motion to represent MVD.

[0190] Figure 15 This is an example of UMVE search processing.

[0191] Figure 16 An example of a UMVE search point is shown.

[0192] This technique uses the Merge candidate list as is. However, only candidates of the default Merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE extensions.

[0193] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.

[0194] Table 4 Basic Candidate IDX

[0195] Basic candidate IDX 0 1 2 3 The Nth MVP First MVP Second MVP The third MVP The fourth MVP

[0196] If the number of basic candidates is equal to 1, then no signaling is sent to the basic candidate IDX.

[0197] The distance index is motion amplitude information. The distance index indicates a predefined distance from the starting point. The predefined distances are shown below:

[0198] Table 5 Distance from IDX

[0199] Distance from IDX 0 1 2 3 4 5 6 7 Pixel distance 1 / 4 pixel 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels 16 pixels 32 pixels

[0200] The direction index represents the direction of MVD relative to the starting point. The direction index can represent the four directions shown below.

[0201] Table 6 Directional IDX

[0202] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0203] Immediately after sending the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, the UMVE flag is resolved. If the UMVE flag is equal to 1, the UMVE syntax is resolved. However, if it is not 1, the AFFINE flag is resolved. If the AFFINE flag is equal to 1, it is an AFFINE mode; however, if it is not 1, the skip / merge index of the VTM's skip / merge mode is resolved.

[0204] The additional line buffer due to UMVE candidates is unnecessary because the software's skip / merge candidates are directly used as the base candidates. MV supplementation is determined directly before motion compensation using the input UMVE index. There is no need to reserve a long line buffer for this.

[0205] Under the current normal testing conditions, the first or second Merge candidate in the Merge candidate list can be selected as the base candidate.

[0206] UMVE is also known as Merge with MV difference (MMVD).

[0207] 2.2.5 MERGE based on sub-block technology

[0208] In some embodiments, all sub-block-related motion candidates are placed in a separate Merge list in addition to the regular Merge list of non-sub-block Merge candidates.

[0209] Add the sub-block-related motion candidates to a separate Merge list called "Sub-block Merge Candidate List".

[0210] In one example, the list of sub-block merge candidates includes ATMVP candidates and affine merge candidates.

[0211] Fill the candidate list for sub-block Merge in the following order:

[0212] a. ATMVP candidate (may be available or not);

[0213] b. A list of Affine Merge (including inherited affine candidates; and constructed affine candidates);

[0214] c. Zero-filled MV 4-parameter affine model

[0215] 2.2.5.1 Advanced Temporal Motion Vector Predictor (ATMVP) (also known as Sub-Block Temporal Motion Vector Predictor, SbTMVP)

[0216] The basic idea of ​​ATMVP is to derive multiple sets of temporal motion vector predictors for a block. Each sub-block is assigned a set of motion information. When generating ATMVP Merge candidates, motion compensation is performed at the 8x8 level rather than at the entire block level.

[0217] In the current design, ATMVP predicts the motion vectors of sub-CUs in the CU in two steps, which are described in the following two subsections 2.2.5.1.1 and 2.2.5.1.2 respectively.

[0218] 2.2.5.1.1 Derivation of Initialized Motion Vectors

[0219] The initial motion vector is represented by tempMv. When block A1 is available and not intra-frame encoded (e.g., inter-frame or IBC mode encoded), the following is applied to derive the initial motion vector.

[0220] – If all of the following conditions are true, then tempMv is set to be equal to the motion vector of block A1 in list 1, denoted as mvL1A1:

[0221] – The reference image index for list 1 is available (not equal to -1) and it has the same POC value as the juxtaposed image (e.g., DiffPicOrderCnt(ColPic, RefPicList[1][refIdxL1A1]) equals 0).

[0222] – Compared to the current image, none of the reference images have a larger POC (e.g., for each image aPic in the reference image list for the current strip, DiffPicOrderCnt(aPic, currPic) is less than or equal to 0).

[0223] –The current stripe is equal to stripe B.

[0224] –collocated_from_l0_flag equals 0.

[0225] Otherwise, if all of the following conditions are true, then tempMv is set to be equal to the motion vector of block A1 in list 0, denoted as mvL0A1:

[0226] – The reference image index for list 0 is available (not equal to -1).

[0227] – It has the same POC value as the juxtaposed image (e.g., DiffPicOrderCnt(ColPic,RefPicList[0][refIdxL0A1]) equals 0).

[0228] Otherwise, use the zero motion vector as the initial MV.

[0229] Identify the corresponding block in the juxtaposed image of the signaling notification at the strip header with initialized motion vector (with the center position of the current block plus the rounded MV, which is trimmed to a certain range if necessary).

[0230] If the block is inter-frame encoded, proceed to step two. Otherwise, set the ATMVP candidate to unavailable.

[0231] 2.2.5.1.2 Derivation of Sub-CU Motion

[0232] The second step is to divide the current CU into sub-CUs and obtain the motion information of each sub-CU from the block corresponding to each sub-CU in the juxtaposed image.

[0233] If the corresponding block of a sub-CU is encoded and decoded in inter-frame mode, the final motion information of the current sub-CU is derived using motion information by invoking the derivation process for co-locating the motion vector (MV). This derivation process is no different from that used for regular TMVP processing. Basically, if a corresponding block is predicted from the target list X for unidirectional or bidirectional prediction, the motion vector is used; otherwise, if a corresponding block is predicted from list Y (Y = 1-X) for unidirectional or bidirectional prediction, and NoBackwardPredFlag equals 1, the MV of list Y is used. Otherwise, no motion candidate can be found.

[0234] If the blocks in the juxtaposed image identified by the initial MV and the current sub-CU position are intra-frame or IBC encoded, or if motion candidates cannot be found as described above, then the following rules further apply:

[0235] R will be used to capture and juxtapose images. col The motion vector of the motion field in the equation is represented by MV. col To minimize the impact of MV scaling, the method used to derive the MV is selected as follows. col MV in the airspace candidate list: If the reference image of the candidate MV is a juxtaposed image, then select this MV and use it as the MV. col And without any scaling. Alternatively, select the MV with the reference image that is closest to the juxtaposed image to derive the MV using scaling. col .

[0236] The relevant decoding process for juxtaposing motion vector derivation in some embodiments is described below:

[0237] 8.5.2.12 Derivation of the juxtaposed motion vectors

[0238] The input to this process is:

[0239] – The variable currCb specifies the current encoding / decoding block.

[0240] – The variable colCb specifies the juxtaposition encoding / decoding blocks within the juxtaposition image defined by ColPic.

[0241] – Luminance position (xColCb, yColCb), specifies the top-left sample of the juxtaposed luminance block defined by colCb relative to the top-left luminance sample of the juxtaposed image defined by ColPic.

[0242] – Refer to the index refIdxLX, where X is 0 or 1.

[0243] – The flag sbFlag indicates the candidate for time-domain Merging in the sub-block.

[0244] The output of this process is:

[0245] Motion vector prediction with a precision of -1 / 16 fractional sample points, mvLXCol

[0246] –Availability flag: availableFlagLXCol.

[0247] The variable currPic specifies the current image.

[0248] Set the arrays predFlagL0Col[x][y], mvL0Col[x][y], and refIdxL0Col[x][y] to be equal to the PredFlagL0[x][y], MvDmvrL0[x][y], and RefIdxL0[x][y] of the juxtaposed images specified by ColPic, respectively. Set the arrays predFlagL1Col[x][y], mvL1Col[x][y], and refIdxL1Col[x][y] to be equal to the PredFlagL1[x][y], MvDmvrL1[x][y], and RefIdxL1[x][y] of the juxtaposed images specified by ColPic, respectively.

[0249] The derivation of variables mvLXCol and availableFlagLXCol is as follows:

[0250] – If colCb is encoded and decoded in intra-frame or IBC prediction mode, set both components of mvLXCol to 0 and set availableFlagLXCol to 0.

[0251] – Otherwise, derive the motion vector mvCol, reference index refIdxCol, and reference list identifier listCol as described below:

[0252] – If sbFlag equals 0, then set availableFlagLXCol to equal 1, and apply the following rules:

[0253] – If predFlagL0Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL1Col[xColCb][yColCb], refIdxL1Col[xColCb][yColCb], and L1, respectively.

[0254] Otherwise, if predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb] equals 0, then set mvCol, refIdxCol, and listCol to equal mvL0Col[xColCb][yColCb], refIdxL0Col[xColCb][yColCb], and L0, respectively.

[0255] Otherwise, (predFlagL0Col[xColCb][yColCb] equals 1 and predFlagL1Col[xColCb][yColCb])

[0256] If the value is equal to 1, perform the following allocation:

[0257] – If NoBackwardPredFlag equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX, respectively.

[0258] Otherwise, set mvCol, refIdxCol, and listCol to equal mvLNCol[xColCb][yColCb],

[0259] refIdxLNCol[xColCb][yColCb] and LN, where N is the value of collocated_from_l0_flag.

[0260] Otherwise (sbFlag equals 1), the following rules apply:

[0261] – If PredFlagLXCol[xColCb][yColCb] equals 1, then set mvCol, refIdxCol, and listCol to equal mvLXCol[xColCb][yColCb], refIdxLXCol[xColCb][yColCb], and LX respectively, and set availableFlagLXCol to 1.

[0262] – Otherwise (PredFlagLXCol[xColCb][yColCb] equals 0), the following rule applies:

[0263] – If DiffPicOrderCnt(aPic,currPic) is less than or equal to 0 for each image aPic in the reference image list for the current strip, and PredFlagLYCol[xColCb][yColCb]...

[0264] If the value is 1, then set mvCol, refIdxCol, and listCol to mvLYCol[xColCb][yColCb], refIdxLYCol[xColCb][yColCb], and LY respectively, where Y equals ! X, where X is the value of X used to call this process. Set availableFlagLXCol to 1.

[0265] – Set both components of mvLXCol to 0, and set availableFlagLXCol to equal 0.

[0266] – When availableFlagLXCol is true (TRUE), derive mvLXCol and availableFlagLXCol as follows:

[0267] – If LongTermRefPic(currPic, currCb, refIdxLX, LX) is not equal to LongTermRefPic

[0268] If (ColPic, colCb, refIdxCol, listCol) is used, then both components of mvLXCol will be set to 0, and availableFlagLXCol will be set to 0.

[0269] Otherwise, set the variable availableFlagLXCol to 1 and set refPicList[listCol][refIdxCol] to 1.

[0270] Set the image with reference index refIdxCol in the list of reference images listCol, which has stripes. The stripes contain the codec blocks colCb in the juxtaposed images as specified by ColPic, and the following rules apply:

[0271] colPocDiff=DiffPicOrderCnt(ColPic,refPicList[listCol][refIdxCol])(8-402)

[0272] currPocDiff=DiffPicOrderCnt(currPic,RefPicList[X][refIdxLX]) (8-403)

[0273] – Invoke the temporal motion buffer compression process for juxtaposing motion vectors as specified in Clause 8.5.2.15, with mvCol as input and a modified mvCol as output.

[0274] – If RefPicList[X][refIdxLX] is a long-term reference image, or colPocDiff equals currPocDiff, then the derivation of mvLXCol is as follows:

[0275] mvLXCol=mvCol (8-404)

[0276] – Otherwise, derive mvLXCol as a scaled version of the motion vector mvCol as described below:

[0277] tx= ( 16384 + ( Abs( td ) >> 1 ) ) / td (8-405)

[0278] distScaleFactor = Clip3( -4096, 4095, ( tb * tx + 32 ) >> 6 ) (8-406)

[0279] mvLXCol= Clip3(-131072,131071,(distScaleFactor*mvCol+128 -

[0280] ( distScaleFactor * mvCol >= 0 ) ) >> 8 ) ) (8-407)

[0281] The derivation of td and tb is as follows:

[0282] td = Clip3( -128, 127, colPocDiff ) (8-408)

[0283] tb = Clip3( -128, 127, currPocDiff ) (8-409)

[0284] 2.3 Intra-frame block copying

[0285] The HEVC Screen Content Codec Extension (HEVC-SCC) and the current VVC test model (VTM-4.0) employ Intra-Block Copy (IBC) (also known as Current Picture Reference). IBC extends the concept of motion compensation from inter-frame coding to intra-frame coding. For example... Figure 17 As shown, when IBC is applied, the current block is predicted using a reference block within the same image. Samples in the reference block must be reconstructed before encoding or decoding the current block. While IBC is not sufficiently efficient for most camera-captured sequences, it demonstrates significant encoding / decoding gains for screen content. This is because screen content images contain many repeating patterns, such as icons and text characters. IBC effectively removes redundancy between these repeating patterns. In HEVC-SCC, IBC can be applied to the codec unit (CU) of inter-frame encoding / decoding if the current image is chosen as its reference image. In this case, the MV is renamed to a block vector (BV), and the BV always has integer pixel precision. For compatibility with the main HEVC profile, the current image is marked as the "long-term" reference image in the decoded picture buffer (DPB). It should be noted that, similarly, in multi-view / 3D video codec standards, inter-view reference images are also marked as "long-term" reference images.

[0286] After BV finds its reference block, predictions are generated by copying the reference block. The residual can be obtained by subtracting the reference pixel from the original signal. Transformation and quantization can then be applied as in other codec modes.

[0287] However, when the reference block is outside the image, overlaps with the current block, is outside the reconstructed region, or is outside the valid region subject to certain constraints, some or all pixel values ​​are not defined. Basically, there are two solutions to this problem. One is to prohibit this situation, for example, in terms of bitstream consistency. The other is to apply padding to those undefined pixel values. The following subsections describe the solutions in detail.

[0288] 2.3.1 Single BV List

[0289] In some embodiments, the BV predictors used for Merge mode and AMVP mode in IBC will share a common predictor list, which includes the following elements:

[0290] (1) Two adjacent airspace locations (e.g.) Figure 2 (A1, B1)

[0291] (2) 5 HMVP entries

[0292] (3) Default zero vector

[0293] The number of candidates in the list is controlled by variables derived from the strip header. For Merge mode, a maximum of the first 6 entries of this list will be used; for AMVP mode, the first 2 entries of this list will be used. Furthermore, the list must conform to the shared Merge list region requirement (the same list is shared within SMR).

[0294] In addition to the BV predictor candidate list mentioned above, the pruning operation between the HMVP candidate and the existing Merge candidate (A1, B1) can be simplified. In the simplified process, there will be a maximum of two pruning operations, as it only compares the first HMVP candidate with the spatial Merge candidate.

[0295] In IBC AMVP mode, the MV difference compared to the selected MVP in the list is also transmitted. In IBC Merge mode, the selected MVP is directly used as the MV of the current block.

[0296] 2.3.2 Size Limitations of IBC

[0297] In the latest VVC and VTM5, a syntax constraint was introduced to explicitly disable the 128x128 IBC mode, based on the current bitstream constraints, which was present in previous VTM and VVC versions. This makes the presence of the IBC flag dependent on the CU size <128x128.

[0298] 2.3.3 IBC Shared Merge List

[0299] To reduce decoder complexity and support parallel encoding, in some embodiments, all leaf codec units (CUs) for an ancestor node in the CU partitioning tree share the same merging candidate list to enable parallel processing of CUs with small skip / merge encoding / decoding. The ancestor node is named the Merge Shared Node. A shared merging candidate list is generated at the Merge Shared Node, assuming the Merge Shared Node is a leaf CU.

[0300] More specifically, the following can be applied:

[0301] - If a block has no more than 32 luminance samples (e.g., 4x8 or 8x4) and is divided into two 4x4 child blocks, a shared Merge list is used between very small blocks (e.g., two adjacent 4x4 blocks).

[0302] - If a block has more than 32 luminance samples, but after partitioning, at least one child block is less than the threshold (32), then all child blocks of the partition share the same Merge list (e.g., 16x4 or 4x16 for triple partitioning or 8x8 for quadruple partitioning).

[0303] This restriction applies only to IBC Merge mode.

[0304] 2.3.4 Syntax Table

[0305] 7.3.8.5 Encoding / Decoding Unit Syntax

[0306]

[0307]

[0308] 7.3.8.8 Motion Vector Difference Syntax

[0309]

[0310] 2.4 Maximum Transform Block Size

[0311] The maximum luminance transform size is 64 or 32, enabled only by the SPS level flag. The maximum chrominance transform size is derived from the chrominance sampling ratio relative to the maximum luminance transform size.

[0312] When the CU / CB size is larger than the maximum brightness conversion size, patches from the smaller TU to the larger CU can be used.

[0313] MaxTbSizeY represents the maximum brightness transformation size.

[0314] 7.3.2.3 Sequence Parameter Set (RBSP) Syntax

[0315]

[0316]

[0317] A value of 1 for sps_max_luma_transform_size_64_flag indicates that the maximum transformation size of the luminance sample is 64. A value of 0 for sps_max_luma_transform_size_64_flag indicates that the maximum transformation size of the luminance sample is 32.

[0318] When CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag should be equal to 0.

[0319] Derive the variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSizeY, and MaxTbSizeY as follows:

[0320] MinTbLog2SizeY = 2 (7-27)

[0321] MaxTbLog2SizeY = sps_max_luma_transform_size_64_flag? 6 : 5 (7-28)

[0322] MinTbSizeY = 1 << MinTbLog2SizeY (7-29)

[0323] MaxTbSizeY = 1 << MaxTbLog2SizeY (7-30)

[0324] A value of 0 for `sps_sbt_max_size_64_flag` specifies that the maximum CU width and height allowed for subblock transformation is 32 luminance samples. A value of 1 for `sps_sbt_max_size_64_flag` specifies that the maximum CU width and height allowed for subblock transformation is 64 luminance samples.

[0325] MaxSbtSize= Min(MaxTbSizeY, sps_sbt_max_size_64_flag? 64:32) (7-32)

[0326] 2.4.1 Tools that depend on the maximum transform block size

[0327] The partitioning of binary and ternary trees depends on MaxTbSizeY.

[0328] 7.3.8.5 Encoding / Decoding Unit Syntax

[0329]

[0330]

[0331]

[0332]

[0333]

[0334] 2.4.2 Decoding Processing of Intra-Frame and Inter-Frame Coded Blocks

[0335] When the size of the current intra-codec CU is larger than the maximum transform block size, the CU is recursively divided into smaller blocks (PUs) until both the width and height are no larger than the maximum transform block size. All PUs share the same intra-prediction mode, but the reference samples used for intra-prediction are different. Simultaneously, the TU is set to be equal to the PU. Note that the current block should not be encoded or decoded in Intra-Segmentation (ISP) mode.

[0336] When the size of the current inter-frame codec's CU is larger than the maximum transform block size, the CU is recursively divided into smaller blocks (TUs) until both the width and height are no larger than the maximum transform block size. All TUs share the same motion information, but use different reference samples for intra-frame prediction.

[0337] 8.4.5 Intra-frame block decoding processing

[0338] 8.4.5.1 General Decoding Processing of Intra-Frame Blocks

[0339] The input to this process is:

[0340] –Sample position (xTb0, yTb0), specifies the top-left sample of the current transform block relative to the top-left sample of the current image.

[0341] – The variable nTbW specifies the width of the current transform block.

[0342] – The variable nTbH specifies the height of the current transform block.

[0343] – The variable predModeIntra specifies the intra-frame prediction mode.

[0344] – The variable cIdx specifies the color component of the current block.

[0345] The output of this process is the reconstructed image after the loop filtering.

[0346] The maximum transform block width (maxTbWidth) and height (maxTbHeight) are derived as follows:

[0347] maxTbWidth = (cIdx = = 0)? MaxTbSizeY : MaxTbSizeY / SubWidthC(8-41)

[0348] maxTbHeight = (cIdx = = 0)? MaxTbSizeY : MaxTbSizeY / SubHeightC (8-42)

[0349] Derive the brightness sample point positions based on the following:

[0350] (xTbY,yTbY)=(cIdx==0)? (xTb0,yTb0):(xTb0*SubWidthC,yTb0*SubHeightC)(8-43)

[0351] Depending on maxTbSize, the following rules apply:

[0352] – If Inter-SubpartitionsSplitType equals ISP_NO_SPLIT, and nTbW is greater than maxTbWidth or nTbH is greater than maxTbHeight, then the following sequential steps apply.

[0353] 1. Derive the variables newTbW and newTbH according to the following:

[0354] newTbW = (nTbW > maxTbWidth)? ( nTbW / 2 ) : nTbW (8-44)

[0355] newTbH = (nTbH > maxTbHeight)? ( nTbH / 2 ) : nTbH (8-45)

[0356] 2. Taking the position (xTb0, yTb0), transform block width nTbW set to equal newTbW, transform block height nTbH set to equal newTbH, intra-prediction mode predModeIntra, and variable cIdx as input, the general decoding process of the intra-block as specified in this clause is invoked, and the output is the modified reconstructed image before loop filtering.

[0357] 3. If nTbW is greater than maxTbWidth, then set the position (xTb0, yTb0) to be equal to (xTb0 + newTbW, yTb0), the transform block width nTbW to be equal to newTbW, and the transform block height nTbH.

[0358] Set to newTbH, intra-prediction mode predModeIntra, and variable cIdx as input, invoke the general decoding process for intra-blocks as specified in this clause, and output the modified reconstructed image before loop filtering.

[0359] 4. If nTbH is greater than maxTbHeight, then the general decoding process for the intra-block specified in this clause is called with the position (xTb0, yTb0) set to equal (xTb0, yTb0+newTbH), the transform block width nTbW set to equal newTbW, the transform block height nTbH set to equal newTbH, the intra-prediction mode predModeIntra, and the variable cIdx as input. The output is the modified reconstructed image before loop filtering.

[0360] 5. If nTbW is greater than maxTbWidth and nTbH is greater than maxTbHeight, then the position (xTb0, yTb0) is set to be equal to (xTb0+newTbW, yTb0+newTbH), the transform block width nTbW is set to be equal to newTbW, the transform block height nTbH is set to be equal to newTbH, and the intra-prediction mode predModeIntra is set to...

[0361] The variable cIdx is used as input to call the general decoding process of the intra-block as specified in this clause, and the output is the modified reconstructed image before loop filtering.

[0362] – Otherwise, apply the following sequential steps (standard intra-frame prediction processing):

[0363]

[0364]

[0365] 8.5.8 Decoding Processing of Residual Signals of Codec Blocks Encoded and Decoded in Inter-Frame Prediction Mode

[0366] The input to this process is:

[0367] –Sample position (xTb0, yTb0), specifies the top-left sample of the current transform block relative to the top-left sample of the current image.

[0368] – The variable nTbW specifies the width of the current transform block.

[0369] – The variable nTbH specifies the height of the current transform block.

[0370] – The variable cIdx specifies the color component of the current block.

[0371] The output of this process is an array of (nTbW)×(nTbH) resSamples.

[0372] The maximum transform block width (maxTbWidth) and height (maxTbHeight) are derived as follows:

[0373] maxTbWidth = (cIdx = = 0)? MaxTbSizeY : MaxTbSizeY / SubWidthC(8-883)

[0374] maxTbHeight = (cIdx = = 0)? MaxTbSizeY : MaxTbSizeY / SubHeightC (8-884)

[0375] Derive the brightness sample point positions based on the following:

[0376] (xTbY,yTbY)=(cIdx==0)? (xTb0,yTb0):(xTb0*SubWidthC,yTb0*SubHeightC)(8-885)

[0377] Depending on maxTbSize, the following rules apply:

[0378] – If nTbW is greater than maxTbWidth or nTbH is greater than maxTbHeight, then the following sequential steps apply.

[0379] 1. Derive the variables newTbW and newTbH according to the following:

[0380] newTbW = (nTbW > maxTbWidth)? ( nTbW / 2 ) : nTbW (8-886)

[0381] newTbH = (nTbH > maxTbHeight)? ( nTbH / 2 ) : nTbH (8-887)

[0382] 2. Taking the position (xTb0, yTb0), transform block width nTbW set to equal newTbW, transform block height nTbH set to equal newTbH, and variable cIdx as input, the decoding process of the residual signal of the codec unit in the inter-frame prediction mode as specified in this clause is invoked. The output is the modified reconstructed image before loop filtering.

[0383] 3. When nTbW is greater than maxTbWidth, the codec processing procedure for the residual signal of the codec unit in the inter-frame prediction mode is called with the position (xTb0, yTb0) set to equal (xTb0+newTbW, yTb0), the transform block width nTbW set to equal newTbW, the transform block height nTbH set to equal newTbH, and the variable cIdx as input. The output is the modified reconstructed image.

[0384] 4. When nTbH is greater than maxTbHeight, the position (xTb0, yTb0) is set to be equal to (xTb0, yTb0 + newTbH), the transform block width nTbW is set to be equal to newTbW, the transform block height nTbH is set to be equal to newTbH, and the variable cIdx is used as input. The decoding process of the residual signal of the codec unit in the inter-frame prediction mode codec specified in this clause is called. The output is the modified reconstructed image before loop filtering.

[0385] 5. When nTbW is greater than maxTbWidth and nTbH is greater than maxTbHeight, the codec processing procedure for the residual signal of the codec unit in the inter-frame prediction mode, as specified in this clause, is called with the position (xTb0, yTb0) set to equal (xTb0+newTbW, yTb0+newTbH), the transform block width nTbW set to equal newTbW, the transform block height nTbH set to equal newTbH, and the variable cIdx as input. The output is the modified reconstructed image before loop filtering.

[0386] Otherwise, if cu_sbt_flag equals 1, the following rules apply:

[0387] – Derive the variables sbtMinNumFourths, wPartIdx, and hPartIdx as follows:

[0388] sbtMinNumFourths = cu_sbt_quad_flag? 1 : 2 (8-888)

[0389] wPartIdx = cu_sbt_horizontal_flag? 4: sbtMinNumFourths (8-889)

[0390] hPartIdx = ! cu_sbt_horizontal_flag? 4: sbtMinNumFourths (8-890)

[0391] –Derive the variables xPartIdx and yPartIdx as follows:

[0392] – If cu_sbt_pos_flag equals 0, then set xPartIdx and yPartIdx to equal 0.

[0393] Otherwise (cu_sbt_pos_flag equals 1), derive the variables xPartIdx and yPartIdx as follows:

[0394] xPartIdx = cu_sbt_horizontal_flag? 0 : ( 4 - sbtMinNumFourths ) (8-891)

[0395] yPartIdx = ! cu_sbt_horizontal_flag? 0 : ( 4 - sbtMinNumFourths )(8-892)

[0396] –Derive the variables xTbYSub, yTbYSub, xTb0Sub, yTb0Sub, nTbWSub, and nTbHSub as follows:

[0397] xTbYSub=xTbY+((nTbW*((cIdx==0)?1:SubWidthC)*xPartIdx / 4)

[0398] (8-893)

[0399] yTbYSub=yTbY+((nTbH*((cIdx==0)?1:SubHeightC)*yPartIdx / 4)

[0400] (8-894)

[0401] xTb0Sub = xTb0 + ( nTbW * xPartIdx / 4 ) (8-895)

[0402] yTb0Sub = yTb0 + ( nTbH * yPartIdx / 4 ) (8-896)

[0403] nTbWSub = nTbW * wPartIdx / 4 (8-897)

[0404] nTbHSub = nTbH * hPartIdx / 4 (8-898)

[0405] – Taking the luminance position (xTbYSub, yTbYSub), variables cIdx, nTbWSub, and nTbHSub as input, invoke the scaling and transformation processing specified in Item 8.7.2, and the output is (nTbWSub) × (nTbHSub).

[0406] Array resSamplesTb.

[0407] – Set the residual samples resSamples[x][y] with x = 0..nTbW–1 and y = 0..nTbH–1 to equal 0.

[0408] – Based on the following, we derive the following: x=xTb0Sub..xTb0Sub+nTbWSub–1、

[0409] The residual samples resSamples[x][y] of y = yTb0Sub..yTb0Sub+nTbHSub–1 are:

[0410] resSamples[ x ][ y ] = resSamplesTb[ x – xTb0Sub ][ y – yTb0Sub ](8-899)

[0411] Otherwise, with the brightness position (xTbY, yTbY), variable cIdx, transformation width nTbW, and transformation height nTbH as input, the scaling and transformation processing specified in clause 8.7.2 is invoked, and the output is an array of (nTbW)×(nTbH) resSamples.

[0412] 2.5 Low-Frequency Inseparable Transform (LFNST) (also known as Reduced Quadratic Transform (RST) / Inseparable Quadratic Transform (NSST))

[0413] In JEM, a quadratic transform is applied between the positive master transform and quantization (at the encoder) and between dequantization and the inverse master transform (at the decoder). For example... Figure 18 As shown, the 4×4 (or 8×8) quadratic transformation is performed depending on the block size. For example, for each 8×8 block, the 8×8 quadratic transformation is applied to the larger block (e.g., minimum value (width, height) > 4), and the 4×4 quadratic transformation is applied to the smaller block (e.g., minimum value (width, height) < 8).

[0414] The quadratic transformation employs an inseparable transformation, hence it is also called the Inseparable Quadratic Transformation (NSST). There are 35 transformation sets in total, and each transformation set uses 3 inseparable transformation matrices (kernels, each kernel having a 16×16 matrix).

[0415] A reduced quadratic transform (RST, also known as low-frequency inseparable transform (LFNST)) is introduced, along with four transform sets (instead of 35) mappings based on the intra-frame prediction direction. In this paper, 16×48 and 16×16 matrices are used for 8×8 and 4×4 blocks, respectively. For ease of annotation, the 16×48 transform is represented as RST8×8, and the 16×16 transform as RST4×4. This approach has recently been adopted in VVC.

[0416] Figure 19 An example of the reduced quadratic transform (RST) is shown.

[0417] The second-order forward transform and the second-order inverse transform are processing steps independent of the main transform.

[0418] For the encoder, the main forward transform is performed first, followed by a second forward transform and quantization, and then CABAC bit encoding. For the decoder, CABAC bit decoding, inverse quantization, and a second inverse transform are performed first, followed by the main inverse transform. RST is only applied to the TU of intra-frame encoding and decoding.

[0419] 3. Example technical problems solved by the technical solutions disclosed in this document

[0420] Currently, VVC design has the following problems in terms of IBC mode and transformation design:

[0421] (1) Signaling notification of the IBC mode flag depends on the block size limit, but this is not the case for the IBC skip mode flag, which makes the worst case unchanged (e.g., a large block, such as N×128 or 128×N, can still select IBC mode).

[0422] (2) In the current VVC, it is assumed that the transform block size is always equal to 64×64, and the transform block size is always set to the codec block size. How to handle the case where the IBC block size is larger than the transform block size is unknown.

[0423] (3) IBC uses unfiltered build samples in the same video unit (e.g., stripe). However, filtered samples (e.g., via deblocking filter / SAO / ALF) can have less distortion compared to unfiltered samples. Using filtered samples can provide additional codec gain.

[0424] (4) Context modeling of the LFNST matrix depends on the master transformation and the segmentation type (single-tree or double-tree). However, from our analysis, there is no explicit dependency between the selected RST matrix and the master transformation.

[0425] 4. Examples of technologies and embodiments

[0426] In this application, Intra-Block Copying (IBC) is not limited to current IBC techniques, but can be interpreted as a technique that uses reference samples within the current strip / piece / brick / sub-picture / picture / other video unit (e.g., CTU line), excluding conventional intra-frame prediction methods. In one example, the reference samples are those reconstructed samples where loop filtering processes (e.g., deblocking filter, SAO, ALF) are not invoked.

[0427] The VPDU size can be represented as vSizeX * vSizeY. In one example, vSizeX = vSizeY = min(ctbSizeY, 64), where ctbSizeY is the width / height of the CTB. The VPDU is a vSize * vSize block with a top-left position (m * vSize, n * vSize) relative to the top-left corner of the image, where m and n are integers. Alternatively, vSizeX / vSizeY can be a fixed number, such as 64.

[0428] The following detailed inventions should be considered as examples for interpreting general concepts. These inventions should not be interpreted in a narrow sense. Furthermore, these inventions can be combined in any manner.

[0429] Signaling regarding IBC notify and use

[0430] Assume the current block size is represented as W. curr ×H curr The maximum allowed IBC block size is represented by W. IBCMax ×H IBCMax .

[0431] 1. IBC codec blocks can be divided into multiple TB / TUs without signaling notification of the division information. The current IBC code block size is represented as W. curr ×H curr .

[0432] a. In one example, W curr Greater than the first threshold and / or H curr If the value exceeds the second threshold, a partition can be initiated. Therefore, the TB size is smaller than the current block size.

[0433] i. In one example, the first and / or second thresholds are the maximum transform size, denoted as...

[0434] MaxTbSizeY.

[0435] b. In one example, recursive partitioning can be invoked until the width or height of TU / TB is no greater than a threshold.

[0436] i. In one example, the recursive partitioning can be the same as the partitioning used by the residual blocks in inter-frame encoding and decoding, for example, in each iteration, if the width is greater than MaxTbSizeY, the width is divided by half, and if the height is greater than MaxTbSizeY, the height is divided by half.

[0437] ii. Alternatively, recursive partitioning can be invoked until the partition width and height of TU / TB are each no greater than the threshold.

[0438] c. Alternatively, all or some TBs / TUs may share the same motion information (e.g., BV).

[0439] 2. Whether to enable or disable IBC may depend on the maximum transform size (e.g., MaxTbSizeY in the specification).

[0440] a. In one example, IBC can be disabled when the current luma block dimension (width and / or height) or (when the current block is a chroma block) the corresponding luma block dimension (width and / or height) is greater than the maximum transform size in the luma sample.

[0441] i. In one example, IBC can be disabled when both the width and height of the block are greater than the maximum transform size in the luminance sample.

[0442] ii. In one example, IBC can be disabled when the width or height of the block is greater than the maximum transform size in the luminance sample.

[0443] b. In one example, whether and / or how signaling notifications use IBC may depend on the dimensions of the block (width and / or height) and / or the maximum transform size.

[0444] i. In one example, the indication of using IBC can include IBC in a slice / strip / tile / subpicture.

[0445] Skip flags (e.g., cu_skip_flag).

[0446] 1) In one example, when W curr Greater than MaxTbSizeY and / or H curr When the value is greater than MaxTbSizeY, no signaling notification is required to cu_skip_flag.

[0447] 2) Alternatively, when W curr Not greater than MaxTbSizeY and / or H curr If the value is not greater than MaxTbSizeY, cu_skip_flag can be notified via signaling.

[0448] 3) In one example, when W curr Greater than MaxTbSizeY and / or H curr When the value is greater than MaxTbSizeY, cu_skip_flag can be signaled, but it should be equal to 0 in the consistent bitstream.

[0449] 4) For example, the above method can also be called when the current slice / strip / tile / sub-image is an I slice / strip / tile / sub-image.

[0450] ii. In one example, the indication for using IBC may include an IBC mode flag (e.g., pred_mode_ibc_flag).

[0451] 1) In one example, when W curr Greater than MaxTbSizeY and / or H curr When the value is greater than MaxTbSizeY, no signaling notification to pred_mode_ibc_flag is required.

[0452] 2) Alternatively, when W curr Not greater than MaxTbSizeY and / or H curr If the value is not greater than MaxTbSizeY, pred_mode_ibc_flag can be notified via signaling.

[0453] In one example, when W curr Greater than MaxTbSizeY and / or H curr When the value is greater than MaxTbSizeY, a signaling notification can be sent to pred_mode_ibc_flag, but it should be equal to 0 in its consistency bitstream.

[0454] iii. In one example, the residual blocks of an IBC codec block can be forced to all be zero when certain zero-forcing conditions are met.

[0455] 1) In one example, when W curr Greater than MaxTbSizeY and / or H curr When the value is greater than MaxTbSizeY, the residual blocks of the IBC codec block can be forced to be all zero.

[0456] 2) In one example, when W curr Greater than a first threshold (e.g., 64 / vSizeX) and / or H curr When the value is greater than a second threshold (e.g., 64 / vSizeY), the residual blocks of the IBC codec block can be forced to be all zero.

[0457] 3) Alternatively, in the above cases, the codec block flag can be skipped (e.g.,

[0458] Signaling notifications for cu_cbf, tu_cbf_cb, tu_cbf_cr, and tu_cbf_luma.

[0459] 4) Alternatively, in addition, codec block flags (e.g., cu_cbf, tu_cbf_cb, tu_cbf_cr, ...)

[0460] The signaling notifications for tu_cbf_luma can remain unchanged; however, the consistent bitstream should satisfy that the flag is equal to 0.

[0461] iv. In one example, whether signaling notification codec block flags (e.g., cu_cbf, tu_cbf_cb, tu_cbf_cr, tu_cbf_luma) are used may depend on the use of IBC and / or a threshold regarding the allowed IBC size.

[0462] 1) In one example, the current block is in IBC mode, and W curr Greater than MaxTbSizeY

[0463] and / or H curr When the value is greater than MaxTbSizeY, the codec block flags (e.g., cu_cbf) can be skipped.

[0464] Signaling notifications for tu_cbf_cb, tu_cbf_cr, and tu_cbf_luma.

[0465] 2) In one example, the current block is in IBC mode, and W curr Greater than the first threshold

[0466] (e.g., 64 / vSizeX) and / or H curr When the value exceeds the second threshold (e.g., 64 / vSizeY), signaling notifications for codec block flags (e.g., cu_cbf, tu_cbf_cb, tu_cbf_cr, tu_cbf_luma) can be skipped.

[0467] 3. Whether to signal the use of IBC skip mode can depend on the current block dimension (width and / or height) and the maximum allowed IBC block size (e.g., vSizeX / vSizeY).

[0468] a. In one example, when W curr Greater than W IBCMax and / or H curr Greater than H IBCMax At this time, it is possible to notify cu_skip_flag without signaling.

[0469] i. Alternatively, when W curr Not greater than W IBCMax and / or H curr Not greater than H IBCMax At that time, a signaling notification can be sent to cu_skip_flag.

[0470] b. Alternatively, the above method can also be invoked when the current slice / strip / tile / sub-image is an I slice / strip / tile / sub-image.

[0471] c. In one example, W IBCMax and H IBCMax Both equal 64.

[0472] d. In one example, W IBCMax Set to vSizeX, H IBCMax Set it to vSizeY.

[0473] e. In one example, W IBCMax Set the width of the maximum transform block, H IBCMax Set the height to be equal to the maximum transform block size.

[0474] Recursive partitioning of IBC / inter-frame codec blocks

[0475] 4. When an IBC / inter-frame codec block (CU) is divided into multiple TB / TUs and no signaling notification of the division information is given (e.g., the block size is larger than the maximum transform block size), but motion information is signaled to the entire CU once, multiple signaling notifications or motion information derivation are proposed.

[0476] a. In one example, all TBs / TUs can share the same AMVP / Merge flag, which can be signaled once.

[0477] b. In one example, a motion candidate list can be built once for the entire CU; however, different candidates can be assigned to different TBs / TUs in the list.

[0478] i. In one example, the assigned candidate indices can be encoded and decoded in a bitstream.

[0479] c. In one example, the motion candidate list can be constructed without using motion information of neighboring TBs / TUs within the current CU.

[0480] i. In one example, a list of motion candidates can be constructed by accessing motion information of spatially neighboring blocks (adjacent or non-adjacent) relative to the current CU.

[0481] 5. When an IBC / inter-frame codec block (CU) is divided into multiple TBs / TUs without signaling notification of the division information (e.g., the block size is larger than the maximum transform block size), the first TU / TB can be predicted by at least one reconstruction or prediction sample in the second TU / TB.

[0482] a. BV verification checks can be performed individually for each TU / TB.

[0483] 6. When an IBC / inter-frame codec block (CU) is divided into multiple TBs / TUs without signaling notification of the division information (e.g., the block size is larger than the maximum transform block size), the first TU / TB cannot be predicted by any reconstruction or prediction sample in the second TU / TB.

[0484] a. It can perform BV verification checks on the entire codec block.

[0485] Prediction block generation for IBC codec blocks

[0486] 7. It is proposed to use the filtered reconstructed reference samples for IBC prediction block generation.

[0487] a. In one example, filtered reconstructed reference samples can be generated by applying a deblocking filter / bilateral filter / SAO / ALF.

[0488] b. In one example, for an IBC codec block, at least one reference sample is filtered, and

[0489] At least one reference sample is from before filtering.

[0490] 8. Both filtered and unfiltered reconstructed reference samples can be used for IBC prediction block generation.

[0491] a. In one example, filtered reconstructed reference samples are generated by applying a deblocking filter / bilateral filter / SAO / ALF.

[0492] b. In one example, for the prediction of an IBC block, part of the prediction may come from filtered samples, while another part of the prediction may come from unfiltered samples.

[0493] 9. For IBC prediction block generation, whether to use the filtered or unfiltered value of the reconstructed reference sample depends on the location of the reconstructed sample.

[0494] a. In one example, the decision may depend on the relative position of the reconstructed sample point with respect to the current CTU and / or the CTU that covers the reference sample point.

[0495] b. In one example, if the reference sample point is outside the current CTU / CTB, or in a different CTU / CTB

[0496] If the sample is in the middle and has a large distance from the CTU boundary (e.g., 4 pixels), then the corresponding filtered sample can be used.

[0497] c. In one example, if the reference sample is within the current CTU / CTB, the unfiltered reconstructed sample can be used to generate the prediction block.

[0498] d. In one example, if the reference sample is outside the current VPDU, the corresponding filtered sample can be used.

[0499] e. In one example, if the reference sample is within the current VPDU, the unfiltered reconstructed sample can be used to generate the prediction block.

[0500] About MVD encoding and decoding

[0501] The motion vector (or block vector), or motion vector difference or block vector difference, is represented by (Vx, Vy).

[0502] 10. It is proposed to divide the absolute value of the motion vector component (e.g., Vx or Vy) represented by AbsV into two parts, where the first part is equal to AbsV – ((AbsV >> N) << N) (e.g., AbsV & (1 << N), the least significant N bits), and the other part is equal to (AbsV >> N) (e.g., the remaining most significant bits). Each part can be encoded and decoded separately.

[0503] a. In one example, each component of a motion vector or motion vector difference can be encoded and decoded separately.

[0504] i. In one example, the first part is encoded and decoded using fixed-length encoding, e.g., N bits.

[0505] 1) Alternatively, a first flag can be encoded and decoded to indicate whether the first part is equal to 0.

[0506] a. Alternatively, in addition, if not, the value of the first part minus 1 is encoded and decoded.

[0507] ii. In one example, the second part with MV / MVD flag information can be encoded and decoded using the current MVD encoding and decoding method.

[0508] b. In one example, the first part of each component of a motion vector or motion vector difference can be jointly encoded and decoded, while the second part can be encoded and decoded separately.

[0509] i. In one example, the first parts of Vx and Vy can be formed into a new positive value with 2N bits (e.g., (Vx << N) + Vy).

[0510] 1) In one example, a first flag can be encoded and decoded to indicate whether the new positive value is equal to 0.

[0511] a. Alternatively, in addition, if not, the new positive value minus 1 is encoded and decoded.

[0512] 2) Alternatively, fixed-length encoding or exp-golomb encoding (e.g.,

[0513] EG-0 th To encode and decode new positive values.

[0514] ii. In one example, the current MVD encoding / decoding method can be used to encode and decode the second part of each component of a motion vector with MV / MVD tagging information.

[0515] c. In one example, the first part of each component of a motion vector can be jointly encoded and decoded, and the second part can also be jointly encoded and decoded.

[0516] d. In one example, N is a positive value, such as 1 or 2.

[0517] e. In one example, N can depend on the MV precision used for motion vector storage.

[0518] i. In one example, if the precision of the MV used for motion vector storage is 1 / 2 M If the pixel is N, then N is set to M.

[0519] 1) In one example, if the precision of the MV used for motion vector storage is 1 / 16 pixel, then N equals 4.

[0520] 2) In one example, if the precision of the MV used for motion vector storage is 1 / 8 pixel, then N equals 3.

[0521] 3) In one example, if the precision of the MV used for motion vector storage is 1 / 4 pixel, then N equals 2.

[0522] f. In one example, N can depend on the MV of the current motion vector (e.g., the block vector in the IBC).

[0523] Precision.

[0524] i. In one example, if BV is 1 pixel precision, N can be set to 1.

[0525] Regarding SBT

[0526] 11. Whether the signaling notifies the maximum CU width and / or height allowed for subblock transformation may depend on the maximum allowed transformation block size.

[0527] a. In one example, whether the maximum CU width and height allowed for sub-block transformation is 64 or 32 luma samples (e.g., sps_sbt_max_size_64_flag) can depend on the value of the maximum allowed transform block size (e.g., sps_max_luma_transform_size_64_flag).

[0528] i. In one example, only when sps_sbt_enabled_flag and

[0529] The signaling is only sent to sps_sbt_max_size_64_flag when both sps_max_luma_transform_size_64_flag are true.

[0530] Regarding the LFNST matrix index

[0531] 12. Propose using two contexts to encode and decode LFNST / LFNST matrix indices (e.g., lfnst_idx), where the selection is entirely based on the use of the split tree.

[0532] g. In one example, one context is used for single-tree splitting and another context is used for dual-tree splitting, without considering the dependency on the main transformation.

[0533] 13. Propose the use of dual-tree and / or stripe type and / or color component to encode and decode LFNST / LFNST matrix indices (e.g., lfnst_idx).

[0534] 5. Additional Example Implementation

[0535] by Bold, underline, italic The changes are highlighted using the [[]] brackets. Deleted text is marked with [[]].

[0536] 5.1 Example #1

[0537] Table 9-82 shows how to assign ctxInc to syntax elements using a context-encoded container.

[0538]

[0539] 5.2 Example #2

[0540] 7.3.8.5 Encoding / Decoding Unit Syntax

[0541]

[0542] Figure 20This is a block diagram illustrating an example video processing system 2000, in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 2000. System 2000 may include an input 2002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 2002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0543] System 2000 may include codec component 2004, which can implement various codec or encoding methods described in this application. Codec component 2004 can reduce the average bit rate of the video from input 2002 to the output of codec component 2004 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 2004 can be stored or transmitted via connected communication, as shown in component 2006. The stored or communicated bitstream (or codec) representation of the video received at input 2002 can be used by component 2008 to generate pixel values ​​or to send displayable video to display interface 2010. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations for the reverse codec results will be performed by the decoder.

[0544] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The techniques described herein can be implemented in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0545] Figure 21This is a block diagram of a video processing device 2100. Device 2100 can be used to implement one or more methods described herein. Device 2100 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Device 2100 may include one or more processors 2102, one or more memories 2104, and video processing hardware 2106. Processor 2102 can be configured to implement one or more methods described in this application. Memory (multiple memories) 2104 can be used to store data and code for implementing the methods and techniques described herein. Video processing hardware 2106 can be used to implement some of the techniques described in this application in hardware circuitry.

[0546] In some embodiments, the following solutions may be implemented as preferred solutions.

[0547] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 1).

[0548] 1. A video processing method (e.g., Figure 22 The method described in the image (2200) includes: for a conversion between a codec representation of a video block in a video region and a video block based on an intra-block copying tool, determining (2202) that the video block can be divided into multiple transformation units, wherein the determination is based on the codec conditions of the video block and the codec representation omits signaling notification of the division; and performing (2204) the conversion based on the division.

[0549] 2. The method described in Solution 1, wherein the encoding / decoding conditions correspond to the dimensions of the video block.

[0550] 3. The method described in Solution 2, wherein the division is allowed if the width of the video block is greater than a first threshold, or if the height of the video block is greater than a second threshold.

[0551] 4. The method as described in Solution 3, wherein the first threshold and / or the second threshold are equal to the maximum transform size of the video region.

[0552] 5. The method as described in any one of solutions 1-4, wherein the partitioning is recursively applied to video blocks.

[0553] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 2).

[0554] 6. A video processing method, comprising: for the conversion between the codec representation of a video block of a video region and the video block, determining whether an intra-block copy (IBC) tool is enabled for the conversion of the video block based on the maximum transform size of the video region; and performing the conversion based on the determination.

[0555] 7. The method as described in Solution 6, wherein the video block is a chroma block, and wherein the maximum transform size corresponds to the size of the corresponding luma block, wherein IBC is disabled because the maximum transform size is greater than a threshold.

[0556] 8. The method as described in any one of solutions 6-7, wherein the syntax elements in the codec representation indicate whether the IBC tool is enabled because the dimension of the video block or the maximum transform size meets the standard.

[0557] 9. The method as described in any one of solutions 6-8, wherein, when determining that the IBC tool is enabled, the conversion further includes, based on a zero-forcing condition, forcing all residual blocks of the video block to zero.

[0558] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 3).

[0559] 10. A video processing method comprising: for a conversion between a codec representation of a video block of a video region and the video block, determining whether a signaling notification of an intra-block copy (IBC) tool for the conversion is included in the codec representation; and performing the conversion based on the determination, wherein the determination is based on the width and / or height of the video block and the maximum permissible IBC block size of the video region.

[0560] 11. The method as described in Solution 10, wherein signaling notifications are not included because the width is greater than a first threshold or the height is greater than a second threshold.

[0561] The following solutions can be implemented in conjunction with the additional techniques described in the items listed in the previous section (e.g., Item 4).

[0562] 12. A video processing method comprising: determining, for a conversion between a codec representation of a video block in a video region and a video block using an intra-block copy (IBC) tool, that the video block can be divided into multiple transform units (TUs) for conversion; and performing the conversion based on the determination, wherein the conversion includes using individual motion information on the multiple TUs.

[0563] 13. The method as described in Solution 12, wherein multiple TUs are constrained to share the same advanced motion vector predictor or Merge pattern flag.

[0564] 14. The method as described in solution 12, wherein, during the conversion, motion information of one of the plurality of TUs is determined without using motion information of another of the plurality of TUs.

[0565] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 5).

[0566] 15. The method of any one of solutions 12-14, wherein a first TU is predicted from a second TU among a plurality of TUs.

[0567] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 6).

[0568] 16. The method of any one of solutions 12-15, wherein the first TU among the plurality of TUs is predicted to be disabled from the second TU among the plurality of TUs.

[0569] The following solutions can be implemented in conjunction with the additional techniques described in the items listed in the previous section (e.g., Item 7).

[0570] 17. A video processing method comprising: for a conversion between a codec representation of a video block of a video region and the video block, determining that an intra-block copying tool is enabled for the conversion; and using the intra-block copying tool to perform the conversion, wherein the prediction of the video block is performed using filtered reconstructed samples of the video region.

[0571] 18. The method of solution 17, wherein the filtered reconstructed samples include samples generated by applying a loop filter to the reconstructed samples, wherein the loop filter is a deblocking filter, a bilateral filter, a sample adaptive offset, or an adaptive loop filter.

[0572] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 8).

[0573] 19. The method of any one of solutions 17-18, wherein, based on rules, the prediction also selectively uses unfiltered reconstructed samples of the video region.

[0574] The following solutions can be implemented in conjunction with the additional techniques described in the items listed in the previous section (e.g., Item 9).

[0575] 20. The method described in solution 19, wherein the rules depend on the location of the reconstructed samples.

[0576] 21. The method as described in Solution 20, wherein the position of the reconstructed sample for the rule is the relative position of the reconstructed sample with respect to the codec tree unit of the video block.

[0577] The following solutions can be implemented in conjunction with the additional techniques described in the projects listed in the previous section (e.g., Project 10).

[0578] 22. A video processing method, comprising: performing a conversion between a video including a plurality of video blocks and an encoded / decoded representation of the video, wherein at least some of the video blocks are encoded / decoded using motion vector information, and wherein the motion vector information is represented as a first part in the encoded / decoded representation based on the least significant bits of the absolute value of the motion vector information, and the motion vector information is represented as a second part in the encoded / decoded representation based on the remaining more significant bits more significant than the least significant bits.

[0579] 23. The method according to solution 22, wherein the absolute value of the motion vector V represented as AbsV is divided into a first part and a second part, the first part being equal to AbsV–((AbsV>>N)<<N) or AbsV&(1<<N), and the second part being equal to (AbsV>>N), where N is an integer.

[0580] 24. The method according to any one of solutions 22-23, wherein the first part and the second part are separately encoded / decoded in the encoded / decoded representation.

[0581] 25. The method according to any one of solutions 22-24, wherein the first part of the x motion vector and the first part of the y motion vector are jointly encoded / decoded at the same position.

[0582] 26. The method according to solution 25, wherein the second part of the x motion vector and the second part of the y motion vector are separately encoded / decoded.

[0583] 27. The method according to any one of solutions 22-26, wherein N is a function of the precision for storing the motion vector.

[0584] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., item 11).

[0585] 28. A video processing method, comprising: for a conversion between an encoded / decoded representation of a video block in a video region and the video block, determining whether a sub-block transform tool is enabled for the conversion; and performing the conversion based on the determination, wherein the determination is based on the maximum allowable transform block size of the video region, and wherein signaling is conditionally included in the encoded / decoded representation based on the maximum allowable transform block size.

[0586] The following solutions can be implemented together with the additional techniques described in the items listed in the previous section (e.g., items 12, 13).

[0587] 29. A video processing method comprising: for a conversion between a codec representation of a video block of a video region and the video block, determining whether to use a low-frequency inseparable transform (LFNST) during the conversion; and performing the conversion based on the determination, wherein the determination is based on codec conditions applied to the video block, and wherein two contexts are used to codec the matrix index of the LFNST in the codec representation.

[0588] 30. The method as described in Solution 29, wherein the encoding / decoding conditions include the segmentation tree used, the stripe type of the video block, or the color component identifier of the video block.

[0589] 31. The method of any one of solutions 1-30, wherein the video region includes an encoding / decoding unit, and wherein the video block corresponds to the luminance or chrominance component of the video region.

[0590] 32. The method of any one of solutions 1-31, wherein the transformation includes generating a bitstream representation from the current video block.

[0591] 33. The method of any one of solutions 1-31, wherein the transformation includes generating samples of the current video block from the bitstream representation.

[0592] 34. A video processing apparatus, including a processor configured to implement the method as described in any one or more of solutions 1-33.

[0593] 35. A computer-readable medium having code stored thereon that, when run, causes a processor to perform the methods described in any one or more of solutions 1-33.

[0594] 36. A method, system, or apparatus as described herein.

[0595] Figure 23 This is a flowchart representation of a video processing method 2300 according to the present technology. Method 2300 includes: in operation 2310, determining, for a conversion between a current block in a video region of a video and a codec representation of the video, that it is permissible to divide the current block into multiple transform units based on the characteristics of the current block. Signaling notification of the division is omitted in the codec representation. Method 2300 includes: in operation 2320, performing the conversion based on the determination.

[0596] In some embodiments, the characteristics of the current block include the dimensions of the current block. In some embodiments, partitioning is invoked for transformation if the width of the current block is greater than a first threshold. In some embodiments, partitioning is invoked for transformation if the height of the current block is greater than a second threshold. In some embodiments, partitioning the current block includes: recursively partitioning the current block until the dimension of one of a plurality of transform units is equal to or less than a threshold. In some embodiments, the width of the current block is recursively divided by half until the width of the current block is less than or equal to the first threshold. In some embodiments, the height of the current block is recursively divided by half until the height of the current block is less than or equal to the second threshold. In some embodiments, the current block is recursively partitioned until the width of the current block is less than or equal to the first threshold and the height of the current block is less than or equal to the second threshold. In some embodiments, the first threshold is equal to the maximum transform size of the video region. In some embodiments, the second threshold is equal to the maximum transform size of the video region.

[0597] In some embodiments, at least a subset of multiple transform units share the same motion information. In some embodiments, the same motion information includes syntax flags for an advanced motion vector prediction mode or a merge mode, the syntax flags being signaled once in the codec representation of the current block. In some embodiments, the motion information of the current block is determined multiple times in the codec representation. In some embodiments, a motion candidate list is constructed once for the current block, and different candidates are assigned to multiple transform units in the motion candidate list. In some embodiments, the indices of the different candidates are signaled in the codec representation. In some embodiments, for transform units of multiple transform units, a motion candidate list is constructed without using motion information of neighboring transform units within the current block. In some embodiments, a motion candidate list is constructed using motion information of spatially neighboring blocks of the current block. In some embodiments, a first transform unit is reconstructed based on at least one reconstruction sample in a second transform unit. In some embodiments, a verification check is performed individually for each of the multiple transform units. In some embodiments, a first transform unit is predicted without using any reconstruction samples from other transform units. In some embodiments, a verification check is performed on the block after reconstruction.

[0598] Figure 24 This is a flowchart representation of a video processing method 2400 according to the present technology. Method 2400 includes: in operation 2410, for a conversion between a current block in a video region of a video and a codec representation of the video, determining predicted samples of the current block based on filtered reconstructed reference samples in the video region using an intra-block copy (IBC) model. Method 2400 includes: in operation 2420, performing the conversion based on the determination.

[0599] In some embodiments, filtered reconstructed reference samples are generated based on filtering operations, including at least one of the following: deblocking filtering, bilateral filtering, Sample Adaptive Offset (SAO) filtering, or adaptive loop filtering. In some embodiments, determining the prediction samples of the current block using the IBC model is also based on unfiltered reconstructed reference samples in the video region. In some embodiments, prediction samples in a first portion of the current block are determined based on filtered reconstructed reference samples, and prediction samples in a second portion of the current block are determined based on unfiltered reconstructed reference samples. In some embodiments, whether to use filtered or unfiltered reconstructed reference samples to determine the prediction samples of the block is based on the position of the prediction sample relative to the current codec tree unit or relative to the codec tree unit covering the sample. In some embodiments, when the prediction sample is located outside the current codec tree unit, the filtered reconstructed reference samples are used to reconstruct the prediction sample. In some embodiments, when the prediction sample is located inside the codec tree unit and there is a distance D equal to or greater than 4 pixels between the prediction sample and one or more boundaries of the codec tree unit, the filtered reconstructed reference samples are used to reconstruct the prediction sample. In some embodiments, when the predicted sample is located inside the current codec tree unit, the predicted sample is reconstructed using unfiltered reconstruction reference samples. In some embodiments, when the predicted sample is located outside the current virtual pipeline data unit, the predicted sample is reconstructed using filtered reconstruction reference samples. In some embodiments, when the predicted sample is located inside the current virtual pipeline data unit, the predicted sample is reconstructed using unfiltered reconstruction reference samples.

[0600] Figure 25 This is a flowchart representation of a video processing method 2500 according to the present technology. Method 2500 includes: in operation 2510, for the conversion between the current block of the video and the codec representation of the video, determining, according to rules, whether a syntax element indicating a skip mode using an Intra-Block Copy (IBC) codec model is included in the codec representation. The rules specify that the signaling notification of the syntax element is based on the dimension of the current block and / or based on the maximum allowed dimension of the block encoded using the IBC codec model. Method 2500 also includes: in operation 2520, performing the conversion based on the determination.

[0601] In some embodiments, the rule specifies that syntax elements are omitted from the codec representation if the width of the current block is greater than the maximum allowed width. In some embodiments, the rule specifies that syntax elements are omitted from the codec representation if the height of the current block is greater than the maximum allowed height. In some embodiments, the rule specifies that syntax elements are included in the codec representation if the width of the current block is less than or equal to the maximum allowed width. In some embodiments, the rule specifies that syntax elements are included in the codec representation if the height of the current block is less than or equal to the maximum allowed height.

[0602] In some embodiments, the current block is within a video region of the video, and the rules apply to situations where the video region includes I-pieces, I-strips, I-tiles, or I-sub-pictures. In some embodiments, the maximum allowed width or maximum allowed height is 64. In some embodiments, the maximum allowed width or maximum allowed height is equal to the dimension of the virtual pipeline data unit. In some embodiments, the maximum allowed width or maximum allowed height is equal to the maximum dimension of the transform unit.

[0603] Figure 26 This is a flowchart representation of a video processing method 2600 according to the present technology. Method 2600 includes: in operation 2610, for the conversion between a current block of video and a codec representation of the video, determining at least one context for associating an index with a Low Frequency Inseparable Transform (LFNST) codec model. The LFNST codec model includes: applying a positive quadratic transform between a positive principal transform and a quantization step during encoding; or applying an inverse quadratic transform between a dequantization step and an inverse principal transform during decoding. The size of the positive quadratic transform and the inverse quadratic transform is smaller than the size of the current block. At least one context is determined based on the segmentation type of the current block without considering the positive principal transform or the inverse principal transform. Method 2600 also includes: in operation 2620, performing a conversion based on the determination.

[0604] In some embodiments, when the partition type is single-tree partitioning, only one context is used for encoding and decoding the index. In some embodiments, when the partition type is dual-tree partitioning, two contexts are used for encoding and decoding the index.

[0605] In some embodiments, an index indicating the use of a Low Frequency Indivisible Transform (LFNST) codec model is included in the codec representation based on characteristics associated with the block, including the segmentation type, stripe type, or color component associated with the current block.

[0606] Figure 27This is a flowchart representation of a video processing method 2700 according to the present technology. Method 2700 includes: in operation 2710, a conversion between a current block of a video region and the codec representation of the video, determining whether to enable the Intra-Block Copy (IBC) codec model based on the maximum transform unit size applicable to the video region. Method 2700 further includes: in operation 2720, performing the conversion based on the determination.

[0607] In some embodiments, the IBC codec model is disabled if the dimension of the current block or the dimension of the luma block corresponding to the current block is greater than the maximum transform unit size. In some embodiments, the IBC codec model is disabled if the width of the current block or the width of the luma block corresponding to the current block is greater than the maximum transform unit width. In some embodiments, the IBC codec model is disabled if the height of the current block or the height of the luma block corresponding to the current block is greater than the maximum transform height. In some embodiments, the signaling notification of the use of the IBC codec model in the codec representation is based on the dimension of the current block and the maximum transform unit size. In some embodiments, the signaling notification of the use of the IBC codec model includes a syntax element indicating a skip mode for the IBC codec model. In some embodiments, the syntax element is included in an I-slice, I-strip, I-tile, or I-subpicture in the codec representation.

[0608] In some embodiments, signaling notification of the use of the IBC codec model includes syntax elements indicating the IBC mode. In some embodiments, if the width of the current block is greater than the maximum transform unit width or the height of the current block is greater than the maximum transform unit height, the syntax element is omitted in the codec representation. In some embodiments, if the width of the current block is less than or equal to the maximum transform unit width or the height of the current block is less than or equal to the maximum transform unit height, the syntax element is included in the codec representation and set to 0. In some embodiments, if the width of the current block is greater than the maximum transform unit width or the height of the current block is greater than the maximum transform unit height, the syntax element is included in the codec representation and set to 0.

[0609] In some embodiments, when the IBC encoding / decoding model is enabled for the current block, according to the rule, the samples in the residual block corresponding to the current block are set to 0. In some embodiments, the rule stipulates that: when the width of the current block is greater than the maximum transform unit width or the height of the current block is greater than the maximum transform unit height, the samples are set to 0. In some embodiments, the rule stipulates that: when the width of the current block is greater than the first threshold or the height of the current block is greater than the second threshold, the samples are set to 0. In some embodiments, the first threshold is 64 / vSizeX, where vSizeX is the width of the virtual pipeline data unit. In some embodiments, the second threshold is 64 / vSizeY, where vSizeY is the height of the virtual pipeline data unit. In some embodiments, the syntax flag of the encoding / decoding block is omitted in the encoding / decoding representation. In some embodiments, the syntax flag of the encoding / decoding block is set to 0 in the encoding / decoding representation.

[0610] In some embodiments, in the encoding / decoding representation, signaling that the syntax flag for the encoding / decoding block is based on the use of the IBC encoding / decoding model. In some embodiments, when the IBC encoding / decoding model is enabled for the current block and the dimension of the block is greater than the maximum transform unit size, the syntax flag is omitted in the encoding / decoding representation. In some embodiments, when the IBC encoding / decoding model is enabled for the current block and the dimension of the block is greater than the threshold associated with the virtual pipeline data unit, the syntax flag is omitted in the encoding / decoding representation.

[0611] Figure 28 is a flowchart representation of a video processing method 2800 according to the present technology. Method 2800 includes, at operation 2810, for the conversion between the current block of a video and the encoding / decoding representation of the video, determining that the absolute value of the component of the motion vector of the current block is divided into two parts. The motion vector is represented as (Vx, Vy), the component is represented as Vi, and Vi is Vx or Vy. The first part of the two parts is equal to |Vi|-((|Vi|>>N)<<N), the second part of the two parts is equal to |Vi|>>N, and N is a positive integer. The two parts are separately encoded / decoded in the encoding / decoding representation. Method 2800 further includes: at operation 2820, performing the conversion based on the determination.

[0612] In some embodiments, two components Vx and Vy are signaled separately in the codec representation. In some embodiments, the first part is coded in a fixed length of N bits. In some embodiments, the first parts of component Vx and component Vy are jointly coded, and the second parts of component Vx and component Vy are coded separately. In some embodiments, the first parts of component Vx and component Vy are coded into a value with a length of 2N bits. In some embodiments, the value is equal to (Vx << N) + Vy. In some embodiments, a fixed length coding process or an exp - golomb coding process is used to code the value.

[0613] In some embodiments, the first parts of component Vx and component Vy are jointly coded, and the second parts of component Vx and component Vy are jointly coded. In some embodiments, a syntax flag is included in the codec representation to indicate whether the first part is equal to 0. In some embodiments, the first part has a value K, where K ≠ 0, and the value (K - 1) is coded in the codec representation. In some embodiments, a motion vector difference coding process is used to code the second part of each component with flag information. In some embodiments, N is 1 or 2. In some embodiments, N is determined based on the motion vector precision for motion vector data storage. In some embodiments, when the motion vector precision is 1 / 16 pixel, N is 4. In some embodiments, when the motion vector precision is 1 / 8 pixel, N is 3. In some embodiments, when the motion vector precision is 1 / 4 pixel, N is 2. In some embodiments, N is determined based on the motion vector precision of the current motion vector. In some embodiments, when the motion vector precision is 1 pixel, N is 1.

[0614] Figure 29 is a flowchart representation of a video processing method 2900 according to the present technology. Method 2900 includes, at operation 2910, for the conversion between the current block of a video and the codec representation of the video, determining information about the maximum dimension of the current block that allows sub - block transformation of the current block based on the maximum allowed dimension of the transform block. Method 2900 further includes: at operation 2920, performing the conversion based on the determination.

[0615] In some embodiments, the maximum dimension of the current block that allows sub - block transformation of the current block corresponds to the maximum allowed dimension of the transform block. In some embodiments, the maximum dimension of the current block is 64 or 32. In some embodiments, when a first syntax flag in the slice parameter set indicates enabling sub - block transformation and a second syntax flag in the slice parameter set indicates that the maximum allowed dimension of the transform block is 64, a syntax flag indicating the maximum dimension of the current block is included in the codec representation.

[0616] In some embodiments, the transformation includes generating a codec representation from the current block. In some embodiments, the transformation includes generating samples of the current block from the codec representation.

[0617] In the above solutions, performing a conversion involves using the results of previously determined steps (e.g., using or not using certain encoding / decoding steps) during the encoding or decoding operation to obtain the conversion result.

[0618] Some embodiments of the disclosed technology include making a decision or determination to enable a video processing tool or mode. In one example, when a video processing tool or mode is enabled, the encoder will use or implement the tool or mode in the processing of video blocks, but may not necessarily modify the resulting bitstream based on the use of the tool or mode. That is, when a video processing tool or mode is enabled based on a decision or determination, the conversion from video blocks to a bitstream representation of the video will be performed using that video processing tool or mode. In another example, when a video processing tool or mode is enabled, the decoder will utilize knowledge that the bitstream has already been modified based on the video processing tool or mode to process the bitstream. That is, the conversion from the bitstream representation of the video to video blocks will be performed using the video processing tool or mode enabled based on a decision or determination.

[0619] Some embodiments of the disclosed technology include making a decision or determination to disable a video processing tool or mode. In one example, when a video processing tool or mode is disabled, the encoder does not use the tool or mode in the conversion between the video block-to-video bitstream representation. In another example, when a video processing tool or mode is disabled, the decoder processes the bitstream using the knowledge that the bitstream has not been modified using a video processing tool or mode that was enabled based on the decision or determination.

[0620] The disclosures and other solutions, examples, embodiments, modules, and functional operations described in this application can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed in this application and their structural equivalents, or combinations thereof. The embodiments disclosed in this application, and other embodiments, can be implemented as one or more computer program products, such as one or more modules of computer program instructions encoded on a tangible and non-volatile computer-readable medium for execution by a data processing device or for controlling the operation of the data processing device. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that influences machine-readable propagated signals, or a combination thereof. The term "data processing device" includes all devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computer groups. In addition to hardware, the device may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0621] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to that program, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed and executed on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.

[0622] The processing and logic flows described in this application can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuits, and the device can also be implemented as special-purpose logic circuits, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0623] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as one or more of any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by, or merged into, special-purpose logic circuitry.

[0624] While this patent document contains numerous details, it should not be construed as limiting the scope of any invention or claim, but rather as a description of features of specific embodiments of a particular invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in certain circumstances, one or more features from a combination of claims may be removed from the combination, and a combination of claims may refer to a sub-combination or a variation of a sub-combination.

[0625] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence shown to perform such operations, or all the described operations, in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0626] Only some implementations and examples are described; other implementations, enhancements, and variations can be made based on the content described and illustrated in this patent document.

Claims

1. A method for processing video data, comprising: For the first conversion between the first video block of the video and the bitstream of the video, determine whether the width of the first video block is greater than a first threshold, or whether the height of the first video block is greater than a second threshold; as well as When the width of the first video block is greater than the first threshold or the height of the first video block is greater than the second threshold, the segmentation operation of the first video block is invoked to obtain two or more transform blocks. Based on the determination, the first conversion is performed. In this context, the signaling for the call and the signaling for the segmentation operation are omitted from the bitstream. Wherein, the first threshold and the second threshold are equal to a first value, the first value specifying the maximum transformation size in the luminance sample, and The method further includes: Determine whether the bitstream contains the third syntax element of the first video block. The third syntax element specifies whether to use a Low Frequency Inseparable Transform (LFNST) kernel from the selected transform set, and which LFNST kernel from the selected transform set to use. Specifically, a single context selected from contexts with an index increment of 0 or 1 is used for the first binary bit of the third syntax element, and The index increment is determined solely based on whether the tree type of the first video block is a single tree.

2. The method according to claim 1, further comprising: For the second conversion between the second video block of the video and the bitstream of the video, it is determined whether the first syntax element is included in the bitstream according to the first rule; as well as At least based on the determination, the second conversion is performed. The first syntax element indicates whether the prediction mode is applied to the second video block. In the prediction mode, predicted samples are derived based on blocks of sample values ​​for the same video region determined by block vectors. The first rule stipulates that when the width of the second video block is greater than the third threshold and the height of the second video block is greater than the fourth threshold, the first syntax element is not included in the bitstream.

3. The method according to claim 2, further comprising: The second rule determines whether the second syntax element indicating whether to apply the skip mode to the second video block is included in the bitstream. The second rule specifies that when the second video block is not a tree-type chroma block with a dual tree, the sequence enable flag indicates that the prediction mode is enabled, the width of the second video block is less than or equal to the third threshold, and the height of the second video block is less than or equal to the fourth threshold, the second syntax element is included in the bitstream.

4. The method according to claim 3, wherein, The third threshold and the fourth threshold are 64, or when the size of the codec tree block is greater than 64, the third threshold and the fourth threshold are equal to the size of the virtual unit.

5. The method according to claim 1, wherein, In response to the first video block having a tree type of the single tree, the index increment is 0, and in response to the first video block having a tree type of the single tree, the index increment is 1.

6. The method according to claim 1, wherein, The third syntax element is associated with the LFNST encoding / decoding model. The LFNST encoding / decoding model includes applying a forward quadratic transform between the forward master transform and the quantization step during encoding, or applying an inverse quadratic transform between the inverse quantization step and the inverse master transform during decoding. The index increment is determined without referring to the forward master transform or the reverse master transform.

7. The method according to claim 1, wherein, In the segmentation operation, it is also determined whether to segment the width of the first video block based on the height of the first video block, and it is also determined whether to segment the height of the first video block based on the width of the first video block.

8. The method according to claim 7, wherein, In the segmentation operation, if the width of the first video block is greater than the first value and the height of the first video block, the width of the first video block is halved.

9. The method according to claim 8, wherein, In the segmentation operation, if the width of the first video block is greater than the first value and the height of the first video block is greater than the height of the first video block, the height of the first video block remains unchanged.

10. The method according to claim 7, wherein, In the segmentation operation, if the height of the first video block is greater than the first value, and the width of the first video block is less than or equal to the first value or the height of the first video block, the height of the first video block is halved.

11. The method according to claim 10, wherein, In the segmentation operation, if the height of the first video block is greater than the first value, and the width of the first video block is less than or equal to the first value or the height of the first video block, the width of the first video block remains unchanged.

12. The method according to claim 1, wherein, The segmentation operation is called recursively until the width of the first video block is less than or equal to the first threshold and the height is less than or equal to the second threshold.

13. The method according to claim 1, wherein, When the first value equals 32, and the first video block is derived based on the codec block of the prediction mode, the segmentation operation is invoked. In the prediction mode, predicted samples are derived based on blocks of sample values ​​from the same video region determined by the block vector. Wherein, the width and height of the codec block are less than or equal to 64 and greater than 32.

14. The method according to claim 1, wherein, The first conversion includes encoding the video into the bitstream.

15. The method according to claim 1, wherein, The first conversion includes decoding the video from the bitstream.

16. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor: For the first conversion between the first video block of the video and the bitstream of the video, determine whether the width of the first video block is greater than a first threshold, or whether the height of the first video block is greater than a second threshold; as well as When the width of the first video block is greater than the first threshold or the height of the first video block is greater than the second threshold, the segmentation operation of the first video block is invoked to obtain two or more transform blocks. Based on the determination, the first conversion is performed. In this context, the signaling for the call and the signaling for the segmentation operation are omitted from the bitstream. Wherein, the first threshold and the second threshold are equal to a first value, the first value specifying the maximum transformation size in the luminance sample, and When the instruction is executed by the processor, it also causes the processor to: Determine whether the bitstream contains the third syntax element of the first video block. The third syntax element specifies whether to use a Low Frequency Inseparable Transform (LFNST) kernel from the selected transform set, and which LFNST kernel from the selected transform set to use. Specifically, a single context selected from contexts with an index increment of 0 or 1 is used for the first binary bit of the third syntax element, and The index increment is determined solely based on whether the tree type of the first video block is a single tree.

17. A non-transitory computer-readable storage medium storing instructions that cause a processor to: For the first conversion between the first video block of the video and the bitstream of the video, determine whether the width of the first video block is greater than a first threshold, or whether the height of the first video block is greater than a second threshold; as well as When the width of the first video block is greater than the first threshold or the height of the first video block is greater than the second threshold, the segmentation operation of the first video block is invoked to obtain two or more transform blocks. Based on the determination, the first conversion is performed. In this context, the signaling for the call and the signaling for the segmentation operation are omitted from the bitstream. Wherein, the first threshold and the second threshold are equal to a first value, the first value specifying the maximum transformation size in the luminance sample, and The instruction also causes the processor to: Determine whether the bitstream contains the third syntax element of the first video block. The third syntax element specifies whether to use a Low Frequency Inseparable Transform (LFNST) kernel from the selected transform set, and which LFNST kernel from the selected transform set to use. Specifically, a single context selected from contexts with an index increment of 0 or 1 is used for the first binary bit of the third syntax element, and The index increment is determined solely based on whether the tree type of the first video block is a single tree.

18. A non-transitory computer-readable recording medium storing a program and a bit stream, said program being implemented when executed by a processor: For the first video block, determine whether the width of the first video block is greater than a first threshold, or whether the height of the first video block is greater than a second threshold; and When the width of the first video block is greater than the first threshold or the height of the first video block is greater than the second threshold, the segmentation operation of the first video block is invoked to obtain two or more transform blocks. The bit stream is generated based on the determination. in, The signaling for the call and the signaling for the segmentation operation are omitted from the bitstream. Wherein, the first threshold and the second threshold are equal to a first value, the first value specifying the maximum transformation size in the luminance sample, and When the program is executed by the processor, it also implements: Determine whether the bitstream contains the third syntax element of the first video block. The third syntax element specifies whether to use a Low Frequency Inseparable Transform (LFNST) kernel from the selected transform set, and which LFNST kernel from the selected transform set to use. Specifically, a single context selected from contexts with an index increment of 0 or 1 is used for the first binary bit of the third syntax element, and The index increment is determined solely based on whether the tree type of the first video block is a single tree.

19. A method for storing a video bitstream, comprising: For the first video block of the video, determine whether the width of the first video block is greater than a first threshold, or whether the height of the first video block is greater than a second threshold. as well as When the width of the first video block is greater than the first threshold or the height of the first video block is greater than the second threshold, the segmentation operation of the first video block is invoked to obtain two or more transform blocks. The bit stream is generated based on the determination; as well as The bitstream is stored in a non-transitory computer-readable storage medium. In this context, the signaling for the call and the signaling for the segmentation operation are omitted from the bitstream. Wherein, the first threshold and the second threshold are equal to a first value, the first value specifying the maximum transformation size in the luminance sample, and The method further includes: Determine whether the bitstream contains the third syntax element of the first video block. The third syntax element specifies whether to use a Low Frequency Inseparable Transform (LFNST) kernel from the selected transform set, and which LFNST kernel from the selected transform set to use. Specifically, a single context selected from contexts with an index increment of 0 or 1 is used for the first binary bit of the third syntax element, and The index increment is determined solely based on whether the tree type of the first video block is a single tree.

Citation Information

Patent Citations

  • Recursive partitioning of video coding blocks

    CN114365488A