Coding and decoding intra-block copies

By employing advanced block splitting and transformation techniques, including IBC and LFNST, the methods address bandwidth challenges in video coding, achieving improved compression and reduced bitrate.

JP7791295B2Active Publication Date: 2025-12-23DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024206541
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-09
Filing Date
2024-11-27
Publication Date
2025-12-23
Estimated Expiration
2040-09-09

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in efficiently compressing digital video data due to increasing bandwidth demands, particularly in the context of intra-block copy tools, which require improved methods for block splitting, transformation, and motion vector handling.

Method used

The proposed methods involve splitting blocks into multiple transform units based on block characteristics, using intra-block copy (IBC) models, low-frequency non-separable transforms (LFNST), and advanced motion vector representations to enhance video coding efficiency.

Benefits of technology

These methods improve video coding efficiency by reducing bitrate requirements and enhancing compression performance, aligning with emerging standards like VVC, while maintaining computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007791295000016
    Figure 0007791295000016
  • Figure 0007791295000017
    Figure 0007791295000017
  • Figure 0007791295000018
    Figure 0007791295000018
Patent Text Reader

Abstract

To provide devices, systems and methods related to video and image coding and decoding in which an intra block copy tool is used for coding or decoding.SOLUTION: A method of video processing includes determining, for a conversion between a current block of a video and a coded representation of the video, whether a syntax element indicating usage of a skip mode for an intra-block copy (IBC) coding model is included in the coded representation according to a rule that specifies that signaling of the syntax element is based on a dimension of the current block and / or a maximum allowed dimension for a block that is coded using the IBC coding model. The method also includes performing the conversion based on the determination.SELECTED DRAWING: Figure 25
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS Under the applicable patent laws and / or regulations under the Paris Convention, this application Priority and interest of International Patent Application No. PCT / CN2019 / 104869, filed September 9, 2019 The purpose of this application is to timely assert the benefits of the invention. The entire disclosure of which is incorporated by reference as part of the present disclosure.

[0002] This patent document relates to video and image coding and decoding. [Background technology]

[0003] Despite advances in video compression, digital video is still used extensively across the internet and It accounts for the largest bandwidth usage of any other digital communications network. As the number of connected user devices capable of displaying and displaying digital content increases, Bandwidth demands for video usage are expected to continue to increase. Summary of the Invention

[0004] Regarding digital video coding, in particular, coding an intra-block copy tool Video and image coding and decoding devices for use in encoding or decoding , systems and methods.

[0005] In one exemplary embodiment, a method of image processing is disclosed, the method comprising: For conversion between the current block in the domain and the coding representation of the image, Allows splitting of blocks into multiple transform units based on the characteristics of the current block. In the coding representation, the signaling of the split is omitted. The method also includes performing a transformation based on the determination.

[0006] In another exemplary aspect, a method of image processing is disclosed, the method comprising: for transforming a current block in the image domain to a coding representation of the image, Intra-block copy (IBC) model based on filtered reconstructed reference samples The method includes determining a predicted sample for the current block using the code. This also includes performing a conversion based on the

[0007] In another exemplary aspect, a method of video processing is disclosed, the method comprising: For the conversion between blocks and this video coding representation, the intra-block This syntax element indicates the use of the Skip Mode for the Interleaved Block Copy (IBC) coding model. This rule determines whether the syntax element signaling the current block dimensions and / or the IBC coding model. The method specifies that the maximum allowable size for a block to be coded is based on: It also includes performing a conversion based on the determination.

[0008] In another exemplary aspect, a method of video processing is disclosed, the method comprising: Low Frequency Non-Separable Transform (LFNST) for conversion between block and video coding representations At least one coding model for coding the associated index This LFNST coding model involves determining one context. applying a forward quadratic transform between the forward linear transform and the quantization step during the quantization; During decoding, the inverse secondary transform is applied between the inverse quantization step and the inverse primary transform. The size of the forward and inverse secondary transforms is greater than the size of the current block. The at least one context is a forward linear transform or an inverse linear transform. The method is based on the partition type of the current block, without considering the partition conversion. It also includes performing a conversion according to the determination.

[0009] In another exemplary aspect, a method of image processing is disclosed, the method comprising: A coding representation of the image is converted into a coding representation of the image. Based on the maximum transform unit size, the Intra Block Copy (IBC) coding model The method includes determining whether the controller is enabled. The method performs the conversion according to the determination. This also includes the following.

[0010] In another exemplary aspect, a method of video processing is disclosed. The method comprises: For the conversion between the block and the image and the coding representation of this image, the motion vector for the current block is The motion vector is calculated by dividing the absolute values ​​of the components of the vector into two parts. The voltage is expressed as (Vx, Vy), and this component is expressed as Vi, where Vi is either Vx or Vy. The first of the two parts is |Vi|-((|Vi|>>N)< <N)に and the second of the two parts is equal to |Vi|>>N,N, where N is a positive integer. The two parts are coded separately in the coding representation. This also includes performing conversion by

[0011] In another exemplary aspect, a method of video processing is disclosed. The method comprises: For conversion between image and video, the maximum allowable size of the conversion block is used to determine the current block. Determine information about the maximum dimensions of the current block that allows sub-block transformations in The method also includes performing a transformation according to the determination.

[0012] In another exemplary aspect, a method of video processing is disclosed. Coding Representation of Video Blocks in the Video Domain Based on a Block Copy Tool and Video Block Allows splitting of video blocks into multiple transform units for transformation with blocks determining whether or not the video block is encoded, the determination being made based on a coding condition of the video block; ,The coding representation omits signaling of the split, ,determining and transforming based on the split. and performing the steps of:

[0013] In another exemplary aspect, a method of video processing is disclosed. The method comprises: For the conversion between the coding representation of the image block and the video block, the maximum conversion size in the video domain is Intra Block Copy (IBC) tool for converting video blocks based on size determining whether the parameter is enabled and performing the conversion based on the determination. nothing. In another exemplary aspect, a method of video processing is disclosed. The method comprises: For the conversion between the coding representation of the image block and the video block, an interface for this conversion is provided. Whether the signaling of the Interactive Block Copy (IBC) tool is included in the coding expression determining whether the image is a The width and / or height of the block and the maximum allowable IBC block size for the video area. and performing a transformation based on the

[0014] In another exemplary aspect, a method of video processing is disclosed. The Interleaved Block Copy (IBC) tool is used to create a coding representation of the video blocks in the video domain. Divide a video block into multiple transform units (TUs) for transformation to and from the video block. determining whether to allow the conversion to a plurality of TUs; and and performing a transformation, including using separate motion information for the image.

[0015] In another exemplary aspect, a method of video processing is disclosed. The method comprises: Intra-block copying is used to convert between image block coding representations and video blocks. determining that the conversion tool is enabled for intra-block copying; and performing a transformation using a tool to generate a filtered reconstruction of the image region. The samples are used to make a prediction for the video block.

[0016] In another exemplary aspect, a method of video processing is disclosed. The method comprises: It includes a video containing a lock and a coding representation and conversion of this video, and at least Some video blocks are coded using motion vector information, and the motion vectors The information is a first portion based on the first least significant bit of the absolute value of the motion vector information, and a second part based on the remaining more significant bits that are more significant than the first least significant bit. and is expressed in coding representation.

[0017] In another exemplary aspect, a method of video processing is disclosed. The method comprises: Sub-block transformation tools are used to transform between image block coding representations and video blocks. Determines whether a rule is enabled for conversion and performs the conversion based on the determination. and determining based on a maximum allowable transform block size for the video domain. and conditionally signal to this coding representation based on the maximum allowed transform block size. Include knowledge.

[0018] In another exemplary aspect, a method of video processing is disclosed. The method comprises: For the conversion between the coding representation of the image block and the video block, a low frequency non-separable transform (L FNST) is used during the conversion and the conversion is performed based on this determination. and determining whether the video block is a video block that is a plurality of video blocks, the determining being based on a coding condition applied to the video block. The matrix index for LFNST and LFNST is calculated by It is coded in a coding representation using the

[0019] In yet another exemplary embodiment, the method is provided in the form of processor executable code. The method is embodied in and stored on a computer-readable program medium.

[0020] In yet another exemplary embodiment, a device configured or operable to perform the method described above. A device capable of performing the method is disclosed. The device is programmed to implement the method. A processing unit may be included.

[0021] In yet another exemplary aspect, a video decoder device includes: The method may be implemented.

[0022] These and other aspects and features of the disclosed technology are set forth in the drawings, description and claims. is explained in more detail. [Brief explanation of the drawings]

[0023] [Figure 1] 10 illustrates an example of a derivation process for building a merge candidate list. [Figure 2] 10 shows examples of spatial merge candidate locations. [Figure 3] 10 shows examples of candidate pairs that are considered for redundancy check of spatial merge candidates. [Figure 4] Examples of locations for the second PU for Nx2N and 2NxN divisions are shown. [Figure 5] FIG. 10 is an illustration of motion vector scaling for temporal merge candidates. [Figure 6] 1 shows example candidate positions for temporal merge candidates, C0 and C1. [Figure 7] 1 illustrates an example of a combined bidirectional prediction merge candidate. [Figure 8] 10 shows an example of a process for deriving motion vector prediction candidates. [Figure 9] 10 illustrates a description of motion vector scaling for spatial motion vector candidates. [Figure 10] 10 shows examples of candidate positions for affine merge mode. [Figure 11] 10 shows a modified merge list construction process. [Figure 12] 1 illustrates an example of triangulation-based inter prediction. [Figure 13] 10 illustrates an example of a central processing unit applying a first group of weighting factors. [Figure 14] 1 shows an example of a motion vector storage device. [Figure 15] An example of UMVE search processing is shown below. [Figure 16] An example of UMVE search points is shown below. [Figure 17] FIG. 10 illustrates the operation of an intra block copy tool. [Figure 18]An example of a quadratic transformation in JEM is shown below. [Figure 19] An example of a reduced quadratic transformation (RST) is shown below. [Figure 20] FIG. 1 is a block diagram illustrating an exemplary video processing system in which the disclosed techniques can be implemented. [Figure 21] FIG. 1 is a block diagram illustrating an example of a video processing device. [Figure 22] 1 is a flowchart illustrating an example of a video processing method. [Figure 23] 1 is a flowchart illustrating a video processing method according to the present technology. [Figure 24] 10 is a flowchart illustrating another video processing method according to the present technology. [Figure 25] 10 is a flowchart illustrating another video processing method according to the present technology. [Figure 26] 10 is a flowchart illustrating another video processing method according to the present technology. [Figure 27] 10 is a flowchart illustrating another video processing method according to the present technology. [Figure 28] 10 is a flowchart illustrating another video processing method according to the present technology. [Figure 29] 10 is a flowchart illustrating yet another video processing method according to the present technology. DETAILED DESCRIPTION OF THE INVENTION

[0024] Embodiments of the disclosed technology utilize existing video coding techniques to improve compression performance. This specification may be applied to standards (e.g., HEVC, H.265) and future standards. uses chapter headings to improve the readability of the description, and This does not limit the scope (and / or implementation) to each chapter.

[0025] 1. Summary of the invention

[0026] This specification relates to video coding technology. Specifically, intra block copy ( IBC (also known as Current Picture Reference, CPR) coding. It may be applied to existing video coding standards or may be applied to standards (Versatile Video The present invention may be applied to determine future video coding. It can also be applied to any video standard or video codec.

[0027] 2. Background technology

[0028] Video coding standards are primarily developed through the well-known ITU-T and ISO / IEC standards. ITU-T created H.261 and H.263, and ISO / IEC developed MP EG-1 and MPEG-4 Visual, and both organizations are H.262 / MPEG-2 V ideo and H.264 / MPEG-4 AVC (Advanced Video Cod) ing) and co-created the H.265 / HEVC standard. The standard uses a hybrid video coding structure that utilizes temporal prediction and transform coding. In 2015, to explore future video coding technologies beyond HEVC, is a joint project of VCEG and MPEG called JVET (Joint Video Exploration Since then, many new methods have been adopted by JVET. The reference software called JEM (Joint Exploration Mode) In April 2018, VCEG (Q6 / 16) and ISO / IE C JTC1 SC29 / WG11(MPEG) Joint Video Exp The Joint Virtuoso Team (JVET) was launched, achieving a 50% bitrate reduction compared to HEVC. We are working to develop VVC standards with this goal in mind.

[0029] 2.1 Inter Prediction in HEVC / H.265

[0030] For inter-coding coding units (CUs), the partition model Depending on the code, it may be coded with one prediction unit (PU) or two PUs. Each inter-predicted PU has motion parameters for one or two reference picture lists. The motion parameters include a motion vector and a reference picture index. The use of one of the two reference picture lists is inter_ The motion vector may be signaled using pred_idc. It may also be explicitly coded as a delta.

[0031] If one CU is coded in skip mode, one PU is responsible for this CU. The motion vector differentials are also coded based on the reference picture. There is no merge mode, which specifies the motion parameter for the current PU. The parameters are obtained from neighboring PUs, including spatial and temporal candidates. Can be applied to any inter-predicted PU, not just for skip mode An alternative to the merge mode is to explicitly send the motion parameters, and the motion vectors ( More precisely, the motion vector difference (MVD) compared to the motion vector predictor, The corresponding reference picture index in the reference picture list, and the usage status of the reference picture list. Such a mode is referred to as advanced motion vector mode in this disclosure. This is called AMVP.

[0032] If the signaling indicates that one of two reference picture lists is to be used, one Generate PUs from blocks of samples, which is called "uni-prediction." P slices and Uniprediction is available for both B slices.

[0033] If the signaling indicates that both reference picture lists are to be used, the block of the two samples Generate PU from block. This is called "bidirectional prediction". Bidirectional prediction is only performed on B slices. is available.

[0034] The inter prediction modes defined in HEVC will be explained in detail below. The following describes the mode.

[0035] 2.1.1 Reference Picture List

[0036] In HEVC, the term inter prediction refers to prediction of a picture with references other than the current decoded picture. A prediction derived from data elements (e.g., sample values ​​or motion vectors) of the reference picture. Similar to H.264 / AVC, it is used to indicate a single frame from multiple reference pictures. The reference pictures used for inter prediction can be one or more reference pictures. The reference index is used to identify any reference picture in the list. The prediction signal is generated using the following architecture:

[0037] One reference picture list, List0, is used for P slices, and two reference picture lists, List1, List0 and List1 are used for B slices. The reference pictures used can come from past and future pictures in terms of capture / display order. It's okay to have it.

[0038] 2.2.1 Merge Mode

[0039] 2.1.2.1. Deriving Merge Mode Candidates

[0040] When predicting PU using merge mode, the merge candidate list is generated from the bitstream. parse the index pointing to an entry in The construction of this list is specified in the HEVC standard and follows this sequence of steps: It can be summarized based on the ●Step 1: Derive initial candidates Step 1.1: Spatial candidate derivation Step 1.2: Spatial candidate redundancy check Step 1.3: Temporal candidate derivation ● Step 2: Insert additional candidates Step 2.1: Creating bidirectional prediction candidates Step 2.2: Insertion of zero motion candidates

[0041] These steps are also shown diagrammatically in Figure 1, which shows the steps for constructing the merge candidate list. An exemplary derivation process is shown. For spatial merge candidate derivation, We select up to four merge candidates from the candidates. At most one merge candidate is selected from the four candidates. Since we are assuming candidates, the number of candidates obtained in step 1 is signaled in the slice header. If the maximum number of merge candidates (MaxNumMergeCand) is not reached, additional candidates are generated. Since the number of candidates is fixed, we use truncated unary thresholding (TU) to find the best merge. If the CU size is equal to 8, encode the index of the candidate. The PU has one merge candidate list that is the same as the merge candidate list for the 2N × 2N prediction units. share.

[0042] The operations associated with the above steps are now described in detail.

[0043] 2.1.2.2 Spatial candidate derivation

[0044] In deriving spatial merge candidates, up to four merge candidates are selected from the candidates located in the positions shown in Figure 2. Select a page candidate. The order of derivation is A1, B1, B0, A0, B2. Position A1, B If any of PUs 1, B0, or A0 is not available (for example, another slice or Position B2 is considered only if it belongs to a file) or if it is intra-coded. After adding the candidate at position A1, the remaining candidates are added and subjected to a redundancy check. This allows candidates with the same motion information to be reliably removed from the list, improving coding efficiency. To reduce the computational complexity, the redundancy checks mentioned above can be performed using In this case, we do not consider all possible candidate pairs. Instead, we use the arrows in Figure 3. Only pairs linked by the same motion information are considered, and the corresponding candidates used for redundancy check are If a candidate does not have any other source of overlapping motion information, it is added to the list. A sub-process is a "second PU" associated with a division different from 2N x 2N. Figure 4 shows the second PU for the cases of N × 2N and 2N × N, respectively. In the case of a 2N division, the candidate at position A1 is not considered in the list construction. By doing so, two prediction units with the same motion information are derived, and one It is redundant to have only one PU in the current coding unit. When dividing the PU into 2NxN, position B1 is not taken into consideration.

[0045] 2.1.2.3 Temporal candidate derivation

[0046] In this step, only one candidate is added to the list. In the derivation of the merge candidates, the co-located pictures in the co-located pictures Based on the PU, a scaled motion vector is derived. 5 is an explanatory diagram of scaling of motion vectors for image candidates. As shown by the dotted line in FIG. The scaled motion vectors of the temporal merge candidates are obtained, which are the POC distance t b and td are used to scale the motion vector of the PU at the same position. tb is defined as the POC difference between the reference picture of the current picture and the current picture. , td is defined as the POC difference between the reference picture of the co-located PU and the co-located picture. Set the reference picture index of the temporal merge candidate equal to zero. A practical implementation of ring processing is described in the HEVC specification [1]. In the case of a reference picture list, two motion vectors are used: one for reference picture list 0 and the other for reference picture list 1. One for reference picture list 1 and by combining these , forming bi-predictive merge candidates.

[0047] 2.1.2.4 Co-located Pictures and Co-located PUs

[0048] TMVP is enabled (e.g. slice_temporal_mvp_en If the collocated_flag is 1, a variable C representing the collocated picture olPic is derived as follows:

[0049] - The current slice is a B slice and the signal collocated_from_l0 If _flag is 0, ColPic is RefPicList1[collocat ed_ref_idx].

[0050] - Otherwise (slice_type is B and collocated_fr om_l0_flag is 1 or slice_type is P), ColP Set ic to RefPicList0[collocated_ref_idx].

[0051] where collocated_ref_idx and collocated_fro m_l0_flag are two syntax elements that can be signaled in the slice header .

[0052] In the same position PU(Y) belonging to the reference frame, as shown in Figure 6, candidate C0 and Select a temporal candidate position between candidates C1 and C2. If a PU at position C0 is not available, If intra-coded or in the current coding tree unit (CT U, also known as LCU (Largest Coding Unit) If it is outside the line, position C1 is used Otherwise, position C0 is used to derive temporal merge candidates.

[0053] The relevant syntax elements are described below.

[0054] [Table 1]

[0055] 2.1.2.5 Deriving MVs for TMVP Candidates

[0056] In some embodiments, to derive TMVP candidates, the following operations are performed.

[0057] 1) In the list X, set the reference picture list X=0 and set the target reference picture Set the reference picture with index 0 (for example, curr_ref). Call the derivation process of the placed motion vector, and the MV of the list X that points to curr_ref is obtain.

[0058] 2) If the current slice is a B slice, set the reference picture list X=1; The target reference picture is the reference picture with index 0 in list X ( For example, curr_ref). Calls the derivation process of the motion vectors arranged at the same position. and get the MV of list X that points to curr_ref.

[0059] The process of deriving co-located motion vectors is described in the next subsection 2.1. .2.5.1 explains.

[0060] 2.1.2.5.1 Derivation of motion vectors located at the same position

[0061] For co-located blocks, it is either uni-predictive or bi-predictive intra-predictive. It may be coded or inter-coded. If the TMVP candidate is disabled, it is set to be equal to unavailable.

[0062] In the case of single prediction from list A, the motion vector of list A is used as the reference picture list. Scale to X.

[0063] In bi-prediction, if the target reference picture list is X, then the motion vector of list A is The vector is scaled to the target reference picture list X, and A is scaled according to the following rules: It is decided.

[0064] - None of the reference pictures has a larger POC value compared to the current picture. If not, A is set equal to X.

[0065] - Otherwise, A is equal to collocated_from_l0_flag It is set.

[0066] 2.1.2.6 Additional Candidate Inserts

[0067] Besides spatial-temporal merge candidates, there are two additional types of merge candidates: There are bi-predictive merge candidates and zero merge candidates. The joint bidirectional prediction merge candidate is generated by the B slice. The first reference picture list motion parameters of the first candidate and the second reference picture list motion parameters of another candidate are used only for the first candidate. By combining the two reference picture lists and motion parameters, joint bidirectional prediction candidates are generated. If these two tuples provide different motion hypotheses, they are used to generate a new Figure 7 shows an example of a combined bidirectional prediction merge candidate. Figure 7 shows the mvL0 and refIdxL0 in the original list (left), Or using two candidates with mvL1 and refIdxL1, the final list (right ) to generate combined bidirectional predictive merge candidates. There are various rules for the combinations that are considered to generate candidates.

[0068] By inserting zero motion candidates and filling the remaining entries in the merge candidate list , hits the MaxNumMergeCand capacity. These candidates have a spatial displacement of zero. The number of reference pictures starts from zero and increases each time a new zero motion candidate is added to the list. Finally, no redundancy check is performed on these candidates. stomach.

[0069] 2.1.3 Advanced Motion Vector Prediction (AMVP)

[0070] AMVP exploits the spatial-temporal correlation between motion vectors and neighboring PUs, It is used to explicitly transmit motion parameters. For each reference picture list, first, the left, top, Check the availability of PU positions in the time neighborhood of , remove redundant candidates, and select the zero vector By adding the ,the length of the candidate list is kept constant, and the motion vector candidate list is constructed. The encoder then selects the best predictor from the candidate list and writes the selected candidate as The corresponding index indicating the merge index can be transmitted. Similarly, the index of the best motion vector candidate is encoded using a shortened unary In this case, the maximum number of coding targets is 2 (see Figure 8). The process of deriving vector prediction candidates will now be described in detail.

[0071] 2.1.3.1 Derivation of AMVP Candidates

[0072] FIG. 8 summarizes the process of deriving motion vector prediction candidates.

[0073] In motion vector prediction, there are two types of motion vector candidates: spatial and temporal. Two types of motion vector candidates are considered: Then, as shown in Figure 2, the best position is calculated based on the motion vectors of each PU at five different positions. Ultimately, two motion vector candidates are derived.

[0074] To derive candidate temporal motion vectors, A motion vector candidate is selected from two candidates derived based on the spatial-temporal After creating the initial list of candidates, remove duplicate motion vector candidates in the list. If the number of candidates is greater than two, the reference picture in the associated reference picture list is used. Remove from the list any motion vector candidates with a vector index greater than 1. If the number of potential motion vector candidates is less than two, add additional zero motion vector candidates to the list. Add.

[0075] 2.1.3.2 Spatial Motion Vector Candidates

[0076] In deriving spatial motion vector candidates, they are derived from the PUs located as shown in Figure 2. Of the five possible candidates, up to two candidates that are in the same position as the motion merge are selected. The order of derivation for the left side of the current PU is A0, A1, scaled The derivation order for the upper side of the current PU is defined as A0, A1. The order is B0, B1, B2, scaled B0, scaled B1, scaled Therefore, for each side, the vectors that can be used as motion vector candidates are There are four cases where spatial scaling can be used: two where spatial scaling is not required, and There are two cases where dynamic scaling is used. The four different cases can be summarized as follows: It becomes like this.

[0077] ● No spatial scaling -(1) The same reference picture list and the same reference picture index (same POC ) (2) Different reference picture lists but the same reference picture (same POC) ● Spatial scaling (3) Same reference picture list but different reference pictures (different POC) (4) Different reference picture lists and different reference pictures (different POCs)

[0078] First check the non-spatial scaling case, then do the spatial scaling. Regardless of the reference picture list, the POC selects the reference pictures of neighboring PUs and the reference picture of the current PU. If the picture is different from the left one, spatial scaling is taken into account. All PUs of the left candidate are available. If not available or if intra-coded, the upper motion vector is The ring helps in parallel derivation of left and upper MV candidates. Otherwise, the upper MV No spatial scaling is allowed for the vectors.

[0079] In the spatial scaling process, as shown in Figure 9, the same process is carried out as in the temporal scaling. The main difference is that the reference picture of the current PU is The actual scaling process is performed by taking the time scale and the index as input. The point is that it is the same as the ring.

[0080] 2.1.3.3 Temporal Motion Vector Candidates

[0081] Processing for deriving temporal merge candidates other than deriving reference picture indexes are all the same as the processes for deriving spatial motion vector candidates (see FIG. 6). The reference picture index is signaled to the decoder.

[0082] 2.2 Inter-Prediction Method in VVC

[0083] New coding tools to improve inter prediction include signaling MVD Adaptive Motion Vector Difference Resolution (AMVR), Merge with Motion Vector Difference (MMVD), Triangular Prediction Mode (TPM), Combined Intra-Inter Prediction (CIIP), Advanced TM VP (ATMVP, also known as SbTMVP), affine prediction mode, generalized bidirectional prediction (GB I), Decoder-side Motion Vector Refinement (DMVR), Bidirectional Optical Flow (BIO , also known as BDOF).

[0084] There are two different merge list building processes supported by VVC:

[0085] (1) Sub-block merge candidate list: includes ATMVP and affine merge candidates. One merge list construction process is used for both affine and ATMVP modes. The ATMVP and affine merge candidates may be added in order. The size of the lock merge list is signaled in the slice header and has a maximum value of 5.

[0086] (2) Normal merge list: For inter-coding blocks, one merge list is used. The spatial / temporal merge candidate,HMVP,pairs are combined into a common structure. Merge candidates and zero motion candidates may be inserted in order. is signaled in the slice header and has a maximum value of 6. MMVD, TPM, CIIP depends on the normal merge list.

[0087] Similarly, the AMVP lists supported by VVC are: (1) Affine AMVP Candidate List (2) Regular AMVP candidate list

[0088] 2.2.1 Coding Block Structure in VVC

[0089] In VVC, a quad tree / binary tree / ternary tree (QT / BT / TT) structure is adopted, and the picture Divide the image into square or rectangular blocks.

[0090] Besides QT / BT / TT, there is a separate tree (aka dual codec) for I-frames. A separate tree is used for coding. The block structure is signaled separately for the luminance and chrominance components.

[0091] Also, CU is an intra-subpartition (e.g., PU is equal to TU but is smaller than CU). and inter-coding where PU is equal to CU but TU is less than PU. Two specific coding methods (such as sub-block transformation of coding blocks) Blocks other than the assigned block are set to PU and TU.

[0092] 2.2.2 Merging Whole Blocks

[0093] 2.2.2.1 Merge list construction in translational normal merge mode

[0094] 2.2.2.1.1 History-Based Motion Vector Prediction (HMVP)

[0095] Unlike the merge list design, VVC uses history-based motion vector prediction (HMV P) method is adopted.

[0096] HMVP stores previously coded motion information. The motion information of the block is defined as an HMVP candidate. The new table is maintained on the fly during the encoding / decoding process. When starting to encode / decode a new tile / LCU row / slice, the HMVP table is emptied. If there are inter-coding blocks and non-sub-blocks in non-TPM mode, At any time, add the relevant motion information as a new HMVP candidate to the last entry in the table. The overall coding flow is shown in Figure 10.

[0097] 2.2.2.1.2 Normal Merge List Construction Process

[0098] A typical merge list construction (for translation) follows this sequence of steps: We can summarize it as follows.

[0099] Step 1: Derive spatial candidates Step 2: Inserting HMVP candidates Step 3: Inserting pairwise average candidates Step 4: Default motion candidates

[0100] HMVP candidates can be used in both AMVP and merge candidate list construction processes. Figure 11 shows the modified merge candidate list construction process (highlighted in blue). After inserting, if the merge candidate list is not full, the HM stored in the HMVP table VP candidates can be used to fill in the merge candidate list. A block is usually From the viewpoint of motion information, it is considered that the block has a high correlation with its nearest neighbor. Considering this, we insert the HMVP candidates in the table in descending order of their indices. Add it to the list first, then add the first entry to the end. Similarly, add redundant entries to the HMVP candidates. The total number of available merge candidates is used to signal which merge candidates are mergable. Once the maximum possible number is reached, the merge candidate list construction process ends.

[0101] Note that all spatial / temporal / HMVP candidates are coded in non-IBC mode. If not, it will not be allowed to be added to the regular merge candidate list.

[0102] The HMVP table contains up to five canonical move candidates, each unique.

[0103] 2.2.2.1.2.1 Pruning Process

[0104] Only if the corresponding candidates used for redundancy check do not have the same motion information The candidate is added to the list. This comparison process is called pruning.

[0105] The pruning process among spatial candidates depends on the use of the TPM of the current block.

[0106] The current block is not coded in TPM mode (e.g., normal merge, In the case of MMVD, CIIP, HEVC pruning process for spatial merge candidates (e.g. For example, 5 prunings are used.

[0107] 2.2.3 Triangle Prediction Mode (TPM)

[0108] In VVC, triangle partition mode (triangl) is used for inter prediction. Triangulation modes are supported, with 8x8 and above being supported. and applies only to CUs coded in merge mode, MMVD or CII Not applicable in P mode. For CUs that meet these conditions, the CU level flag is signaled. Indicates whether the triangulation mode is applied.

[0109] When using this mode, either diagonal or anti-diagonal division is performed, as shown in Figure 11. Divide one CU into two equal triangular partitions using The rectangular partitions are inter-predicted using their own motion, and each partition Only uni-prediction is allowed for each partition, i.e., each partition can only be predicted by one motion vector. As with conventional bi-prediction, there are two motion vectors per CU and one reference index. To ensure that only compensated prediction is required, a unipredictive motion constraint is applied.

[0110] FIG. 12 shows an example of inter prediction based on triangulation.

[0111] The CU level flag indicates that the current CU is coded in triangle partition mode. If it indicates that the triangle partition is diagonal or anti-diagonal, a flag indicating the direction of the triangle partition (diagonal or anti-diagonal) and and two more merge indexes (one for each partition). After predicting each of the rectangular partitions, we use adaptive weighted blending to blend the diagonal and This is the predicted signal for the entire CU, and the other samples along the edge of the diagonal are adjusted. Similarly to the prediction mode of , the transformation and quantization process is performed on the entire CU. The motion fields of the predicted CUs are stored in 4x4 units using the partition mode.

[0112] A regular merge candidate list can merge triangles without extra pruning of motion vectors. Reused for split-merge prediction. For each merge candidate in the regular merge candidate list, , only one of the L0 or L1 motion vectors is used for triangular prediction. The order of selecting L1 pair motion vectors is based on their merge index parity. According to the scheme, a regular merge list can be used directly.

[0113] 2.2.3.1 TPM Merge List Construction Process

[0114] Basically, you can make some modifications to the normal merge list construction process. The following applies to:

[0115] (1) How the pruning process is performed depends on the TPM usage of the current block. Depends on the law.

[0116] - If the current block is not TPM coded, it is suitable for spatial merging candidates. This is called HEVC5 pruning.

[0117] - Otherwise (if the current block is coded in the TPM), the new When adding new spatial merge candidates, full pruning is applied. Compare with A1, compare B0 with A1, B1, compare A0 with A1, B1, compare B0, B2 Compare with A1, B1, A0, and B0.

[0118] (2) The condition for checking the motion information from B2 is Depends on the use of the TPM.

[0119] - If the current block is not coded by the TPM, access B2 and Check B2 only if there are less than four spatial merge candidates before checking B2 do.

[0120] - Otherwise (if the current block is TPM coded), B2 B2 is always accessed before adding B1, regardless of the number of spatial merge candidates available. , check.

[0121] 2.2.3.2 Adaptive Weighting

[0122] After predicting each triangle prediction unit, adapt to the diagonal edge between two triangle prediction units. The weighting process is performed to derive the final prediction for the entire CU. The two weighting coefficient groups are as follows: It is defined as follows.

[0123] The luminance and chrominance samples are respectively set to {7 / 8, 6 / 8, 4 / 8, 2 / 8 ,1 / 8} and the first weighting factor group of {7 / 8,4 / 8,1 / 8} is used.

[0124] The luminance and chrominance samples are respectively {7 / 8, 6 / 8, 5 / 8, 4 / 8 ,3 / 8,2 / 8,1 / 8} and the second weighting factor group of {6 / 8,4 / 8,2 / 8} Use.

[0125] Selecting a weighting factor group based on a comparison of the motion vectors of the two triangular prediction units The second group of weighting factors is used if any one of the following conditions is true: .

[0126] - The reference pictures of the two triangular prediction units are different from each other. - The absolute value of the difference between the horizontal values ​​of the two motion vectors is greater than 16 pixels. - The absolute value of the difference between the vertical values ​​of the two motion vectors is greater than 16 pixels.

[0127] If not, the first weighting factor group is used. An example is shown in FIG.

[0128] 2.2.3.3 Motion Vector Storage

[0129] The motion vectors of the triangular prediction unit (Mv1, Mv2 in Fig. 14) are For each 4x4 grid, the CU stores the 4x4 grid position. Based on the result, the uni-predictive or bi-predictive motion vector is stored. For a 4x4 grid that is in the finding region (i.e., not on the diagonal edge), M The single predicted motion vector v1 or Mv2 is stored. For the 4x4 grid located in the frame, the bidirectional predicted motion vectors are stored according to the following rules: Therefore, a bidirectional predicted motion vector is derived from Mv1 and Mv2.

[0130] (1) Mv1 and Mv2 have motion vectors in different directions (L0 or L1). In this case, Mv1 and Mv2 are simply combined to form a bidirectional motion vector predictor. do.

[0131] (2) If both Mv1 and Mv2 are coming from the same L0 (or L1) direction, then: is. - The reference picture of Mv2 is a picture in the L1 (or L0) reference picture list If Mv1 is equal to Mv2, Mv2 is scaled to that picture. The bidirectional motion vector predictor is formed by combining the obtained Mv1 with the obtained Mv2. - The reference picture of Mv1 is a picture in the L1 (or L0) reference picture list. If Mv1 is equal to the picture, Mv1 is scaled to that picture. v1 and Mv2 are combined to form a bidirectional motion vector predictor. - Otherwise, only Mv1 is stored for the weighting region.

[0132] 2.2.4 Merging with Motion Vector Difference (MMVD)

[0133] In some embodiments, the proposed motion vector representation method is Ultimate Motion Vector Representation (UMVE, MMV) for either mode or merge mode D) is used.

[0134] UMVE is a merge candidate list, just like the ones in the regular merge candidate list in VVC. Reuse candidates: Among the merge candidates, you can select a base candidate and reuse the proposed behavior. This is further extended by the vector representation method.

[0135] UMVE provides a new motion vector difference (MVD) representation method, where the starting point, The magnitude and direction of the motion are used to represent a single MVD.

[0136] FIG. 15 shows an example of the UMVE search process.

[0137] FIG. 16 shows an example of UMVE search points.

[0138] This technique uses the merge candidate list as is. However, to extend UMVE, Only consider candidates for the default merge type (MRG_TYPE_DEFAULT_N) do.

[0139] The base candidate index defines the starting point. Among the candidates, the best candidates are shown below.

[0140] [Table 2]

[0141] If the number of base candidates is 1, the base candidate IDX is not signaled.

[0142] The distance index is the information of the magnitude of the movement. The predefined distances are as follows:

[0143] [Table 3]

[0144] The direction index represents the direction of the MVD relative to the starting point. Four directions can be represented as shown in

[0145] [Table 4]

[0146] The UMVE flag is signaled immediately after sending the skip or merge flag. If the skip or merge flag is true, parse the UMVE flag. If lag is 1, the UMVE syntax is parsed. But if it is not 1, AFFI Parse the NE flag. If the AFFINE flag is 1, i.e., AFFINE If the mode is 1, but not 1, the VTM skip / merge mode is used. Parse the page index.

[0147] No additional line buffers are required due to the UMVE candidate. This is because the split / merge candidates are used directly as base candidates. to determine the MV completion just before motion compensation. A long line buffer is used for this purpose. There's no need to hold it.

[0148] The first or second merge candidate in the merge candidate list under the current common test conditions may be selected as a base candidate.

[0149] UMVE is also known as Merge with MV Difference (MMVD).

[0150] 2.2.5 Merging for Sub-Block Based Techniques

[0151] In some embodiments, all sub-block related motion candidates are non-sub-block In addition to the regular merge list of merge candidates, they are put into a separate merge list.

[0152] Put the sub-block related motion candidates into a separate merge list and call it the “sub-block merge candidate list.” "Strike".

[0153] In one example, the subblock merge candidate list includes ATMVP candidates and affine mergers. Includes candidates.

[0154] The sub-block merge candidate list is filled with candidates in the following order: a. ATMVP candidates (possibly available or unavailable). b. Affine merge list (including inheritance affine candidates and construction affine candidates). c.0MV4 Padding as a parameter affine model

[0155] 2.2.5.1 Advanced Temporal Motion Vector Prediction Module (ATMVP) (also known as Subblock Temporal Motion Vector Predictor (SbTMVP)

[0156] The basic idea of ​​ATMVP is to use multiple temporal motion vector predictions for one block. Each sub-block is assigned a set of motion information. When generating ATMVP merge candidates, we use 8x8 level instead of whole block level. Motion compensation is performed in the

[0157] In the current design, ATMVP is defined in the following two subsections: 2.2.5.1.1 and 2.2.5.1.2, respectively, of the sub-CUs within the CU. Predict the motion vectors.

[0158] 2.2.5.1.1 Deriving Initial Motion Vectors

[0159] Here, the initialized motion vector is tempMv. Block A1 is available. and not intra-coded (e.g., Inter or IBC mode) , the following is used to derive the initialized motion vectors: Applies. - tempMv is represented by mvL1A1 if all of the following conditions are true: , is set equal to the motion vector of block A1 from list 1. - The reference picture index in list 1 is available (not -1) and is the same have the same POC value as the picture placed in one position (e.g., DiffPicOrder rCnt(ColPic,RefPicList[1][refIdxL1A1)] is 0 ), - All reference pictures have no large POC compared to the current picture (e.g. DiffPicOrderCnt(aPic,currPic) is the order of all the pixels in the current slice. (It is less than or equal to 0 for all pictures aPic in all reference picture lists). - The current slice is the B slice. - collocated_from_l0_flag equals 0. - Otherwise, if all of the following conditions are true, tempMv is added to mvL0A1: and is set equal to the motion vector of block A1 from list 0. - Reference picture index in list 0 is available (not -1). - It has the same POC value as the co-located picture (e.g., Diff PicOrderCnt(ColPic,RefPicList[0][refIdxL 0A1]) is equal to 0). Otherwise, the zero motion vector is used as the initialized MV.

[0160] The initialized motion vectors are assigned to the co-located slices signaled in the slice header. In the image, add the rounded MV to the center position of the corresponding block (current block). , clipping to a range if necessary).

[0161] If the block is inter-coded, go to the second step; otherwise If so, the ATMVP candidate is set to unavailable.

[0162] 2.2.5.1.2 Sub-CU Motion Derivation

[0163] The second step is to split the current CU into sub-CUs and assign them to co-located pictures. The motion information of each sub-CU is obtained from the block corresponding to each sub-CU in the sub-CU block.

[0164] When the corresponding block for one sub-CU is coded in inter mode Using this motion information, we can generate co-located images, which is not different from the conventional TMVP processing. The final motion information of the current sub-CU is obtained by invoking the derivation process for the selected MV. Basically, if the corresponding block is a target list for uni-prediction or bi-prediction, If the motion vector is predicted from the input X, then it is used, otherwise it is either single or bidirectional. For prediction, it is predicted from list Y (Y=1-X) and NoBackwardPred If Flag is 1, the MV for list Y is used. Otherwise, the movement candidate I can't find the complement.

[0165] Co-located blocks in the picture identified by the initialized MV and and whether the current sub-CU position is intra-coded or IBC-coded If the motion candidate is not found as described above, the following is further This applies to:

[0166] Co-located Picture R col is used to extract the motion region in Motion vectors are expressed as MVcol In order to minimize the impact of MV scaling, MV col The MVs in the spatial candidate list for deriving the MVs are the reference points of the candidate MVs. If the image is a co-located picture, select this MV and do not scale it. MV without col otherwise, use the most appropriate Select the MV with the closest reference picture and scale it accordingly. col is derived.

[0167] Decoding process for deriving motion vectors arranged at the same position in an embodiment of the present invention is explained below.

[0168] 8.5.2.12 Derivation of Co-located Motion Vectors

[0169] The inputs to this process are: - the variable currCb that defines the current coding block, -ColPic specifies a co-located code in a co-located picture. The variable colCb that specifies the coding block, - the top left luminance sample of the co-located picture specified by ColPic for the co-located luminance coding block specified by colCb. the luminance position (xColCb, yColCb) that specifies the top-left sample of the block, - a reference index refIdxLX, where X is 0 or 1; - Flag indicating sub-block temporal merging candidate sbFlag.

[0170] The output of this process is: - Motion vector prediction mvLXCol with 1 / 16 fractional sample accuracy. -Availability flag availableFlagLXCol.

[0171] The variable currPic specifies the current picture. Arrays predFlagL0Col[x][y], mlvL0Col[x][y], and and refIdxL1Col[x][y] are the same as the ColPic specified by PredFlagL0[x][y], MvDmvrL0[x RefIdxL0[x][y], and the array predFlagL 1Col[x][y], mvL1Col[x][y], and refIdxL1Col[ x][y] are the Pr of the picture placed at the same position specified by ColPic. edFlagL1[x][y], MvDmvrL1[x][y] and RefIdxL1 Set equal to [x][y].

[0172] The variables mvLXCol and availableFlagLXCol are derived as follows: It is served. -mvLXC if colCb is coded in intra or IBC prediction mode Both modules of ol are set equal to 0, and availableFlagLXCol is set equal to 0. Otherwise, the motion vector mvCol, the reference index refIdxCol, and and the reference list identifier listCol is derived as follows: - If sbFlag is 0, availableFlagLXCol is equal to 1 is set and the following applies: -predFlagL0Col[xColCb][yColCb] is 0, mvCol, refIdxCol and listCol are mvL1Col[ xColCb][yColCb], refIdxL1Col[xColCb][yCol Cb] and L1. - otherwise, predFlagL0Col[xColCb][yColCb] is equal to 1 and predFlagL1Col[xColCb][yColCb] is 0. If mvCol, refxIdCol and listCol are mvL0 and mvL1 respectively. Col[xColCb][yColCb], refIdxL0Col[xColCb][ yColCb] and L0. - else (predFlagL0Col[xColCb][yColCb] is equal to 1 and predFlagL1Col[xColCb][yColCb] is 1 ), and make the following allocations: -If NoBackwardPredFlag is 1, mvCol, refI dxCol and listCol are mvLXCol[xColCb][yC olCb], refIdxLXCol[xColCb][yColCb], and LX are set equal. - Otherwise, mvCol, refIdxCol, and listCol, respectively. mvLNCol[xColCb][yColCb], refIdxLNCol[xCol Cb][yColCb] and LN, and set N equal to collocated_from Let it be the value of m_l0_flag. Otherwise (sbFlag is 1), the following applies: -If PredFlagLXCol[xColCb][yColCb] is 1, mv Col, refxIdCol and listCol are mvLXCol[xC olCb][yColCb], refIdxLXCol[xColCb][yColCb ] and LX are set equal to 1, and availableFlagLXCol is set equal to 1. It is determined. - else (PredFlagLXCol[xColCb][yColCb] is 0) is applied as follows: -DiffPicOrderCnt(aPic,currPic) is the current slide In all reference picture lists aPic of the device, the If agLYCol[xColCb][yColCb] is 1, then mvCol, ref xIdCol and listCol are mvLYCol[xColCb][ yColCb], refIdxLYCol[xColCb][yColCb], and L Set equal to Y, where Y is !X, and X is the value of X for which this operation is called, and a The availableFlagLXCol is set equal to 1. Both modules of -mvLXCol are set equal to 0, and availableFl agLXCol is set equal to 0. If -availableFlagLXCol is TRUE, mvLXCol and and availableFlagLXCol are derived as follows: -LongTermRefPic(currPic,currCb,refIdxL X,LX) is LongTermRefPic(ColPic,colCb,refIdx Col, listCol), the modules of mvLXCol are both set to 0, and a The availableFlagLXCol is set equal to 0. - Otherwise, set the variable availableFlagLXCol to 1 and re fPicList[listCol][refIdxCol] by ColPic Contains the coding block colb in the specified co-located picture Reference index refIdxC in the slice's reference picture list listCol ol and apply the following: colPocDiff=DiffPicOrderCnt(ColPic,ref PicList[listCol][refIdxCol]) (8-402) currPocDiff=DiffPicOrderCnt(currPic,R efPicList[X][refIdxLX]) (8-403) -mvCol as input and modified mvCol as output, 8.5.2.15 Temporal motion buffer compression of co-located motion vectors as defined in paragraph is called. -If RefPicList[X][refIdxLX] is a long-term reference picture If colPocDiff is currPocDiff, or if colPocDiff is currPocDiff, then mvLXCol is derived as follows: mvLXCol=mvCol (8-404) - Otherwise, mvLXCol is the scaled version of the motion vector mvCol. The modified version is derived as follows: tx=(16384+(Abs(td)>>1)) / td (8-405) distScaleFactor=Clip3(-4096,4095,(tb* tx+32)>>6) (8-406) mvLXCol=Clip3(-131072,131071,(distSca leFactor*mvCol+128-(distScaleFactor*mvCo l>=0))>>8)) (8-407) Here, td and tb are derived as follows: td=Clip3(-128,127,colPocDiff) (8-408) tb=Clip3(-128,127,currPocDiff) (8-409 )

[0173] 2.3 Intra-block copy

[0174] Intra Block Copy (IBC), also known as the current picture reference, is used in HEVC scripts. HEVC Content Coding Extensions (HEVC-SCC) and the current VVC test model IBC is an inter-frame coding scheme that uses the concept of motion compensation. As shown in Figure 17, indicates the operation of intra block copy, and the current block is It is predicted by one reference block in the same picture. Before decoding, the samples in the reference block must already be reconstructed. IBC is not very efficient for most camera-captured sequences. Although not ideal, it does show significant coding gains for screen content. The reason is that in the screen content picture, there are repeated patterns of icons, characters, etc. IBC can effectively remove redundancies between these repeating patterns. In HEVC-SCC, the inter-coding unit (CU) is currently If a picture is selected as its reference picture, IBC can be applied. In this case, MVs are renamed block vectors (BVs), and BVs always have integer pixel precision. To conform to the Main Profile HEVC, the current picture is stored in the decoded picture buffer. It is marked as a "long-term" reference picture in the DPB. In the 3D / 4D video coding standard, inter-view reference pictures are also "long-term" references. It is marked as a picture.

[0175] After the BV finds the reference block, it generates a prediction by copying this reference block. The residual can be obtained by subtracting the reference pixel from the original signal. And, like other coding modes, transforms and quantization can be applied. can.

[0176] However, if the reference block is outside the picture or overlaps with the current block, If the area is too small, or if it is outside the reconstructed area, or if it is limited by some constraint, If the pixel value is outside the specified valid area, some or all of the pixel value is undefined. There are two solutions to deal with such problems. The other is to not allow for data stream conformance with these undefined pixel values. The solution is to apply coding. The following subsections explain the solution in detail.

[0177] 2.3.1 Single BV List

[0178] In some embodiments, for merge mode and AMVP mode in IBC, The BV predictors for each share a common predictor list that contains the following elements:

[0179] (1) Two spatially adjacent locations (A1 and B1 as in Figure 2) (2) Five HMVP entries (3) Default zero vector

[0180] The number of candidates in the list is controlled by a variable derived from the slice header. In merge mode, up to the first six entries in this list are used, and in AMVP mode In the case of a shared merge, the first two entries in this list are used. Complies with the requirements for list areas (sharing the same list within the SMR).

[0181] In addition to the BV predictor candidate list mentioned above, we also consider the HMVP candidate and the existing merge candidate (A1, B1). The pruning operation can be simplified. In this case, the first HMVP candidate Since the process only compares the spatial merge candidates with the original data, a maximum of two pruning operations are sufficient. stomach.

[0182] In IBC AMVP mode, the mv difference between the selected MVP and the list is further In IBC merge mode, the selected MVP is sent to the current block. It will be used as the music video.

[0183] 2.3.2 IBC Size Limits

[0184] In recent VVC and VTM5, in previous VTM and VVC versions ,In addition to the current bitstream constraints, to disable 128x128 IBC mode It is proposed to explicitly use the syntax constraint of , which would allow the presence of the IBC flag to be used as a CU size < 128x128.

[0185] 2.3.3 IBC Shared Merge List

[0186] To reduce decoder complexity and support parallel coding, in some embodiments Enables parallel processing of small skip / merge coded CUs To do this, we first divide the coding units of all the leaves of an ancestor node in the CU partition tree. The same merge candidate list is shared for each CU. The merge-share node is called a leaf CU. Generate a shared merge candidate list in

[0187] Specifically, the following may apply:

[0188] - The block has 32 or fewer luminance samples (e.g., 4x8 or 8x4) and two If a pixel is divided into 4x4 child blocks, a very small block (e.g., two adjacent Use merge lists between blocks (4x4 blocks).

[0189] - However, if the number of luminance samples in this block is greater than 32, then one division After this, all the splits are removed because at least one child block is smaller than the threshold (32). All child blocks share the same merge list (e.g., 16x4 or 4x3). x16, or 8x8 divided into 4).

[0190] Such restrictions apply only to IBC merge mode.

[0191] 2.3.4 Syntax Tables

[0192] [Table 5] [Table 6]

[0193] [Table 7]

[0194] 2.4 Maximum Transform Block Size

[0195] Only the maximum luminance transform size, which is either 64 or 32, is affected by the flags at the SPS level. The chroma sampling ratio for the maximum luma transform size is Derive the Roma conversion size.

[0196] If the CU / CB size is larger than the maximum luminance transform size, the size is calculated by a smaller TU. You can call a tiling of CUs.

[0197] The maximum luminance transform size is represented by MaxTbSizeY.

[0198] [Table 8]

[0199] sps_max_luma_transform_size_64_fla equal to 1 g specifies the maximum transform size of luma samples equal to 64. sps equal to 0 _max_luma_transformation_size_64_flag is This specifies that the maximum transform size of degree samples is equal to 32. If CtbSizeY is less than 64, sps_max_luma_transform The value of rm_size_64_flag is equal to 0. Variables MinTbLog2SizeY, MaxTbLog2SizeY, MinTbSi zeY and MaxTbSizeY are derived as follows: MinTbLog2SizeY=2 (7-27) MaxTbLog2SizeY=sps_max_luma_transform _size_64_flag?6:5 (7-28) MinTbSizeY=1< <MinTbLog2SizeY (7-29) MaxTbSizeY=1< <MaxTbLog2SizeY (7-30) sps_sbt_max_size_64_flag=0 enables sub-block conversion Specifies that the maximum CU width and maximum CU height to achieve is 32 luma samples. sps_sbt_max_size_64_flag=1 allows sub-block transformation. Specifies that the maximum CU width and height allowed is 64 luma samples. MaxSbtSize=Min(MaxTbSizeY,sps_sbt_max _size_64_flag ? 64:32) (7-32)

[0200] 2.4.1 Tools that rely on maximum conversion block size

[0201] The splitting of binary and ternary trees depends on MaxTbSizeY.

[0202] [Table 9] [Table 10] [Table 11] [Table 12]

[0203] 2.4.2 Recovery of intra- and inter-coded blocks Encryption processing

[0204] If the size of the current intra-coding CU is larger than the maximum transform block size , the CU is then resized until both its width and height are no larger than the maximum transform block size. , and then recursively divided into smaller blocks (PUs). All PUs have the same intra prediction Although they share the same prediction mode, the reference samples for intra prediction are different. Note that the current block is an intra-subpartition (ISP) It must not be coded in mode.

[0205] If the size of the current inter-coding CU is larger than the maximum transform block size , the CU is then resized until both its width and height are no larger than the maximum transform block size. Then, the image is recursively divided into smaller blocks (TUs). All TUs share the same motion information. However, the reference samples for intra prediction are different.

[0206] 8.4.5 Intrablock Decoding Process

[0207] 8.4.5.1 General Decoding Process for Intra Blocks

[0208] The inputs to this process are: - Defines the top-left sample of the current transform block relative to the top-left sample of the current picture. The sample position (xTb0, yTb0) to be determined, - variable nTbW, which specifies the width of the current transformation block; - variable nTbH, which specifies the height of the current transformation block; - a variable predModeIntra that specifies the intra prediction mode, - A variable cIdx that defines the color component of the current block.

[0209] The output of this processing is the corrected reconstructed image before in-loop filtering. The maximum transform block width maxTbWidth and height maxTbHeight are as follows: It is derived as follows. maxTbWidth=(cIdx==0)?MaxTbSizeY:MaxTb SizeY / SubWidthC (8-41) maxTbHeight=(cIdx==0)?MaxTbSizeY:MaxT bSizeY / SubHeightC (8-42) The luminance sample positions are derived as follows: (xTbY,yTbY)=(cIdx==0)?(xTb0,yTb0):(xT b0*SubWidthC,yTb0*SubHeightC) (8-43) Based on maxTbSize, the following applies: - IntraSubPartitionsSplitType is ISP_NO_SPLIT and nTbW is greater than maxTbWidth, or nTbH is greater than max If it is greater than TbHeight, the following ordered steps are applied:

[0210] 1. The variables newTbW and newTbH are derived as follows: newTbW=(nTbW>maxTbWidth)?(nTbW / 2):nTb W (8-44) newTbH=(nTbH>maxTbHeight)?(nTbH / 2):nT bH (8-45)

[0211] 2. The general decoding process for intra blocks specified in this section is 0, yTb0), the transform block width nTbW set equal to newTbW, and new The height nTbH is set equal to TbH, and the intra prediction mode predModeIntr a and variable cIdx as inputs, and the output is This is the corrected reconstructed picture.

[0212] 3. If nTbW is greater than maxTbWidth, the intrablock defined in this section The general decoding process for a block is set equal to (xTb0 + newTbW, yTb0). The specified position (xTb0, yTb0), the transform block width set equal to newTbW The height nTbH is set equal to nTbW and newTbH, and the intra prediction mode pr It is called with edModeIntra and the variable cIdx as input, and the output is 1 is the modified reconstructed picture before intra-filtering.

[0213] 4. If nTbH is greater than maxTbHeight, the intra-Tb The general decryption process for a lock is equal to (xTb0, yTb0 + newTbH). Set position (xTb0, yTb0), set transformation block equal to newTbW The width nTbW and height nTbH set equal to newTbH, the intra prediction mode p It is called with the inputs redModeIntra and the variable cIdx, and the output is 1 is the modified reconstructed picture before intra-frame filtering.

[0214] 5. nTbW is greater than maxTbWidth and nTbH is greater than maxTbHeight If the .intg.DELTA..times ... The position (xTb0) is set equal to xTb0+newTbW, yTb0+newTbH. ,yTb0), the transform block width nTbW set equal to newTbW and The height nTbH is set equal to bH, and the intra prediction mode predModeIntra , and variable cIdx as input, and the output is the modified This is the corrected reconstructed picture. - Otherwise, the following ordered steps are applied (normal intra prediction process): . - …

[0215] 8.5.8 Residual signal for coding blocks coded in inter prediction mode Decoding process for

[0216] The inputs to this process are: - Defines the top-left sample of the current transform block relative to the top-left sample of the current picture. The sample position (xTb0, yTb0) to be determined, - variable nTbW, which specifies the width of the current transformation block; - variable nTbH, which specifies the height of the current transformation block; - A variable cIdx that defines the color component of the current block.

[0217] The output of this process is the (nTbW) x (nTbH) array resSamples. . The maximum transform block width maxTbWidth and height maxTbHeight are as follows: It is derived as follows. maxTbWidth=(cIdx==0)?MaxTbSizeY:MaxTb SizeY / SubWidthC (8-883) maxTbHeight=(cIdx==0)?MaxTbSizeY:MaxT bSizeY / SubHeightC (8-884) The luminance sample positions are derived as follows: (xTbY,yTbY)=(cIdx==0)?(xTb0,yTb0):(xT b0*SubWidthC,yTb0*SubHeightC) (8-885) Depending on maxTbSize the following applies:

[0218] - nTbW is greater than maxTbWidth or nTbH is greater than maxTbHei If the value is greater than ght, then the following ordered steps are applied:

[0219] 1. The variables newTbW and newTbH are derived as follows: newTbW=(nTbW>maxTbWidth)?(nTbW / 2):nTb W (8-886) newTbH=(nTbH>maxTbHeight)?(nTbH / 2):nT bH (8-887)

[0220] 2. At position (xTb0, yTb0), set the transform block width nT equal to newTbW. bW, the height nTbH set equal to newTbH, and the variable cIdx as inputs. , coding units coded in inter prediction mode as defined in this section The output is the modified signal before in-loop filtering. The reconstructed picture is

[0221] 3. If nTbW is greater than maxTbWidth, (xTb0+newTbW,y Tb0), and the position (xTb0, yTb0) set equal to newTbW. The width of the transformed block nTbW is set equal to newTbH, and the height nTbH is set equal to the variable cI. dx as input, coded in inter prediction mode as specified in this section Invokes the decoding process of the residual signal of a coding unit, and the output is a modified reconstruction It is a picture.

[0222] 4. If nTbH is greater than maxTbHeight, (xTb0, yTb0+ne Set position (xTb0, yTb0) equal to newTbW. The width of the transformed block nTbW is set equal to newTbH, and the height nTbH is set equal to the variable c Idx is used as input to generate a The output is an in-loop flow. 1 is the modified reconstructed picture before filtering.

[0223] 5. nWTbW is greater than maxTbWidth and nTbH is greater than maxTbHei If it is greater than ght, it is equal to (xTb0+newTbW, yTb0+newTbH). Set position (xTb0, yTb0), set transformation block equal to newTbW The width is nTbW, the height is nTbH set equal to newTbH, and the variable cIdx is used as input. , as specified in this section, coding units coded in inter prediction mode The output is the modified signal before in-loop filtering. This is the corrected reconstructed picture.

[0224] - Otherwise, if cu_sbt_flag is equal to 1, the following applies: - The variables sbtMinNumFourth, wPartIdx, and hPartId x is derived as follows: sbtMinNumFourths=cu_sbt_quad_flag?1:2 (8-888) wPartIdx=cu_sbt_horizontal_flag?4:sbt MinNumFourths (8-889) hPartIdx=!cu_sbt_horizontal_flag?4:sb tMinNumFourths (8-890) The variables xPartIdx and yPartIdx are derived as follows: - If cu_sbt_pos_flag is equal to 0, then xPartIdx and y PartIdx is set equal to 0. - Otherwise (cu_sbt_pos_flag is equal to 1), the variable xPa rtIdx and yPartIdx are derived as follows: xPartIdx=cu_sbt_horizontal_flag?0:( 4-sbtMinNumFourths) (8-891) yPartIdx=!cu_sbt_horizontal_flag?0: (4-sbtMinNumFourths) (8-892) - Variables xTbYSub, yTbYSub, xTb0Sub, yTb0Sub, nTb WSubandnTbHSub are derived as follows: xTbYSub=xTbY+((nTbW*((cIdx==0)?1:SubW idthC)*xPartIdx / 4) (8-893) yTbYSub=yTbY+((nTbH*((cIdx==0)?1:SubH eightC)*yPartIdx / 4) (8-894) xTb0Sub=xTb0+(nTbW*xPartIdx / 4) (8-895 ) yTb0Sub=yTb0+(nTbH*yPartIdx / 4) (8-896 ) nTbWSub=nTbW*wPartIdx / 4 (8-897) nTbHSub=nTbH*hPartIdx / 4 (8-898) - Luminance position (xTbYSub, yTbYSub), variables cIdx, nTbWSub, nTbHSub is used as input and scaled and transformed as specified in Section 8.7.2. Call the process and output the (nTbWSub) x (nTbHSub) array resSam It is plesTb. - Residual samples resSamples[x][y](x=0..nTbW-1,y= 0..nTbH-1) equal to 0. - Residual samples resSamples[x][y](x=xTb0Sub..xTb 0Sub+nTbWSub-1,y=yTb0Sub..yTb0Sub+nTbHSu b-1) is derived as follows: resSamples[x][y]=resSamplesTb[x-xTb0S ub][y-yTb0Sub] (8-899) - Or, the luminance position (xTbY, yTbY), the variable cIdx, the transformation width nTbW, and and the transformation height nTbH as inputs, and The transform process is invoked and the output is a (nTbW) x (nTbH) array of resamples.

[0225] 2.5 Low Frequency Non-Separable Transform (LFNST) (also known as Reduced Quadratic Transform (RST) / Non-Separable Quadratic Next transformation (NSST)

[0226] In JEM, the secondary transform is between the forward primary transform and the quantization (at the encoder), and is applied between the inverse quantization and the inverse linear transform (at the decoder side). As shown, 4x4 (or 8x8) secondary transformations are performed depending on the block size. For example, A 4x4 quadratic transform is applied to small blocks (e.g., min(width, height)<8), The 8x8 quadratic transform is performed by dividing the 8x8 blocks by a larger block (e.g., min(width, height)>4).

[0227] Since a non-separable transformation is applied to the quadratic transformation, this is a non-separable quadratic transformation (Non-Sep Also known as the Neural Network Secondary Transform (NSST) There are a total of 35 transformation sets, each containing three non-separable transformation matrices ( kernels, each with a 16x16 matrix) are used.

[0228] According to the intra prediction direction, the reduced 2D transform, also known as low frequency non-separable transform (LFNST), is used. We introduce the first transformation (RST) and use a matching set of four transformations (instead of 35). In this contribution, we introduce the , 16x48 and 16x16 matrices are used, respectively. For convenience of notation, we use 16x48 variables. The 16x16 transformation is represented as RST8x8, and the 16x16 transformation is represented as RST4x4. , which has been adopted by VVC in recent years.

[0229] FIG. 19 shows an example of a reduced quadratic transformation (RST).

[0230] The secondary forward and inverse transforms are separate processing steps from the primary transform. be.

[0231] In the case of the encoder, first a first-order forward transform is performed, then a second-order forward transform and quantization, then a parallel Decoder, CABAC bit decoding, and inverse quantization For intra-coding, the secondary inverse transform is performed first, followed by the primary inverse transform. Applies only to Serving TUs.

[0232] 3. Exemplary technical problems solved by the technical solutions disclosed herein

[0233] The current VVC design has the following problems from the perspective of IBC mode and conversion design:

[0234] (1) The signaling of the IBC mode flag depends on the block size limit. However, the IBC skip mode flag is not an insurmountable worst case scenario (e.g. For example, large blocks such as Nx128 or 128xN still require IBC mode. can be).

[0235] (2) In the current VVC, the size of the transformation block is always equal to 64x64. The size of the conversion block is always set to the size of the coding block. It is unclear how to handle the case where the lock size is larger than the conversion block size.

[0236] (3) IBC is a method for filtering filtered images within the same video unit (e.g., slice). However, filtered samples (e.g. , via deblocking filter / SAO / ALF) is the unfiltered signal. Use filtered samples as they may have less distortion than regular samples. may result in further coding gain.

[0237] (4) Context modeling of the LFNST matrix is ​​performed using linear transformations and partitions. It depends on the type (single or dual tree). However, from our analysis, the There is no clear dependency that the RST matrix is ​​correlated with the linear transformation.

[0238] 4. Examples of Techniques and Embodiments

[0239] In this specification, intra-block copy (IBC) is not limited to current IBC technology. It is not used for the current slice / tile / bridge, except for the traditional intra prediction method. Reference samples within a picture / subpicture / picture / other video unit (e.g., CTU row) In one example, the reference sample is an in-loop filter. Reconstruction where filtering processes (e.g., deblocking filter, SAO, ALF) are not invoked This is a construction sample.

[0240] The size of the VPDU can be expressed as vSizeX*vSizeY. Then, vSizeX=vSizeY=min(ctbSizeY,64) and ctbS Let sizeY be the width / height of the CTB. The VPDU is positioned so that the top left corner is at the left edge of the picture. vSize*vSize blocks, where m*vSize, n*vSize are the blocks above. where m and n are integers. Alternatively, vSizeX / vSizeY can be a fixed number such as 64. It is a constant.

[0241] The following detailed inventions should be considered examples to illustrate the general concept. These inventions should not be construed narrowly. Furthermore, these inventions are not intended to be limiting. can be combined in a manner

[0242] IBC signalling and usage

[0243] Let the size of the current block be W curr ×H curr The maximum allowable IBC block size is Size W IBCMax ×H IBCMax Let us express it as:

[0244] 1. IBC coding blocks can be split into multiple TB / T The current IBC code block size may be divided into W curr ×H curr in represent. In one example, W curr is greater than the first threshold, and / or H curr If is greater than a second threshold, splitting can be invoked. is smaller than the current block size. i. In one example, the first and / or second threshold is MaxTbSizeY The maximum transform size that can be represented. b. In one example, recursion is performed until either the width or height of TU / TB is less than or equal to a threshold. It may also call objective division. i. In one example, recursive partitioning is used for inter-coded residual blocks. For example, every time the width is greater than MaxTbSizeY, If the height is greater than MaxTbSizeY, the width is halved. can. ii. Alternatively, the division width and height of the TU / TB are respectively reduced to or below the threshold. This may be called recursive splitting. c. Alternatively, all or some of the multiple TBs / TUs may share the same motion information (e.g., BV ) may be shared.

[0245] 2. Enabling or disabling IBC depends on the maximum conversion size (e.g., this specification MaxTbSizeY in the document). In one example, the dimensions (width and / or height) of the current luminance block or the corresponding The dimensions (width and / or height) of the luma block to be If the maximum transform size in luma samples (if any) is greater than the maximum transform size in luma samples, IBC is disabled. obtain. i. In one example, both the width and height of the block are the maximum transformation in luminance samples. If it is larger than the size, the IBC may be invalidated. ii. In one example, either the width or height of a block is the maximum in luminance samples. For larger than large transform sizes, IBC may be disabled. b. In one example, whether and / or how the IBC signals the use of the IBC. The notification is based on the block dimensions (width and / or height) and / or maximum variable It may depend on the exchange size. i. In one example, the indication of IBC usage is an IBC skip flag (e.g., cu_ skip_flag) may be included in the tile / slice / brick / subpicture. 1) In one example, cu_skip_flag is W curr MaxTbSi Greater than zeY and / or H curr is greater than MaxTbSizeY , may not be signaled. 2) Alternatively, cu_skip_flag can be set to W curr MaxTbSize Y or less and / or H curr If is less than or equal to MaxTbSizeY, the signal You may be notified. 3) In one example, cu_skip_flag is W curr MaxTbSi Greater than zeY and / or H curr is greater than MaxTbSizeY , may be signaled, but must be 0 in the conformance bitstream. . 4) Additionally, for example, if the current tile / slice / brick / subpicture is When the image is a row / slice / brick / subpicture, the above method may be called. ii. In one example, the indication of IBC use is set by an IBC mode flag (e.g., pre d_mode_ibc_flag). 1) In one example, pred_mode_ibc_flag is W curr M axTbSizeY and / or H curr is greater than MaxTbSizeY If .times. ... 2) Alternatively, pred_mode_ibc_flag is set to W curr Max is less than or equal to TbSizeY and / or H curr is less than or equal to MaxTbSizeY If so, it may be signaled. In one example, pred_mode_ibc_flag is curr M axTbSizeY and / or H curr is greater than MaxTbSizeY If ,is larger, it may be signaled, but in the conformance bitstream it is set to 0. It is necessary to iii. In one example, if certain zero-forcing conditions are met, the IBC Code The residual blocks of the coding block may be forced to all zeros. 1) In one example, W curris greater than MaxTbSizeY and / or is H curr If is greater than MaxTbSizeY, the remainder of the IBC coding block The difference block may be all zeroed. 2) In one example, W curr is greater than the first threshold (e.g., 64 / vS izeX) and / or H curr is greater than a second threshold (e.g., 64 / vSi zeY), the residual block of the IBC coding block is forced to all zeros. That's fine. 3) Alternatively, in the above case, the coding block flags (e.g., cu_cb f,tu_cbf_cb,tu_cbf_cr,tu_cbf_luma) signal notification You can skip it. 4) Alternatively, coding block flags (e.g., cu_cbf, tu_c bf_cb,tu_cbf_cr,tu_cbf_luma) signaling remains unchanged. A conforming bitstream must satisfy this flag to be 0, although it may remain Let's say. iv. In one example, coding block flags (e.g., cu_cbf, tu_ cbf_cb,tu_cbf_cr,tu_cbf_luma) signaling may depend on the usage of the IBC and / or the threshold of acceptable IBC size. 1) In one example, the current block is in IBC mode and W curr Max TbSizeY is larger than and / or H curr is greater than MaxTbSizeY If not, coding block flags (e.g., cu_cbf, tu_cbf_cb, t u_cbf_cr,tu_cbf_luma) may be skipped. 2) In one example, the current block is in IBC mode and W curr is the first Threshold (e.g., 64 / vSizeX) and / or H greater than curr is a second threshold. If the value is greater than 64 / vSizeY, the signal of the coding block flags No. notification (e.g., cu_cbf, tu_cbf_cb, tu_cbf_cr, tu_cb f_luma) can be skipped.

[0246] 3. Whether to signal the use of IBC skip mode depends on the size of the current block. modulus (width and / or height) and maximum allowed IBC block size (e.g., vSize X / vSizeY). In one example, cu_skip_flag is curr W IBCMax Greater than If it is too large, and / or H curr H IBCMax If it is greater than It's not necessary. i. Alternatively, W curr W IBCMax is less than or equal to and / or H curr but H IBCMax cu_skip_flag may not be signaled if: . b. Alternatively, the current tile / slice / brick / subpicture is a tile / slice / brick / subpicture, the above method may be called. c. In one example, W IBCMax and H IBCMax are both equal to 64. d. In one example, W IBCMax Set to vSizeX and H IBCMax vSi Set equal to zeY. e. In one example, W IBCMaxis set to the width of the largest transform block, and H IBCM ax is set equal to the height of the largest transform block size.

[0247] On recursive division of IBC / inter-coding blocks

[0248] 4. IBC / Inter-coding block (CU) signals partition information (e.g., block size is larger than the maximum transform block size), multiple TB / T When a CU is divided into Us, instead of signaling motion information once for the entire CU, multiple It is proposed to signal or derive motion information. In one example, all TB / TUs are signaled once using the same AMVP / Me The rge flag may be shared. b. In one example, the motion candidate list may be constructed once for the entire CU, but Different TBs / TUs may be assigned different candidates in the list. In one example, the index of the assigned candidate is It may be coded as follows. c. In one example, the motion candidate list includes motion information of neighboring TBs / TUs within the current CU. It may be constructed without it. i. In one example, the motion candidate list is a list of spatially neighboring blocks for the current CU. Constructed by accessing motion information (adjacent or non-adjacent) That's fine.

[0249] 5. IBC / Inter-coding block (CU) signals partition information It is divided into multiple TBs / TUs (for example, the block size is the maximum transform block size) If the number of reconstructed or pre-constructed TU / TBs is greater than 1, then at least one reconstructed or pre-constructed TU / TB in the second TU / TB is The first TU / TB may be predicted based on the measured sample. a.BV validation may be performed separately for each TU / TB.

[0250] 6. IBC / Inter-coding block (CU) signals partition information It is divided into multiple TBs / TUs (for example, the block size is the maximum transform block size) If the second TU / TB is larger than the second TU / TB, then any reconstructed or predicted samples The first TU / TB cannot be predicted due to the pull. a.BV validation may be performed on the entire coding block.

[0251] On generating predicted blocks for IBC coding blocks

[0252] 7. Use filtered reconstructed reference samples for IBC prediction block generation It is proposed that In one example, the filtered reconstructed reference samples are deblocked It may be generated by applying a bilateral filter / SAO / ALF. b. In one example, for an IBC coding block, at least one reference sample The sample is filtered and at least one reference sample is unfiltered. .

[0253] 8. Filtered and unfiltered reconstructed reference samples Both samples may be used to generate the IBC prediction block. In one example, the filtered reconstructed reference samples are deblocked It may be generated by applying a bilateral filter / SAO / ALF. b. In one example, for predictions for an IBC block, part of the prediction is The rest are unfiltered samples. It may also be from.

[0254] 9. For IBC prediction block generation, the filtered values ​​of the reconstructed reference samples are used Whether to use the filtered or unfiltered values ​​depends on the position of the reconstruction sample. This may depend on the location. In one example, this determination is made based on the CT covering the current CTU and / or reference sample. It may depend on the relative position of the reconstruction sample with respect to U. b. In one example, the reference sample is outside or in a different location from the current CTU / CTB. CTU / CTB and have a large distance (e.g., 4 pixels) to the CTU boundary. If so, the corresponding filtered samples may be used. c. In one example, if the reference sample is within the current CTU / CTB, filtering Unsampled reconstructed samples may be used to generate the prediction block. d. In one example, if the reference sample is outside the current VPDU, the corresponding filter A filtered sample may be used. e. In one example, if the reference sample is within the current VPDU, it is filtered. The unmodified reconstructed samples may be used to generate the prediction block.

[0255] About MVD coding

[0256] Motion vectors (or block vectors), or motion vector differences or block vectors The vector difference is expressed as (Vx, Vy).

[0257] The absolute value of the component of the motion vector represented by AbsV (e.g., Vx or Vy) is The first part is equal to AbsV - ((AbsV >> N) << N) (e.g., AbsV & (1 << N ), the least significant N bits), and the other part is (AbsV >> N) (e.g., the remaining most significant bits) and is proposed to be divided into two parts. Each part may be coded individually. a. In one example, each component of one motion vector or the difference of motion vectors may be coded separately. i. In one example, the first part is coded with a fixed - length coding, e.g., N bits. 1) Alternatively, a first flag may be coded to indicate whether the first part is 0. a. Further alternatively, if not, the value obtained by subtracting 1 from the value of the first part is coded. ii. In one example, the second part having the sign information of MV / MVD may be coded with the current MVD coding method. b. In one example, the first part of each component of one motion vector or the difference of motion vectors may be coded together, and the second part may be coded separately. <0OO1888>i. In one example, the first parts of Vx and Vy may be formed to be a new positive value in 2N bits ( e.g., (Vx << N)+Vy)). 1) In one example, the first flag may be coded to indicate whether the new positive value is equal to 0. a. Further alternatively, if not, the value obtained by subtracting 1 from the new positive value is coded. ​​​​​​​​​​2) Alternatively, this new positive value can be coded using fixed length coding or exponential Golomb coding. It may be coded by a coding (eg, EG-0th). ii. In one example, each component of one motion vector has sign information of MV / MVD. The second part of the minute may be coded using the current MVD coding method. . c. In one example, the first part of each component of one motion vector is coded together. The first part may be coded together with the second part. d. In one example, N is a positive value such as 1 or 2. e. In one example, N may depend on the precision of the MV used to store the motion vectors. i. In one example, the MV precision used to store the motion vectors is 1 / 2M pixels. If so, set N to M. 1) In one example, the MV precision used to store the motion vectors is 1 / 16 pixel. In some cases, N is equal to 4. 2) In one example, the MV precision used to store the motion vectors is 1 / 8 pixel. If N is 3, then N is equal to 3. 3) In one example, the MV precision used to store the motion vectors is 1 / 4 pixel. If N is 2, then N is equal to 2. f. In one example, N is the current motion vector (e.g., block vector in IBC). This may depend on the MV accuracy of the i. In one example, if the BV is accurate to one pixel, then N may be set to 1.

[0258] About SBT

[0259] 11. Signal maximum CU width and / or height indication to allow sub-block transformations Whether to notify may depend on the maximum allowed transform block size. a. In one example, the maximum CU width and maximum CU height to allow for sub-block transformations but 64 or 32 luma samples (e.g. sps_sbt_max_size_64_ flag) depends on the value of the maximum allowed transform block size (e.g., sps_ma x_luma_transform_size_64_flag). In one example, sps_sbt_max_size_64_flag is _sbt_enabled_flag and sps_max_luma_transform _size_64_flag is signaled only if both _size_64_flag and _size_64_flag are true.

[0260] About LFNST matrix indexing

[0261] 12. Using LFNST / LFNST matrix indexing (e.g., lfnst_idx) It is proposed to use two contexts to code the The choice is based purely on the use of the partition tree. g. In one example, for a single tree partition, one context is used. In the case of dual trees, a different context is utilized regardless of the dependency on the primary transformation. do.

[0262] 13. Based on the use of dual trees and / or slice types and / or color components Use of LFNST / LFNST matrix index (e.g., lfnst_idx) It is suggested to code

[0263] 5. Additional Exemplary Embodiments

[0264] [ka]

[0265] 5.1 Embodiment #1

[0266] [Table 13]

[0267] 5.2 Embodiment #2

[0268] [Table 14]

[0269] FIG. 20 illustrates an exemplary video processing system in which various techniques disclosed herein may be implemented. 2000. Various implementations of the modules of system 2000 are shown. The system 2000 may include an input for receiving video content. The video content may include a video output unit 2002. The video content may be in a raw or uncompressed format. , for example, may be received as 8 or 10 bit multi-module pixel values, or may be compressed or The input unit 2002 may receive the network signal in a coding format. It may refer to a network interface, a peripheral bus interface, or a memory interface. Examples of network interfaces are Ethernet, passive optical network, Wired interfaces such as Passive Optical Network (PON) and Wi-Fi or This includes wireless interfaces such as cellular interfaces.

[0270] The system 2000 may implement the various coding or encoding methods described herein. The coding module 2004 may include a coding module 2004 that can The coder 2004 is a block diagram of the input unit 2002 and the output of the coding module 2004. The average bit rate of the video may be reduced to generate a coded representation of the video. This coding technique is sometimes called video compression or video transcoding. The output of the coding module 2004 is represented by module 2006. , may be stored or transmitted via a connected communication. 2. The bitstream (or codec) of the video received, stored or transmitted The representation is used by the module 2008 to display the interface unit. The bitstream may generate pixel values ​​or displayable images that are sent to the host 2010. The process of generating a user-viewable image from a representation is called image expansion. Furthermore, certain video processing operations are sometimes referred to as "coding" operations or tools. However, the coding tools or operations used in the encoder are not necessarily the same as the corresponding decoding tools or operations. It is understood that the operation that reverses the result of coding is performed by the decoder. Let's do it.

[0271] An example of a peripheral bus interface unit or a display interface unit is a Universal Serial Bus (USB) or High-Definition Multimedia Interface (HDMI) I (registered trademark) or DisplayPort, etc. Examples of interfaces include Serial Advanced Technology Attachment (SATA), PCI, The technology described in this specification is applicable to mobile phones, laptops, etc. computer, smartphone, or other device capable of digital data processing and / or image display The present invention may be implemented in a variety of electronic devices, such as a mobile device.

[0272] 21 is a block diagram of a video processing device 2100. The device 2100 is The device 2100 may be used to implement one or more of the methods described above. , tablets, computers, Internet of Things (IoT) receivers, etc. The device 2100 may include one or more processing units 2102, one or more memories 2104, and and video processing hardware 2106. 2 may be configured to implement one or more of the methods described herein. 2104 are used to implement the methods and techniques described herein. The video processing hardware 2106 may be used to store data and code. may be used to implement the techniques described herein in hardware circuitry.

[0273] In the embodiment of the present invention, the following solution can be implemented as a preferred solution.

[0274] The following solutions combine additional techniques described in the previous chapter (e.g., item 1). Both may be implemented.

[0275] 1. A video processing method (e.g., method 2200 shown in FIG. 22), comprising: Coding representation of video blocks in the video domain based on copy tools and video block coding Allows splitting of video blocks into multiple transform units for transformation to and from the video block. and determining whether the video block is a video block or not, the determination being made based on a coding condition of the video block. , this coding representation omits signaling of this division, and and performing a transformation based on the division (2204).

[0276] 2. The method according to Solution 1, wherein the coding conditions correspond to the dimensions of the video blocks. method.

[0277] 3. The width of the video block is greater than a first threshold, thereby permitting the division. or the height of the video block is determined to be greater than a second threshold. Therefore, the method according to Solution 2 is determined to be permitted.

[0278] 4. The first threshold and / or the second threshold are the maximum transformation for the image domain. The size is as described in Solution 3.

[0279] 5. The method according to any one of solutions 1 to 4, wherein the division is applied recursively to the video blocks. How to post.

[0280] The following solutions combine additional techniques described in the previous chapter (e.g., item 2). Both may be implemented.

[0281] 6. A video processing method, comprising: coding representation of a video block in a video domain and a video block; For conversion to and from a block, the transformation of the video block is performed based on the maximum transformation size for the video domain. Determine if the Intra Block Copy (IBC) tool is enabled for and performing a transformation based on the determination.

[0282] 7. The image block is a chroma block, and the maximum transform size is the corresponding luminance block. Correspondingly, IBC is disabled due to the maximum transform size being greater than the threshold. Solution 6:

[0283] 8. Syntax elements in the coding representation specify dimensions or criteria for the video blocks. A solution indicating whether the IBC tool is enabled by meeting the maximum conversion size. The method described in solution 6 or 7.

[0284] 9. If the IBC tool is determined to be enabled, this transformation is performed using zero forcing. Based on a condition, we can force the residual block for a video block to be all zero. 9. The method according to any one of Solutions 6 to 8, further comprising:

[0285] The following solutions combine additional techniques described in the previous chapter (e.g., item 3). Both may be implemented.

[0286] 10. A video processing method, comprising: coding representation of a video block in a video domain; For conversion to and from blocks, the signal communication of the Intra Block Copy (IBC) tool for conversion The first step is to determine whether knowledge is included in the coding representation and to make changes based on this determination. and performing a conversion between the width and / or height of the video block and the video area. Based on the maximum allowable IBC block size for the method.

[0287] 11. The width is greater than the first threshold or the height is greater than the second threshold. Therefore, the method described in Solution 10 does not include signaling.

[0288] The following solutions combine additional techniques described in the previous chapter (e.g., item 4). Both may be implemented.

[0289] 12. A video processing method, comprising: using an intra block copy (IBC) tool: For conversion between coding representation of video blocks in the video domain and video blocks, Determine whether it is allowed to split a block into multiple transform units (TUs) for transformation. and performing a transformation based on the determination, the transformation being performed on the plurality of TUs. The method includes using separate motion information.

[0290] 13. Multiple TUs share the same advanced motion vector predictor or merge mode flag. 13. The method of solution 12, wherein the method is constrained to:

[0291] 14. During the transformation, the motion information for one of the plurality of TUs is 13. The method of claim 12, wherein the determination is made without using motion information of another TU of the

[0292] The following solutions combine additional techniques described in the previous chapter (e.g., item 5). Both may be implemented.

[0293] 15. Solution 1, where a first TU of a plurality of TUs is predicted from a second TU of a plurality of TUs 15. The method according to any one of 2 to 14.

[0294] The following solutions combine additional techniques described in the previous chapter (e.g., item 6). Both may be implemented.

[0295] 16. Prediction of a first TU of the plurality of TUs from a second TU of the plurality of TUs is disabled; The method according to any of Solutions 12-15.

[0296] The following solutions combine additional techniques described in the previous chapter (e.g., item 7). Both may be implemented.

[0297] 17. A video processing method, comprising: coding representation of a video block in a video domain; For conversions to or from a block, the intrablock copy tool is enabled for this conversion. determining that the block is in the block, and performing the conversion using an intra-block copy tool; and making predictions of video blocks using filtered reconstruction samples of the video domain. To do, how to do.

[0298] 18. The filtered reconstructed samples are in-loop The in-loop filter includes samples generated by applying a filter, the in-loop filter being a non-blocking filter. Locked filter, bilateral filter, or sample adaptive offset or adaptation 18. The method of solution 17, which is a loop filter.

[0299] The following solutions combine additional techniques described in the previous chapter (e.g., item 8). Both may be implemented.

[0300] 19. Based on the rules, the prediction is further performed on an unfiltered reconstruction of the image region. The method according to any of Solutions 17-18, selectively using a conjugated sample.

[0301] The following solutions combine additional techniques described in the previous chapter (e.g., item 9). Both may be implemented.

[0302] 20. The method of solution 19, wherein the rule depends on the position of the reconstructed sample. method.

[0303] 21. The position of the reconstructed sample used in the rule is determined by the code of the video block. The relative position of the reconstructed sample with respect to the ing tree unit, as described in Solution 20 The method of.

[0304] The following solutions may be implemented together with the additional techniques described in the items described in the previous chapter (e.g., item 10). May be implemented together with.

[0305] 22. A video processing method including converting a video including a plurality of video blocks and a coding representation of the video, wherein at least some of the video blocks are coded using motion vector information, and the motion vector information is represented in a coding representation as a first part based on the first least significant bit of the absolute value of the motion vector information and as a second part based on the remaining more significant bits higher than the first least significant bit. The method of. The motion vector information is represented in a coding representation as a first part based on the first least significant bit of the absolute value of the motion vector information, and as a second part based on the remaining more significant bits higher than the first least significant bit. As a first part based on the first least significant bit of the absolute value of the motion vector information, and as a second part based on the remaining more significant bits higher than the first least significant bit. The method of.

[0306] 23. Divide the absolute value of the motion vector V represented as AbsV into a first part equal to AbsV - ((AbsV > N) << N) or AbsV & (1 << N), and the second part is equal to (AbsV >> N), where N is an integer, the method described in Solution 22. The method of. The method described in Solution 22, where the second part is equal to (AbsV >> N) and N is an integer.

[0307] 24. The method according to any one of Solutions 22 to 2, where the first part and the second part are separately coded in the coding representation. The method according to any one of Solutions 22 to 23, where the first part and the second part are separately coded in the coding representation.

[0308] 25. The method according to any one of Solutions 22 to 24, where the first part of the x motion vector is coded together at the same position as the first part of the y motion vector. The method according to any one of Solutions 22 to 24, where the first part of the x motion vector is coded together at the same position as the first part of the y motion vector.

[0309] 26. The method according to Solution 25, where the second part of the x motion vector and the second part of the y motion vector are coded separately. The method described in Solution 25, where the second part of the x motion vector and the second part of the y motion vector are coded separately.

[0310] 27. Solutions 22-26 where N is a function of the precision used to store motion vectors. Any of the methods described above.

[0311] The following solutions incorporate additional techniques described in the previous chapter (e.g., item 11). It may also be implemented together with

[0312] 28. For conversion between coding representations of video blocks in the video domain and video blocks, determining whether a sub-block conversion tool is enabled for conversion; performing the conversion based on a determination of a maximum allowable conversion value for the video region; The maximum allowable transformation block size is based on the transformation block size. A video processing method including signaling in this coding representation with

[0313] The following solutions are based on the additional steps described in the previous chapter (e.g., steps 12 and 13). This technique may also be implemented in conjunction with the techniques described above.

[0314] 29. A video processing method, comprising: coding representation of a video block in a video domain; Whether low frequency non-separable transform (LFNST) is used during the conversion to convert to or from the clock. and performing a conversion based on the determination, wherein the determination is Based on the coding conditions applicable to the lock, LFNST and The matrix index of is coded in the coding representation using two contexts. The way it is done.

[0315] 30. The coding conditions may include a partition tree to be used, or Solution 29 includes a slice type of the block, or a color component identity of the video block. The method described.

[0316] 31. The video region includes a coding unit, and the video block is a 31. The method according to any of Solutions 1 to 30, wherein the luminance or chrominance component of

[0317] 32. The transform generates the bitstream representation from the current video block. 32. The method of any of Solutions 1 to 31, comprising:

[0318] 33. The conversion is performed by converting the sample of the current video block from the bitstream representation to 32. The method of any of Solutions 1-31, comprising generating

[0319] 34. A processor configured to implement the method described in any one or more of solutions 1 to 33. A video processing device comprising a processing device.

[0320] 35. A processor implementing the method according to any one or more of solutions 1 to 33 when executed. A computer-readable medium having stored thereon code for causing a

[0321] 36. A method, system or apparatus described herein.

[0322] FIG. 23 is a flow chart illustrating a video processing method 2300 in accordance with the present technology. 300, in step 2310, compares the current block in the image area of ​​the image with the Splitting the current block into multiple transformation units for transformation to and from the coding representation This includes determining whether to allow the current block based on its characteristics. In the block representation, the signaling of the split is omitted. , and performing a transformation based on that determination.

[0323] In some embodiments, the characteristics of the current block include the dimensions of the current block. In some embodiments, if the width of the current block is greater than a first threshold, the division is In some embodiments, the height of the current block is called to transform If it is greater than the threshold, the split is invoked for the transformation. In this case, dividing the current block is performed by dividing the transform units such that one dimension of the transform units is equal to or smaller than a threshold. In some embodiments, this includes recursively splitting the current block until The width of the current block is then recursively halved until it is equal to or less than the first threshold. In some embodiments, if the height of the current block is less than or equal to the second threshold, The height of the current block is recursively divided in half until In the block, the width of the current block is equal to or less than the first threshold value and the height of the current block is equal to or less than the second threshold value. The current block is recursively split until the current block is below the threshold. In some embodiments, the first threshold is equal to the maximum transform size for the image area. The second threshold is equal to the maximum transform size for the image area.

[0324] In some embodiments, at least a subset of the plurality of transform units In some embodiments, the same motion information is shared between the current block and the The advanced motion vector prediction mode or contains syntax flags for merge mode. In some embodiments, the saturation is determined multiple times for the current block in the loading representation. In this case, we build a motion candidate list once for the current block and then apply it to multiple transform units. are assigned different candidates in the motion candidate list. Therefore, the indices of the different candidates are signaled in the coding representation. In this embodiment, for a transform unit of a multiple transform unit, the motion candidate list is currently The motion information of neighboring transform units within the block is not used. In this case, the motion candidate list is generated by using the motion information of blocks spatially neighboring the current block. In some embodiments, the first conversion unit is constructed using a second conversion unit. The reconstructed data is based on at least one reconstructed sample in the set. In the example, a verification check is performed separately for each of the multiple transformation units. In some embodiments, the first transform unit may generate reconstructed samples in other transform units. In some embodiments, after reconstructing the blocks, A validity check is performed on the block.

[0325] Figure 24 is a flow chart illustrating a method 2400 of video processing in accordance with the present technology. 2400, in step 2410, performs a coding operation on the current block of the image area of ​​the image and the image. Filtered reconstructed reference samples in the video domain for conversion to and from the scalar representation Based on this, the Intra Block Copy (IBC) mode is used to copy the current block The method 2400 includes determining the measurement sample. This includes performing a conversion based on the

[0326] In some embodiments, the filtered reconstructed reference sample is a non-blocking Locked filtering operation, bilateral filtering operation, sample adaptive offset At least one of the following: adaptive (SAO) filtering operation or adaptive loop filtering operation In some embodiments, the I Using the BC model to determine the predicted samples of the current block is a In some embodiments, the method is further based on an unfiltered reference sample. , based on the filtered reconstructed reference sample, in the first part of the current block Determine the predicted sample in the filter based on the unfiltered reconstructed reference sample. In some embodiments, the prediction sample for the second portion of the current block is determined by the second portion of the current block. In the filtered reconstructed reference sample or unfiltered The decision whether to use the current sample or the reconstructed reference sample to determine the predicted sample for the block depends on the For a coding tree unit or a sample of a coding tree unit, In some embodiments, the filtered The reconstructed reference sample is the predicted sample located outside the current coding tree unit. In some embodiments, if The filtered reconstructed reference samples are the ones that the predicted samples are based on coding tree units. If the coding tree unit is located within the boundary of the coding tree unit and has a distance D of 4 pixels or more from the boundary of the coding tree unit, In some embodiments, the filter is used to reconstruct the predicted samples. The unsorted reconstructed reference sample is the one that predicts the current coding tree. If it is located within a unit, it is used to reconstruct the predicted sample. In an embodiment, the filtered reconstructed reference sample is a predicted sample relative to the current hypothesis. If it is located outside the virtual pipeline data unit, then it is necessary to use the In some embodiments, the unfiltered reconstruction reference If the predicted sample is located within the current virtual pipeline data unit, is used to reconstruct the predicted samples.

[0327] FIG. 25 is a flow chart illustrating a video processing method 2500 in accordance with the present technology. 500, in step 2510, performs a comparison between the current block of video and the coded representation of the video. To convert, we use skip for the Intra Block Copy (IBC) coding model. Determine whether the syntax element indicating the use of the 'pick mode' is included in the coding expression according to the rules. This rule specifies that the signaling of syntax elements determines the size and / or or the maximum allowed value for a block coded using the IBC coding model. The method 2500 also provides that the determination is based on the volume size in operation 2520. This includes performing conversions based on the specified

[0328] In some embodiments, the rule is that the width of the current block is greater than the maximum allowed width. If the syntax element is not present in the coding representation, the syntax element is omitted. In some embodiments, the rule is: if the height of the current block is greater than the maximum allowed height, In this case, the syntax element is omitted in the coding representation. In an embodiment, the rule is that when the width of the current block is less than or equal to the maximum allowed width, In some embodiments, the coding representation defines a syntax element. The rule is that when the height of the current block is equal to or less than the maximum allowed height, the syntax element It is defined that the coding expression includes:

[0329] In some embodiments, the current block is in the image area of ​​the image, and this rule Applicable if the image region contains I tiles, I slices, I bricks, or I subpictures In some embodiments, the maximum allowable width or height is 64. In some embodiments, the maximum allowable width or maximum allowable height is determined by the virtual pipeline data. Equal to the unit dimension. In some embodiments, the maximum allowable width or maximum allowable height. is equal to the maximum dimension of the conversion unit.

[0330] FIG. 26 is a flowchart illustrating a video processing method 2600 according to the present technology. The method 2600, in step 2610, performs a multiplication of a current block of video and a coding representation of the video. For the transformation between determining at least one context for coding the selected index; This LFNST coding model uses forward linear transform and quantization step during encoding. Applying a forward quadratic transform between the dequantization step and the dequantization step during decoding Applying an inverse secondary transform between a forward secondary transform and an inverse primary transform The size of the directional secondary transformation is smaller than the size of the current block. The context is the performance of the current block, without taking into account the forward or inverse linear transformation. The method 2600 also includes, in operation 2620, determining whether the and then performing the conversion according to that determination.

[0331] In some embodiments, the partition type is a single tree partition. If there is one, only one context is used to code the index. In some embodiments, the partition type is a dual tree partition. When , we use two contexts to code the index.

[0332] In some embodiments, a low frequency non-separable transform (LFNST) coding model The index showing the usage is the feature related to the block, the party related to the current block. Coding representation based on features including segmentation type, slice type or color components Included in.

[0333] FIG. 27 is a flowchart illustrating a video processing method 2700 according to the present technology. The method 2700, in step 2710, performs a process of determining a current block of an image region of an image and a code of the image. For conversion to and from the CG representation, the maximum conversion unit size applicable to the video domain is used. , whether the intra-block copy (IBC) coding model is enabled The method 2700 also includes, in operation 2720, determining, in accordance with the determination: This includes performing a transformation.

[0334] In some embodiments, the dimensions of the current block or the size of the current block If the dimension of the luma block is larger than the size of the largest transform unit, the IBC coding model In some embodiments, the width of the current block or the current block If the width of the luminance block corresponding to the block is greater than the width of the maximum conversion unit, the IBC codec In some embodiments, the current block height or If the height of the luminance block corresponding to the current block is greater than the maximum transformation height, the IBC control In some embodiments, the coding model is invalidated. The method for signaling the use of the IBC coding model is based on the dimensions of the current block and the maximum Based on the size of the large transform unit. In some embodiments, the IBC coding model Signaling the use of the rule is a syntax that indicates skip mode in the IBC coding model. In some embodiments, the syntax element includes an I type in the coding representation. contained in an I-slice, I-brick, or I-subpicture.

[0335] In some embodiments, signaling the use of the IBC coding model comprises: In some embodiments, the syntax element indicates the IBC mode. is greater than the width of the maximum conversion unit, or the height of the current block is greater than the maximum conversion unit If the height is greater than 1, the syntax element is omitted in the coding representation. In this configuration, the width of the current block is less than or equal to the width of the maximum transform unit, or If the block height is less than or equal to the maximum transform unit height, the syntax element is a coding representation In some embodiments, the width of the current block is the width of the largest transform unit. If it is greater than or the height of the current block is greater than the maximum translation unit height, The syntax element is included in the coding representation and is set to 0.

[0336] In some embodiments, the IBC coding model is enabled for the current block. When enabled, the sample in the residual block corresponding to the current block is is set to 0. In some embodiments, the rule is that the width of the current block is the maximum transform The width of the unit is greater than the maximum conversion unit height, or the height of the current block is greater than the maximum conversion unit height. If it is greater than 0, the sample is set to 0. The rule is that the width of the current block is greater than the first threshold, or the height of the current block is greater than the second threshold. It specifies that if it is greater than a threshold of 2, the sample is set to 0. In this case, the first threshold is 64 / vSizeX, where vSizeX is the virtual pipeline data. In some embodiments, the second threshold is 64 / vSizeY. where vSizeY is the height of the virtual pipeline data unit. In this case, the syntax flags for the coding block are omitted in the coding representation. In some embodiments, the syntax flags for a coding block are It is set to 0 in the binning representation.

[0337] In some embodiments, for a coding block in the coding representation The signaling of syntax flags is based on the use of the IBC coding model. In a state where the IBC coding model is enabled for the current block, if the size of the block is larger than the size of the maximum transform unit, the syntax flag in the coding representation can be omitted. In some embodiments, when the IBC coding model is enabled for the current block and the size of the block is larger than the threshold related to the virtual pipeline data unit , the syntax flag is omitted in the coding representation.

[0338] FIG. 28 is a flowchart showing a video processing method 2800 according to the present technology. Method 2 800 includes, in step 2810, determining to divide the absolute value of the component of the motion vector for the current block into two parts in order to transform the current block of the video and the coding representation of the video. Represent the motion vector as (Vx, Vy), represent its component as V i, where Vi is either Vx or Vy. The first part of the two parts is equal to |Vi | - ((|Vi| >> N) << N), and the second part of the two parts is equal to |Vi| >> N, where N is a positive integer. The two parts are coded separately in the coding representation. Method 2800 also includes, in operation 2820, performing the transformation according to the determination.

[0339] In some embodiments, the two components Vx and Vy are signaled separately in the coding representation. In some embodiments, the first part is coded with a fixed length of N bits. [[ID=�5]] In some embodiments, the first part of component Vx and the first part of component Vy are coded together, and the second part of component Vx and the second part of component Vy are coded together. They are coded separately. In some embodiments, the first part of component Vx and the first part of component Vy are coded as a value having a length of 2N bits. In some embodiments, the value is equal to (Vx << N) + Vy). In some examples it is coded using a fixed-length coding process or a Golomb coding process.

[0340] In some embodiments, the first part of component Vx and the first part of component Vy are coded together, and the second part of component Vx and the second part of component Vy are coded together. In some examples, a syntax flag is included in the coding representation to indicate whether the first part is equal to 0. In some examples, the first part has a value of K, where K ≠ 0, and the value of (K - 1) is coded in the coding representation. In some embodiments, the second part of each component having sign information is coded using a motion vector differential coding process. In some examples, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, the second part of each component having sign information is coded using a motion vector differential coding process. In some examples, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, the second part of each component having sign information is coded using a motion vector differential coding process. In some examples, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, the second part of each component having sign information is coded using a motion vector differential coding process. In some examples, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, the second part of each component having sign information is coded using a motion vector differential coding process. In some examples, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, N is 1 or 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, N is determined by the accuracy of the motion vector used for storing the motion vector data. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, when the accuracy of the motion vector is 1 / 16 pixel, N is 4. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is / 4 pixel, N is 2. In some embodiments, when the accuracy of the motion vector is 1 / 8 pixel, N is 3. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is 2. In some embodiments, when the accuracy of the motion vector is 1 / 4 pixel, N is / 2. In some embodiments, N is determined based on the motion vector accuracy of the current motion vector. In some embodiments, when the accuracy of the motion vector is 1 pixel, N is 1.​ is 1.

[0341] FIG. 29 is a flowchart illustrating a video processing method 2900 according to the present technology. The method 2900, in step 2910, performs a multiplication of a current block of video and a coding representation of the video. For conversion to and from, the size of the current block is calculated based on the maximum allowable dimensions of the conversion block. This includes determining information about the maximum size of the current block that allows block-to-block transformation. The method 2900 also includes, at operation 2920, performing a transformation according to the determination. include.

[0342] In some embodiments, allowing sub-block transformations on the current block The maximum size of the current block corresponds to the maximum allowed size of the transformation block. In some embodiments, the maximum size of the current block is 64 or 32. In the example, the first syntax flag in the slice parameter set is the sub-block transform is enabled, and the second syntax flag in the slice parameter set is , indicates that the maximum allowed size of a conversion block is 64, the maximum size of the current block A syntax flag indicating this is included in the coding representation.

[0343] In some embodiments, this transformation generates a coding representation from the current block. In some embodiments, this conversion includes converting the coding representation to a current This involves generating samples for the block.

[0344] In the above solution, the transformation is performed by changing the previous decision step during the encoding or decoding operation. using the results of a step (e.g., using a particular encoding or decoding step, or not) to arrive at the transformation result.

[0345] Some embodiments of the disclosed technology may include enabling video processing tools or modes. In one example, determining or judging whether a video processing tool or mode is enabled. When a video block is processed, the encoder uses this tool or mode to process one video block. use or implement a tool, but results in The resulting bitstream does not necessarily need to be modified; The conversion from the video to a bitstream representation is performed by a video processing tool or Use this video processing tool or mode when the tool or mode is enabled. When a video processing tool or mode is enabled in The bitstream is then modified based on the video processing tool or mode. process the system, i.e., the video processing tools or The mode is used to convert from a bitstream representation of video to blocks of video.

[0346] Some embodiments of the disclosed technology may include disabling video processing tools or modes. In one example, determining or judging that a video processing tool or mode is disabled. When enabled, the encoder converts blocks of video into a bitstream representation of the video. In another example, do not use this tool or mode when using a video processing tool or If the mode is disabled, the decoder will You know that the bitstream has not been modified using any video processing tools or modes. and process the bitstream.

[0347] The disclosed and other solutions, examples, embodiments, modules, The implementation of the rules and functional operations is within the scope of the structures disclosed herein and their structural equivalents. any digital electronic circuit, including computer software, firmware, or may be implemented in hardware, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products. , for example to be implemented by or to control the operation of a data processing apparatus. one of computer program instructions encoded on a computer-readable medium for controlling The computer-readable medium may be implemented as one or more modules. -readable storage device, machine-readable storage substrate, memory device, material that provides a machine-readable propagated signal The "data processing device" may be a composition of the above, or a combination of one or more of these. The term may refer to, for example, a programmable processing device, a computer, or a plurality of processing devices, or any apparatus or device for processing data, including a computer; This device includes a computer program execution environment in addition to the hardware. code that creates the system, e.g., processor firmware, protocol stacks, database management the physical system, operating system, or a combination of one or more of these. A propagated signal may include an artificially generated signal, e.g., a machine-generated signal. A digital signal is an electrical, optical, or electromagnetic signal that encodes information for transmission to an appropriate receiving device. It is generated to

[0348] Computer programs (programs, software, software applications) , script, or code) is a language that is written in a compiled or interpreted It can be written in any form of programming language, including standard Modules suitable for use as standalone programs or in any computing environment Deploy in any form, including as modules, components, subroutines, or other units. A computer program does not necessarily have to access a file in a file system. A program may not respond to requests from other programs or files that hold data. recorded in a part (e.g., one or more scripts stored in a markup language document) The program may be stored in a single file dedicated to that program, or in multiple files. A coordination file (e.g., one or more modules, subprograms, or pieces of code) A computer program may be stored in a file (a file storing the program). A single computer located at a site or a communication network distributed across multiple sites It can also be deployed to run on multiple computers interconnected by do.

[0349] The processes and logic flows described herein operate on input data and produce output. execute one or more computer programs to perform functions by creating The processing and logic flow can be performed by one or more programmable processing units. It also includes application-specific logic circuits, such as FPGAs (Field Programmable Gate Arrays). This can be done by a chip array (chip array) or an ASIC (application specific integrated circuit), and the device can also It can also be implemented as special purpose logic circuitry.

[0350] Suitable processing devices for the execution of a computer program include, for example, both general purpose and special purpose microprocessors. Both the processor and any one or more processors of any kind of digital computer Generally, a processing unit may have a read-only memory or a random access memory. The essential elements of a computer are the a processing unit for executing instructions and one or more memory devices for storing instructions and data Generally, a computer has one or more mass storage devices for storing data. The device may include, or may be a magnetic, magneto-optical, or optical disk. to receive data from or transfer data to these mass storage devices. However, a computer may be operatively coupled to such a device. It is not necessary to have a computer suitable for storing computer program instructions and data. Computer readable media includes all forms of non-volatile memory, media, and memory devices. Includes, for example, EPROM, EEPROM, flash storage, magnetic disks, e.g. Internal hard disk or removable disk, magneto-optical disk, and CD-ROM and semiconductor storage devices such as DVD-ROM disks. may be supplemented by or incorporated into special purpose logic circuits. It may be included.

[0351] This patent specification contains many details, which may not be sufficient to encompass the scope of any subject matter or the scope of any claim. The present invention should not be construed as limiting the scope of the present invention, but rather as being specific to particular embodiments of particular technologies. The description of the features that may be present in the present patent document should be interpreted as a description of the features that may be present in the present patent document. Certain features described in the context may be implemented in combination in one example. In addition, various features described in the context of one example may be used separately in multiple embodiments. or in any suitable subcombination. may be described above as acting in combination and may be initially claimed as such However, one or more features from the claimed combination may, in some cases, be removed from the combination. The claimed combination may be a subcombination or a sub-combination. It may be directed towards variations of the nation.

[0352] Similarly, although operations may be shown in a particular order in the figures, this is not to be construed as a guarantee that a desired result will be achieved. that such actions be performed in the particular order or sequence shown, in order to It should not be understood as requiring that all actions be performed. Also, the separation of the various system components in the examples described in this patent specification It should not be understood that all embodiments require such separation.

[0353] Only some implementations and examples are described and illustrated in this patent document. Other embodiments, extensions, and variations are possible based on the content provided.

Claims

1. 1. A method for processing video data, comprising: For transforming a current transform block of an image to a bitstream of the image, determining whether a width of the current transform block is greater than a first threshold or a height of the current transform block is greater than a second threshold; invoking a split operation on the current transform block if the width of the current transform block is greater than a first threshold or the height of the current transform block is greater than the second threshold; performing the transformation based on at least two or more transformation blocks; signaling of the invocation of the splitting operation and signaling of the splitting operation are omitted in the bitstream; the first threshold and the second threshold are equal to a first value; calling the splitting operation if the first value is equal to 32 and the current transformed block is derived based on a coding block coded with a first prediction mode in which prediction samples are derived from a block of sample values ​​of the same video region as determined by a block vector; the width of the coding block and the height of the coding block are equal to or less than 64 and greater than 32; The splitting operation is recursively invoked until the width of the current transformation block is less than or equal to the first threshold and the height is less than or equal to the second threshold. method.

2. the first value defines a maximum transform size of luma samples; The method of claim 1.

3. In the dividing operation, If the width of the current transform block is greater than the first value, then dividing the width of the current transform block in half; If the height of the current transform block is greater than the first value, then invoking dividing the height of the current transform block in half. The method of claim 2.

4. the converting includes encoding the video into a bitstream; 4. The method according to any one of claims 1 to 3.

5. the converting includes decoding the video from the bitstream.

4. The method according to any one of claims 1 to 3.

6. 1. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions stored thereon, The instructions, when executed by the processing unit, cause the processing unit to: For transforming a current transform block of an image to a bitstream of the image, determining whether a width of the current transform block is greater than a first threshold or a height of the current transform block is greater than a second threshold; invoking a split operation on the current transform block if the width of the current transform block is greater than a first threshold or the height of the current transform block is greater than the second threshold; performing the transformation based on at least two or more transformation blocks; signaling of the invocation of the splitting operation and signaling of the splitting operation are omitted in the bitstream; the first threshold and the second threshold are equal to a first value; calling the splitting operation if the first value is equal to 32 and the current transformed block is derived based on a coding block coded with a first prediction mode in which prediction samples are derived from a block of sample values ​​of the same video region as determined by a block vector; the width of the coding block and the height of the coding block are equal to or less than 64 and greater than 32; The splitting operation is recursively invoked until the width of the current transformation block is less than or equal to the first threshold and the height is less than or equal to the second threshold. Device.

7. A non-transitory computer-readable storage medium storing instructions, comprising: The instructions may include: For transforming a current transform block of an image to a bitstream of the image, determining whether a width of the current transform block is greater than a first threshold or a height of the current transform block is greater than a second threshold; invoking a split operation on the current transform block if the width of the current transform block is greater than a first threshold or the height of the current transform block is greater than the second threshold; performing the transformation based on at least two or more transformation blocks; signaling of the invocation of the splitting operation and signaling of the splitting operation are omitted in the bitstream; the first threshold and the second threshold are equal to a first value; calling the splitting operation if the first value is equal to 32 and the current transformed block is derived based on a coding block coded with a first prediction mode in which prediction samples are derived from a block of sample values ​​of the same video region as determined by a block vector; the width of the coding block and the height of the coding block are equal to or less than 64 and greater than 32; The splitting operation is recursively invoked until the width of the current transformation block is less than or equal to the first threshold and the height is less than or equal to the second threshold. A non-transitory computer-readable storage medium.

8. 1. A method for storing a video bitstream, comprising: determining, for a current transform block of the image, whether a width of the current transform block is greater than a first threshold or a height of the current transform block is greater than a second threshold; invoking a split operation on the current transform block if the width of the current transform block is greater than a first threshold or the height of the current transform block is greater than the second threshold; generating the bitstream based on at least two or more transform blocks; storing the bitstream on a non-transitory computer-readable recording medium; signaling of the invocation of the splitting operation and signaling of the splitting operation are omitted in the bitstream; the first threshold and the second threshold are equal to a first value; calling the splitting operation if the first value is equal to 32 and the current transformed block is derived based on a coding block coded with a first prediction mode in which prediction samples are derived from a block of sample values ​​of the same video region as determined by a block vector; the width of the coding block and the height of the coding block are equal to or less than 64 and greater than 32; The splitting operation is recursively invoked until the width of the current transformation block is less than or equal to the first threshold and the height is less than or equal to the second threshold. method.