Method and device for video processing and medium

By obtaining the NSPT information related to the block size or intra mode of the video block and performing conversion, the problem of insufficient quality and efficiency of video encoding and decoding in the prior art is solved, and a more efficient and high-quality encoding and decoding process is achieved.

CN120226358APending Publication Date: 2025-06-27DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380079940.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-18
Filing Date
2023-11-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have shortcomings in improving the quality and efficiency of encoding and decoding, especially when dealing with different block sizes and intra-frame modes, it is difficult to effectively utilize the indivisible main transform (NSPT) related information.

Method used

A video processing method is proposed to perform conversion based on the indivisible main transform (NSPT) information related to the current block size or intra-frame mode to improve the efficiency and quality of video encoding and decoding.

Benefits of technology

This method can improve the encoding and decoding efficiency and quality, and has significant advantages over conventional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120226358A_ABST
    Figure CN120226358A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method includes: for a conversion between a current block of a video and a bitstream of the video, acquiring information related to an imseparable main transform (NSPT) applied to the current block, the information depending on at least one of a block size of the current block or an intra mode for the current block; and performing a conversion based on the information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to video encoding and decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding and decoding quality of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: obtaining information related to an Inseparable Normalized Transform (NSPT) applied to a current block of a video for conversion between the current block of the video and a bitstream of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and performing the conversion based on the information.

[0005] According to the method of the first aspect of the present disclosure, the information related to the NSPT depends on the block size of the current block and / or the intra mode of the current block. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency and the encoding and decoding quality.

[0006] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing for a video. The method includes: obtaining information related to an Inseparable Normalized Transform (NSPT) applied to a current block of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and generating a bitstream based on the information.

[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: obtaining information related to a non-separable principal transform (NSPT) applied to a current block of a video, the information depending on at least one of a block size of the current block or an intra mode for the current block; generating a bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present invention content is provided to introduce in a simplified form a selection of concepts further described in the detailed implementation below. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Brief Description of the Drawings

[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram showing an exemplary video codec system according to some embodiments of the present disclosure is shown;

[0013] Figure 2 A block diagram showing a first exemplary video encoder according to some embodiments of the present disclosure is shown;

[0014] Figure 3 A block diagram showing an exemplary video decoder according to some embodiments of the present disclosure is shown;

[0015] Figure 4 The positions of spatial Merge candidates are shown;

[0016] Figure 5 Candidate pairs considered for redundancy checking of spatial Merge candidates are shown;

[0017] Figure 6 Motion vector scaling for temporal Merge candidates is shown;

[0018] Figure 7 Candidate positions C0 and C1 for temporal Merge candidates are shown;

[0019] Figure 8 MMVD search points are shown;

[0020] Figure 9 An extended CU region used in BDOF is shown;

[0021] Figure 10 A symmetric MVD mode is shown;

[0022] Figure 11 Shows an affine motion model based on control points;

[0023] Figure 12 Shows the affine MVF of each sub-block;

[0024] Figure 13 Shows the position of the inherited affine motion prediction value;

[0025] Figure 14 Shows control point motion vector inheritance;

[0026] Figure 15 Shows the position for constructing candidate positions of the affine Merge mode;

[0027] Figure 16 Is an illustration of the use of motion vectors for the proposed combination method;

[0028] Figure 17 Shows the sub-block MV VSB and pixel Δv(i,j);

[0029] Figure 18A Shows the spatial neighboring blocks used by ATVMP;

[0030] Figure 18B Shows the derivation of the sub-CU motion field by applying motion displacements from spatial neighbors and scaling the motion information of the corresponding co-located sub-CU;

[0031] Figure 19 Shows the extended CU region used in BDOF;

[0032] Figure 20 Shows motion vector refinement on the decoding side;

[0033] Figure 21 Shows the top neighboring block and left neighboring block used in CIIP weight derivation;

[0034] Figure 22 Shows an example of GPM partitioning grouped at the same angle;

[0035] Figure 23 Shows the unidirectional prediction MV selection for the geometric segmentation mode;

[0036] Figure 24 Shows an exemplary generation of the hybrid weight w0 using the geometric segmentation mode;

[0037] Figure 25 Shows the spatial neighboring blocks for deriving spatial Merge candidates;

[0038] Figure 26Shows template matching performed on a search area around the initial MV;

[0039] Figure 27 Shows the diamond-shaped area in the search area;

[0040] Figure 28 Shows the frequency responses of the interpolation filter and the VVC interpolation filter at the half-pixel phase;

[0041] Figure 29 Shows the template and the reference samples of the template in the reference picture;

[0042] Figure 30 Shows the reference samples of the template and the template for a block with sub-block motion using the motion information of the sub-blocks of the current block;

[0043] Figure 31 Shows the filling candidates for replacing the zero vectors in the IBC list;

[0044] Figure 32 Shows the IBC reference region depending on the current CU position;

[0045] Figure 33 Shows the reference region for IBC when CTU(m, n) is encoded / decoded. The blue block represents the current CTU; the green block represents the reference region; and the white block represents the invalid reference region;

[0046] Figure 34 Shows the first HPT and the second HPT;

[0047] Figure 35 Shows the spatial neighbors for deriving the affine Merge candidate / AMVP candidate;

[0048] Figure 36 Shows the constructed affine Merge candidate / AMVP candidate from non-adjacent neighbors to the first type;

[0049] Figure 37 Shows the low-frequency non-separable transform (LFNST) process;

[0050] Figure 38 Shows the SBT position, type, and transform type;

[0051] Figure 39 Shows the ROI for LFNST 16;

[0052] Figure 40 Shows the ROI for LFNST 8;

[0053] Figure 41 Shows the discontinuity measurement;

[0054] Figure 42 Shows a design using NSPT and LFNST;

[0055] Figure 43 Shows airspace GPM candidates;

[0056] Figure 44 Shows a GPM template;

[0057] Figure 45 Shows a GPM mix;

[0058] Figure 46 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and

[0059] Figure 47 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.

[0060] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description

[0061] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.

[0062] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0063] As used in this disclosure, the terms "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0064] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0065] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Exemplary Environment

[0066] Figure 1 is a block diagram showing an exemplary video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0067] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces that receive video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0068] Video data can include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be directly transmitted to the destination device 120 via the I / O interface 116 over the network 130A. The encoded video data can also be stored on the storage medium / server 130B for access by the destination device 120.

[0069] The destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 can include a receiver and / or a modulator. The I / O interface 126 can obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 can decode the encoded video data. The display device 122 can display the decoded video data to the user. The display device 122 can be integrated with the destination device 120 or can be external to the destination device 120, which is configured to interface with an external display device.

[0070] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.

[0071] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 can be Figure 1 an example of the video encoder 114 in the system 100 shown.

[0072] The video encoder 200 can be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video encoder 200. In some examples, a processor can be configured to execute any or all of the techniques described in the present disclosure.

[0073] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.

[0074] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0075] In addition, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are shown separately in the Figure 2 examples.

[0076] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0077] The mode selection unit 203 may select, for example, one coding mode (intra coding or inter coding) from multiple coding modes based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select an Intra-Inter Combined Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0078] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.

[0079] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on a current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to portions of a picture composed of macroblocks independent of macroblocks within the same picture.

[0080] In some examples, the motion estimation unit 204 may perform uni-directional prediction on a current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0081] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction on a current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find one reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, where the plurality of reference indices indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors for the current video block as the motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0082] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.

[0083] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0084] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0085] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0086] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0087] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0088] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0089] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0090] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0091] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transformed coefficient video block, respectively, to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0092] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce block effect artifacts in the video block.

[0093] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate the entropy encoded data and output a bitstream including the entropy encoded data.

[0094] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 an example of the video decoder 124 in the system 100 shown.

[0095] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0096] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.

[0097] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of the prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.

[0098] The motion compensation unit 302 can generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.

[0099] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of video blocks to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0100] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures describing the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice can be the entire picture or can also be a region of the picture.

[0101] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0102] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.

[0103] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bitrates. 1. Brief Overview Embodiments of the present disclosure relate to video coding and decoding techniques. Specifically, it is about coding and decoding techniques for transformation, screen content coding and decoding, and local illumination compensation in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, ECM, etc. It is also applicable to future video coding and decoding standards or video codecs. 2. Introduction Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly simultaneously, and in the April 2018 JVET meeting, the new video coding standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at this time. Then the VVC working draft and the test model VTM are updated after each meeting. The VVC project achieved Feature Complete (FDIS) at the July 2020 meeting. In January 2021, JVET established the Exploration Experiments (EE) with the goal of enhanced compression efficiency beyond VVC capabilities using novel conventional algorithms. Soon after that, ECM was built as a common software library for the long-term exploration work towards the next-generation video coding standard. 2.1. Inter-frame prediction coding tools For each inter-frame predicted CU, the motion parameters include the motion vector, reference picture index and reference picture list use index, and additional information required for the new decoding features of VVC that will be used for inter-frame predicted sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is decoded using the skip mode, the CU is associated with a PU and has no significant residual coefficients, no decoded motion vector difference or reference picture index. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs (including spatial candidates and temporal candidates, as well as additional lists introduced in VVC). The Merge mode can be applied to any inter-frame predicted CU, not just the skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list use flag for each reference picture list, and other required information are signaled explicitly for each CU. In addition to the inter-frame coding features in HEVC, VVC includes multiple new and refined inter-frame prediction coding tools listed as follows: – Extended Merge prediction; – Merge Mode with MVD (MMVD); – Symmetric MVD (SMVD) signaling; – Affine motion compensation prediction; – Sub-block based temporal motion vector prediction (SbTMVP); – Adaptive motion vector resolution (AMVR); – Motion field storage: 1 / 16th luma sample MV storage and 8×8 motion field compression; – Bi-directional prediction with CU-level weights (BCW); – Bi-directional optical flow (BDOF); – Decoder-side motion vector refinement (DMVR); – Geometric partition mode (GPM); – Intra-inter combined prediction (CIIP). The following text provides details on those inter-prediction methods specified in VVC. 2.1.1. Extended Merge prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVP from spatially neighboring CUs; 2) Temporal MVP from co-located CUs; 3) History-based MVP from the FIFO table; 4) Pairwise-averaged MVP; 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in the Merge mode, rounding-unary binary coding (TU) is used to code the index of the best Merge candidate. The first binary bit of the Merge index is coded using context, and bypass coding is used for the other binary bits. The derivation process for each type of Merge candidate is provided in this session. As done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size of region. 2.1.1.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except for swapping the positions of the first two Merge candidates. At the location Figure 4Select up to four Merge candidates among the candidates at the positions depicted in []. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, A1 are unavailable (e.g., because they belong to another strip or slice) or are intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, so as to improve the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrow in Figure 5 are considered, and a candidate is added to the list only if the corresponding candidate for the redundancy check does not have the same motion information. 2.1.1.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, when deriving this temporal Merge candidate, the scaled motion vector is derived based on the co-located CUs belonging to the co-located reference picture. The reference picture list to be used for deriving the co-located CUs is explicitly signaled in the slice header. As shown by the dashed line in Figure 6 , the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. As depicted in Figure 7 , positions C0 and C1 for the temporal candidate are selected among the candidates. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used to derive the temporal Merge candidate. 2.1.1.3. History-based Merge candidate derivation After the spatial MVP and TMVP, the history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previously coded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a CU that is non-sub-block inter-coded, the associated motion information is added to the last entry of the table as a new HMVP candidate. The size S of the HMVP table is set to 6, which indicates that at most 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if the same HMVP exists in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are shifted forward. HMVP candidates can be used in the Merge candidate list construction process. Check several latest HMVP candidates in the table in order and insert them into the candidate list after the TMVP candidates. Apply a redundancy check to the HMVP candidates for both spatial and temporal Merge candidates. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list construction process from HMVP is terminated. 2.1.1.4. Pairwise-average Merge candidate derivation Pairwise-average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices into the Merge candidate list. Calculate the average motion vector separately for each reference list. If both motion vectors are available in a list, the two motion vectors are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list is kept invalid. When the Merge list is not full after adding pairwise-average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached. 2.1.1.5. Merge estimation region The Merge Estimation Region (MER) allows for independent derivation of the Merge candidate list for a CU within the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for the generation of the Merge candidate list for the current CU. Additionally, the update process for the history-based motion vector prediction value candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.1.2. Merge Mode with MVD (MMVD) In addition to the Merge mode, in cases where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge Mode with Motion Vector Difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after selecting a Merge candidate, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, an index for specifying the motion size, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The Merge candidate flag is signaled to specify which one to use. The distance index specifies the motion size information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1. Table 1 - Relationship between Distance Index and Predefined Offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate the four directions shown in Table 2. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a uni - directional prediction MV or a bi - directional prediction MV where two of the lists point to the same side of the current picture (i.e., both of the two reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 2 specify the sign of the MV offset added to the starting MV. When the starting MV is a bi - directional prediction MV with two MVs pointing to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), the symbols in Table 2 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 2 - Signs of MV Offsets Specified by the Direction Index Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.1. Bi - directional Prediction with CU - level Weights (BCW) In HEVC, a bi - directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi - directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w)*P0+w*P1 + 4) >> 3 (2 - 1) Five weights are allowed in weighted - average bi - directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi - directional prediction CU, the weight w is determined in one of two ways: 1) For non - Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low - latency picture, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion vector precisions. – When combined with affine, if and only if the affine mode is selected as the current best mode, affine ME will be performed for unequal weights. – When the two reference pictures in bi - directional prediction are the same, unequal weights are only conditionally checked. – When specific conditions are met, unequal weights are not searched, which depends on the POC distance, codec QP, and temporal level between the current picture and its reference picture. The BCW weight index is coded using a context coding binary bit, followed by a bypass coding binary bit. The first context coding binary bit indicates whether equal weights are used; and if unequal weights are used, the bypass coding signals additional binary bits to indicate the use of unequal weights. Weighted Prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficient coding of video content in fading situations. Support for WP is also added in the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets corresponding to the reference picture are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled, and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.2.2. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. The BDOF, previously called BIO, is included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is coded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order. – The distances from the two reference pictures to the current picture (i.e., POC differences) are the same. – Both reference pictures are short-term reference pictures. – The CU is coded without using the affine mode or the ATMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then, the motion refinement is used to adjust the bi-predicted sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the difference between two neighboring samples, and i.e., where I (k) (i, j) is the sample value at the coordinate (i, j) of the prediction signal in the list k (k = 0, 1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the auto-correlations and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as: where where Ω is the 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, the motion refinement (v x , v y ) is derived using the cross-correlation terms and auto-correlation terms with the following equation: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on motion refinement and gradients, the following adjustments are calculated for each sample point in a 4×4 sub-block: Finally, the BD - OF sample points of the CU are calculated by adjusting the bi - directionally predicted sample points as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )>>shift(2 - 7) These values are selected such that the multipliers in the BD - OF process do not exceed 15 bits, and the maximum bit - width of the intermediate parameters in the BD - OF process remains within 32 bits. To derive the gradient values, some predicted sample points I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 9 depicted, the BD - OF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating the predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8 - tap motion - compensated interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BD - OF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma sample points, it is divided into sub - blocks with a width and / or height equal to 16 luma sample points, and the sub - block boundaries are considered as CU boundaries during the BD - OF process. The maximum unit size for the BD - OF process is limited to 16×16. For each sub - block, the BD - OF process can be skipped. When the SAD between the initial L0 predicted sample points and the L1 predicted sample points is less than the threshold, the BD - OF process is not applied to the sub - block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub - block width and H indicates the sub - block height. To avoid the additional complexity of SAD calculation, the SAD calculated during the DVMR process between the initial L0 predicted sample points and the L1 predicted sample points is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. When a CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, BDOF is also disabled. 2.1.2.3. Symmetric MVD Coding (SMVD) In VVC, in addition to the normal uni-directional prediction and bi-directional prediction mode MVD signaling, a symmetric MVD mode for bi-directional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indexes for both list 0 and list 1 and the MVD for list 1 is not signaled, but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. – Otherwise, if the nearest reference picture in list -0 and the nearest reference picture in list -1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, then BiDirPredFlag is set to 1, and both the list -0 reference picture and the list -1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bi-directionally predicted and encoded and BiDirPredFlag is equal to 1, then the symmetric mode flag indicating whether to use the symmetric mode is signaled explicitly. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indexes for list 0 and list 1 are set to be equal to the reference picture pair respectively. MVD1 is set to be equal to (-MVD0). The final motion vectors are as follows. Figure 10 The symmetric MVD mode is shown. In the encoder, the symmetric MVD motion estimation starts with an initial MV evaluation. A set of initial MV candidates includes the MVs obtained from the uni-directional prediction search, the MVs obtained from the bi-directional prediction search, and the MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.3. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 11 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: where (mv 0x , mv 0y ) is the motion vector of the upper-left control point, (mv 1x , mv 1y ) is the motion vector of the upper-right control point, and (mv 2x , mv 2y ) is the motion vector of the lower-left control point. To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4×4 luminance sub-block, the motion vector of the central sample of each sub-block is calculated according to the above equations (as Figure 12 shown), and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction for each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of a 4×4 chrominance sub-block is calculated as the average of the MVs of the upper-left luminance sub-block and the lower-right luminance sub-block in the co-located 8×8 luminance region. Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.1.3.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs with both width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates are used to form the affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; – Constructed affine Merge candidate CPMVP derived using the translational MVs of neighboring CUs; – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 13 shown. For the prediction values on the left, the scanning order is A0->A1, and for the prediction values on the upper side, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When the neighboring affine CU is identified, its control point motion vectors are used to derive the CPMV candidates in the affine Merge list of the current CU. As Figure 30 shown, if the neighboring lower left block A is coded using the affine mode, the motion vectors v2, v3, and v4 of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. When block A is coded using the 4-parameter affine model, two CPMVs of the current CU are calculated based on v2 and v3. In the case where block A is coded using the 6-parameter affine model, three CPMVs of the current CU are calculated based on v2, v3, and v4. Figure 14 Shows the inheritance of control point motion vectors. The constructed affine candidates refer to constructing candidates by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from Figure 15 the specified spatial neighbors and temporal neighbors shown in k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, the B2->B3->A2 blocks are checked, and the MV of the first available block is used. For CPMV2, the B1->B0 blocks are checked, and for CPMV3, the A1->A0 blocks are checked. TMVP is used as CPMV4 (if available). After obtaining the MVs of the four control points, affine Merge candidates are constructed based on those motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}. Combinations of 3 CPMVs construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of control point MVs are discarded. After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list. 2.1.3.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with both width and height greater than or equal to 16. In the bitstream, a CU-level affine flag is signaled to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by sequentially using the following four types of CPMV candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; – Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs; – Translational MVs from neighboring CUs. – Zero MV. The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting the inherited affine motion prediction value into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 15 The same checking order as in the construction of affine Merge candidates is used. In addition, the reference picture indices of neighboring blocks are also checked. The first block in the checking order that is inter-coded and has the same reference picture as in the current CU is used. There is only one. When the current CU is coded using the 4-parameter affine mode and both mv0 and mv1 are available, they are added as a candidate in the affine AMVP list. When the current CU is coded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable. If the affine AMVP list candidates are still less than 2 after inserting valid inherited affine AMVP candidates and constructed AMVP candidates, mv0, mv1, and mv2 will be added in order as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.1.3.3. Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is only used to generate the inherited CPMV in the affine Merge mode and the inherited CPMV in the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMV are used for motion compensation, MV derivation in the Merge / AMVP list of translational MVs, and deblocking. To avoid picture line caches for additional CPMVs, the inheritance of affine motion data from the CU above the CTU is processed differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the left-bottom sub-block MV and the right-bottom sub-block MV in the line cache are used for affine MVP derivation instead of the CPMV. In this way, the CPMV is only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is degraded to a 4-parameter model. As Figure 16 shown, along the top boundary of the CTU, the left-bottom sub-block motion vector and the right-bottom sub-block motion vector of the CU are used for affine inheritance of the CU in the bottom of the CTU. 2.1.3.4. Prediction Refinement using Optical Flow (PROF) for Affine Modes Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, Prediction Refinement using Optical Flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equations. PROF is described in the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate the sub-block prediction I(i,j). Step 2) Using a 3-tap filter [-1,0,1], the spatial gradients g x (i,j) and g y (i,j) are calculated at each sample position. The gradient calculation is exactly the same as the gradient calculation in BDOF. g x (i,j) = (I(i+1,j) >> shift1) - (I(i-1,j))shift1) (2-11) g y (i,j) = (I(i,j+1))shift1) - (I(i,j-1))shift1) (2-12) shift1 is used to control the precision of the gradient. The sub-block (i.e., 4×4) prediction extends one sample point on each side of the gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended sample points on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Step 3) The brightness prediction refinement is calculated through the following optical flow equation. ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) (2 - 13) where as Figure 17 shown, Δv(i,j) is the sample point MV calculated for the sample point position (i,j), denoted as v(i,j), and is the difference from the sub-block MV of the sub-block to which the sample point (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 brightness sample point precision. Since the affine model parameters and the sample point position relative to the sub-block center do not change from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i,j) and dy(i,j) be the horizontal and vertical offsets of the sample point position (i,j) to the sub-block center (x SB ,y SB ), and Δv(x,y) can be derived through the following equation. To maintain accuracy, the input of the sub-block (x SB ,y SB ) is calculated as ((W SB –1) / 2,(H SB –1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For the 4-parameter affine model, For the 6-parameter affine model, where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the motion vectors of the upper left, upper right, and lower left control points, and w and h are the width and height of the CU. Step 4) Finally, the luminance prediction refinement ΔI(i,j) is added to the sub-block prediction I(i,j). The final prediction I’ is generated by the following equation. I′(i,j) = I(i,j) + ΔI(i,j) PROF is not applicable to affine-coded CUs in two cases: 1) all control point MVs are the same, which indicates that the CU has only translational motion; 2) the affine motion parameters are greater than the specified limit, because sub-block-based affine MC is degraded to CU-based MC to avoid large memory access bandwidth requirements. Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. In the following two cases, PROF is not applied to the affine motion estimation stage: a) If the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) If the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low-latency picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated. 2.1.4. Sub-block-based Temporal Motion Vector Prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and the Merge mode of the CUs in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; – While TMVP prefetches the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the center block relative to the current CU), SbTMVP applies a motion displacement before prefetching the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in FIG. 18. SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in FIG. 18(a) is checked. If A1 has a motion vector using the co-located picture as its reference picture, then that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the collocated picture as shown in Fig. 18(b). The example in Fig. 18(b) assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample point) in the collocated picture is used to derive the motion information of the sub-CU. After identifying the motion information of the collocated sub-CU, it is converted to the motion vector and reference index of the current sub-CU in a way similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used to signal the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8×8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8. The encoding / decoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P or B slice, additional RD checks are performed to decide whether to use the SbTMVP candidates. 2.1.5. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded / decoded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows: – Normal AMVP mode: quarter luminance sample, half luminance sample, integer luminance sample, or four luminance samples. – Affine AMVP mode: quarter luminance sample, integer luminance sample, or 1 / 16 luminance sample. If the current CU has at least one non-zero MVD component, the MVD resolution indication at the CU level is signaled conditionally. If all MVD components (i.e., both the horizontal MVD and the vertical MVD of reference list L0 and reference list L1) are zero, the quarter-luma sample MVD resolution is assumed. For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half-luma sample or other MVD precision (integer or quarter-luma sample) is used for normal AMVP CUs. In the case of half-luma samples, the half-luma sample positions use a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma sample or quarter-luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer-luma sample MVD precision or 1 / 16-luma sample MVD precision is used. To ensure that the reconstructed MVs have the expected precision (quarter-luma sample, half-luma sample, integer-luma sample, or quarter-luma sample), the motion vector prediction value of the CU is rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., a negative motion vector prediction value is rounded to positive infinity and a positive motion vector prediction value is rounded to negative infinity). The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM13, the RD check for MVD accuracy outside the quarter-luminance samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for the MVD accuracy of the quarter-luminance samples and the RD cost for the MV accuracy of the full-luminance samples are calculated. Then, the RD cost for the MVD accuracy of the full-luminance samples is compared with the RD cost for the MVD accuracy of the quarter-luminance samples to decide whether it is necessary to further check the RD cost for the MVD accuracy of the four-luminance samples. When the RD cost for the MVD accuracy of the quarter-luminance samples is much smaller than the RD cost for the MVD accuracy of the full-luminance samples, the RD check for the MVD accuracy of the four-luminance samples is skipped. Then, if the RD cost for the MVD accuracy of the full-luminance samples is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for the MVD accuracy of the half-luminance samples is skipped. For the affine AMVP mode, if the inter-frame affine mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, the Merge / skip mode, the normal AMVP mode with the MVD accuracy of the quarter-luminance samples, and the affine AMVP mode with the MVD accuracy of the quarter-luminance samples, the MV accuracy of the 1 / 16-luminance samples and the inter-frame affine mode with 1-pixel MV accuracy are not checked. Additionally, in the inter-frame affine modes with 1 / 16-luminance samples and quarter-luminance samples MV accuracy, the affine parameters obtained in the inter-frame affine mode with the quarter-luminance samples MV accuracy are used as the starting search points. 2.1.6. Bi-directional prediction with CU-level weights (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P0 + w * P1 + 4) >> 3 (2 - 18) Five weights are allowed in weighted-average bi-directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luminance samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights are used (w ∈ {3, 4, 5}). – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-latency picture, the unequal weights for 1-pixel and 4-pixel motion vector precisions are only conditionally checked. – When combined with affine, the affine ME for unequal weights is performed if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, the unequal weights are only conditionally checked. – The unequal weights are not searched when certain conditions are met, depending on the POC distance between the current picture and its reference picture, the coding / decoding QP, and the temporal level. The BCW weight index is decoded using a context-coded binary bit followed by a bypass-coded binary bit. The first context-coded binary bit indicates whether equal weights are used; and if unequal weights are used, the bypass-coded binary bit signals additional binary bits to indicate which unequal weight is used. Weighted prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding / decoding of video content in fading situations. Support for WP is also added in the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference picture are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded / decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.7. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. The BDOF, previously known as BIO, was included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is encoded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order. – The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same. – Both of the two reference pictures are short-term reference pictures. – The CU is not encoded and decoded using the affine mode or the SbTMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. Then the motion refinement is used to adjust the bidirectional prediction sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the differences between two neighboring samples, and i.e., where I (k) (i,j) is the sample value at the coordinate (i,j) of the prediction signal in list k (k = 0,1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as: where θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b ) where Ω is a 6×6 window around a 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, the motion refinement (v x , v y ) is derived using the cross - correlation term and the auto - correlation term with the following equation: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on the motion refinement and the gradient, the following adjustment is calculated for each sample point in the 4×4 sub - block: Finally, the BDOF sample points of the CU are calculated by adjusting the bi - directional predicted sample points as follows: pred BDOF (x,y) = (I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset ) >> shift (2 - 24) These values are selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit - width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient value, some predicted sample points I (k) (i,j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 9As depicted, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating out-of-boundary prediction samples, the prediction samples in the extended region (white positions) are generated by directly obtaining reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma samples, it is divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are treated as CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated during the DVMR process between the initial L0 prediction samples and the L1 prediction samples is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, then the BDOF is also disabled. When the CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.8. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the Merge mode MV, decoder-side motion vector refinement based on bilateral matching is applied in VVC. During the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. As Figure 20 shown, the SAD between the red blocks for each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs encoded or decoded using the following modes and features: – CU-level Merge mode with bidirectional prediction MVs. – One reference picture is past and the other reference picture is future with respect to the current picture. – The distances (i.e., POC differences) from the two reference pictures to the current picture are the same. – Both reference pictures are short-term reference pictures. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current block. – The CIIP mode is not used for the current block. The refined MV derived through the DMVR process is used to generate inter-predicted samples and is also used for temporal motion vector prediction in future picture coding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction in future CU coding. Additional functions of DMVR are mentioned in the following sub-articles. 2.1.8.1. Search Scheme In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point checked by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2-25) MV1′ = MV1 - MV_offset (2-26) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. To reduce the influence of DMVR refinement uncertainty, it is proposed to support the original MV in the DMVR process. The SAD between the reference blocks referred to by the initial MV candidates reduces the SAD value by 1 / 4. After integer sample point search, fractional sample point refinement is performed. To save computational complexity, the fractional sample point refinement is derived using the parametric error surface equation instead of performing additional search using SAD comparison. The fractional sample point refinement is conditionally invoked based on the output of the integer sample point search stage. When the integer sample point search stage ends at the center with the minimum SAD in the first iteration or the second iteration search, the fractional sample point refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y)=A(x - x min ) 2 +B(y - y min ) 2 +C (2 - 27) where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min ,y min ) is calculated as: x min =(E(-1,0) - E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2 - 28) y min =(E(0,-1) - E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2 - 29) The values of x min and y min are automatically restricted between -8 and 8 because all cost values are positive and the minimum value is e(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined delta MV. 2.1.8.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, within a 2-sample search range, compared with the normal motion compensation process, DVMR does not access more reference samples. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples of the normal MC process, samples will be filled from those available samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV. 2.1.8.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16×16. 2.1.9. Combined Intra-Inter Prediction (CIIP) In VVC, when a CU is coded / decoded using the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal and the intra prediction signal. The inter prediction signal P in the CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived after the regular intra prediction process with the planar mode. Then, a weighted average is used to combine the intra prediction signal and the inter prediction signal, where the weight values depend on the coding / decoding modes of the top neighboring block and the left neighboring block (depicted in Figure 21 and are calculated as follows: – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) equals 2, then set wt to 3; – Otherwise, if (isIntraLeft + isIntraTop) equals 1, then set wt to 2; – Otherwise, set wt to 1. CIIP prediction is formed as follows: P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (2 - 30) 2.1.10. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. A CU-level flag is used as a type of Merge mode to signal the geometric partitioning mode, where other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. Among a total of 64 partitions, for each possible CU size w×h = 2 m ×2 n , m, n ∈ {3…6}, partitioning is supported by the geometric partitioning mode. When using this mode, the CU is divided into two parts by a geometrically positioned line ( Figure 22 ). The position of the dividing line is mathematically derived from the angular parameter and offset parameter of a specific partition. Each part of the geometric partition in the CU uses its own motion for inter prediction; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure the same as regular bidirectional prediction, and only two motion-compensated predictions are required for each CU. If the current CU uses the geometric partitioning mode, then further signal the geometric partitioning index that indicates the partitioning mode (angle and offset) of the geometric partition and two Merge indices (one Merge index for each partition). The number of maximum GPM candidate sizes is explicitly signaled in the SPS and the syntax binarization of the GPM Merge index is specified. After predicting each part in the parts of the geometric partition, hybrid processing with adaptive weights is used to adjust the sample values along the geometric partition edge. This is the prediction signal for the entire CU, and the transform and quantization processes will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.10.1. Unidirectional Prediction Candidate List Construction Derive the unidirectional prediction candidate list directly from the Merge candidate list constructed according to the extended Merge prediction process. Denote n as the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the n-th extended Merge candidate (where X is equal to the parity of n) is used as the n-th unidirectional prediction motion vector of the geometric partitioning pattern. These motion vectors are marked with "x" in Figure 23 . In the case where there is no corresponding LX motion vector of the n-th extended Merge candidate, the L(1-X) motion vector of the same candidate is used instead of the unidirectional prediction motion vector for the geometric partitioning pattern. 2.1.10.2. Blending along the geometric partitioning edge After using its own motion prediction for each part of the geometric partitioning, blending is applied to the two prediction signals to derive the samples around the geometric partitioning edge. The blending weight for each position of the CU is derived based on the distance between the respective position and the partitioning edge. The distance of the position (x,y) to the partitioning edge is derived as: where i,j are the indices of the angle and offset of the geometric partitioning, which depend on the geometric partitioning index transmitted through the signal. The symbols ρ x,j and ρ y,j depend on the angle index. The weight for each part of the geometric partitioning is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) w1(x,y) = 1 - w0(x,y) partIdx depends on the angle index i. An example of the weight w0 is shown in Figure 24 . 2.1.10.3. Motion field storage for the geometric partitioning pattern Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combined Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded / decoded for the geometric partitioning pattern. The type of motion vector stored for each individual position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx ≤ 0? (1 – partIdx) : partIdx) where motionIdx is equal to d(4x + 2, 4y + 2). partIdx depends on the angle index i. If sType is equal to 0 or 1, store Mv0 or Mv1 in the corresponding motion field, otherwise if sType is equal to 2, store the combined Mv from Mv0 and Mv2. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), simply combine Mv1 and Mv2 to form a bi - directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, store only the uni - directional prediction motion Mv2. 2.1.11. Local Illumination Compensation (LIC) LIC is an inter - prediction technique that models the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, which form a linear equation, i.e., α * p[x]+β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for the AMVP mode to indicate the use of LIC. Local Illumination Compensation is used for uni - directional prediction inter - frame CUs with the following modifications. ● Intra - frame neighboring samples can be used for LIC parameter derivation; ● LIC is disabled for blocks with fewer than 32 luma samples; ● For both non - sub - blocks and affine modes, LIC parameter derivation is performed based on the mode - block samples corresponding to the current CU rather than the partial mode - block samples corresponding to the top - left 16×16 unit; ● The samples of the reference block template are generated by using MC with the block MV without rounding it to integer - pixel precision. 2.1.12. Non - adjacent Spatial Candidates Non - adjacent spatial Merge candidates are inserted after TMVP in the regular Merge candidate list. The mode of the spatial Merge candidate is shown in Figure 25 . The distance between the non - adjacent spatial candidate and the current coded block is based on the width and height of the current coded block. Row - buffer restrictions are not applied. 2.1.13. Template Matching I Template Matching I (TM) is a decoder - side MV derivation method to refine the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top - neighboring block and / or the left - neighboring block of the current CU) and a block in the reference picture (i.e., the same size as the template). As Figure 26As shown, within the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. The template matching method is used with the following modifications: the search step size is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in the Merge mode. In the AMVP mode, the MVP candidate is determined based on the template matching error to select the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM is only performed for this specific MVP candidate for MV refinement. TM refines this MVP candidate by starting with full pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode) within the [-8, +8] pixel search range using iterative diamond search. The AMVP candidate can be further refined by using a cross-shaped search with full pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode), and then sequentially searching by half pixels and quarter pixels depending on the AMVR mode specified in Table 3. This search process ensures that the MVP candidate still maintains the same MV accuracy as the MV accuracy indicated by the AMVR mode after the TM process. Table 3. Search patterns for AMVR and Merge mode with AMVR. In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 3, depending on the alternative interpolation filter (the interpolation filter used when AMVR is in the half-pixel mode) according to the merged motion information, TM can be performed in all ways up to 1 / 8 pixel MVD accuracy or skipped beyond the half-pixel MVD accuracy. Additionally, when the TM mode is enabled, template matching can be used as an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. 2.1.14. Multi-pass decoder-side motion vector refinement (mpDMVR) Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16×16 sub-block within the coded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction. 2.1.14.1. First pass - Block-based bilateral matching MV refinement In the first pass, a refined MV is derived by applying BM to the coding / decoding block. Similar to decoder-side motion vector refinement (DMVR), in the bi-prediction operation, a refined MV is searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between two reference blocks in L0 and L1. BM performs a local search to derive an integer-sample accuracy intDeltaMV. The local search applies a 3×3 square search pattern to loop within a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the distorted DC effect between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current minimum-cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. Existing fractional-sample refinement is further applied to derive the final deltaMV. Then, the refined MVs after the first pass are derived as: ● MV0_pass1 = MV0 + deltaMV; ● MV1_pass1 = MV1 - deltaMV. 2.1.14.2. Second Pass - Sub-block-based Bilateral Matching MV Refinement In the second pass, a refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, a refined MV is searched around two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub-blocks in L0 and L1. For each sub-block, BM performs an exhaustive search to derive an integer-sample accuracy intDeltaMV. The exhaustive search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub - blocks, i.e., bilCost = satdCost * costFactor. The search region (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond - shaped search regions, as Figure 27 shown. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in the order starting from the center of the search region. In each region, the search points are processed in raster - scan order from the upper - left corner of the region to the lower - right corner. When the minimum bilCost within the current search region is less than a threshold equal to sbW * sbH, the full - search of integer pixels is terminated; otherwise, the full - search of integer pixels continues to the next search region until all search points are checked. The existing VVC DMVR fractional - sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MV in the second pass is derived as: ●MV0_pass2(sbIdx2)=MV0_pass1 + deltaMV(sbIdx2); ●MV1_pass2(sbIdx2)=MV1_pass1 - deltaMV(sbIdx2). 2.1.14.3. Third Pass - Sub - block - based Bidirectional Optical Flow MV Refinement In the third pass, the refined MV is derived by applying BDOF to the 8×8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MV of the parent - child sub - blocks in the second pass, BDOF refinement is applied to derive the scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 - sample precision and clipped between - 32 and 32. The refined MV in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: ●MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv; ●MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv. 2.1.15. OBMC When OBMC is applied, the motion information of neighboring blocks with weighted prediction is used to refine the top - boundary pixels and left - boundary pixels of the CU. The conditions for not applying OBMC are as follows: ● When OBMC is disabled at the SPS level. ● When the current block has an intra mode or an IBC mode. ● When LIC is applied to the current block. ● When the current luma block area is less than or equal to 32. Sub - block boundary OBMC is performed by applying the same blend to the top sub - block boundary pixels, left sub - block boundary pixels, bottom sub - block boundary pixels, and right sub - block boundary pixels that use neighboring sub - blocks. It enables the following sub - block - based coding and decoding tools: ● Affine AMVP mode; ● Affine Merge mode and sub - block - based temporal motion vector prediction (SbTMVP); ● Sub - block - based bilateral matching. 2.1.16. Sample - based BDOF In sample - based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The coded block is divided into 8×8 sub - blocks. For each sub - block, it is determined whether to apply BDOF by checking the SAD between two reference sub - blocks against a threshold. If it is decided to apply BDOF to the sub - block, for each sample in the sub - block, a sliding 5×5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample of the window. 2.1.17. Interpolation The 8 - tap interpolation filter used in VVC is replaced by a 12 - tap filter. The interpolation filter is derived from a sine function, where the frequency response is truncated at the Nyquist frequency and clipped by a cosine window function. Table 4 gives the filter coefficients for all 16 phases. Figure 28 The frequency response of the interpolation filter is compared with the VVC interpolation filter, all at the half - pixel phase. Table 4. Filter coefficients of the 12 - tap interpolation filter 2.1.18. Multiple hypothesis prediction (MHP) In the multiple hypothesis inter - prediction mode, in addition to the regular bi - prediction signal, one or more additional motion - compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted sample - by - sample superposition. Using the bi - directional prediction signal p bi and the first additional inter - prediction signal / hypothesis h3, the resulting prediction signal p3 is obtained as follows. p3 = (1 - α)p bi+αh3 According to the following mapping, the weighting factor α is specified by the new syntax element add_hyp_weight_idx. add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 Similarly, more than one additional prediction signal may be used. The resulting overall prediction signal is cumulatively obtained iteratively using each additional prediction signal. p n+1 =(1 - α n+1 )p n +α n+1 h n+1 The resulting overall prediction signal is obtained as the last p n (i.e., having the largest index). Within this EE, up to two additional prediction signals may be used (i.e., n is limited to 2). The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference or implicitly by specifying a Merge index. A separate multi-hypothesis Merge flag differentiates between these two signaling modes. For the inter-frame AMVP mode, if non-equal weights in the BCW are selected in the bi-prediction mode, only MHP is applied. A combination of MHP and BDOF is possible, however BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the normal first two hypotheses). 2.1.19. Adaptive Reordering of Merge Candidates Using Template Matching (ARMC-TM) Merge candidates are adaptively reordered using template matching (TM). The reordering method is applied to the regular Merge mode, the template matching (TM) Merge mode, and the affine Merge mode (excluding SbTMVP candidates). For the TM Merge mode, the Merge candidates are reordered before the refinement process. After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5 for the regular Merge mode and the TM Merge mode. The subgroup size is set to 3 for the affine Merge mode. The Merge candidates within each subgroup are reordered in ascending order according to the template matching cost value. For simplicity, the Merge candidates in the last rather than the first subgroup are not reordered. The template matching cost of the Merge candidates is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate. When a Merge candidate utilizes bi - directional prediction, the reference samples of the template of the Merge candidate are also generated by bi - directional prediction as shown in Figure 29 . For a block - based Merge candidate with a sub - block size equal to Wsub×Hsub, the above - mentioned template includes a number of sub - templates of size Wsub×1, and the left - hand template includes a number of sub - templates of size 1×Hsub. As shown in Figure 30 , the motion information of the sub - blocks in the first row and the first column of the current block is used to derive the reference samples of each sub - template. 2.1.20. Geometric Partitioning Mode (GPM) with Merge Motion Vector Difference (MMVD) The GPM in VVC is extended by applying motion vector refinement to the top of the existing GPM unidirectional MVs. First, the flag of the GPM CU is signaled to specify whether to use this mode. If the mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD. If the MVD is signaled for a geometric partition, after selecting the GPM Merge candidate, the motion of the partition is further refined by the signaled MVD information. All other processes remain the same as GPM. The MVD is signaled as a pair of distance and direction, similar to in MMVD. There are nine candidate distances ( 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixel, 3 pixel, 4 pixel, 6 pixel, 8 pixel, 16 pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) involved in MMVD in GPM (GPM - MMVD). Additionally, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD is left - shifted by 2 in MMVD. 2.1.21. Geometric Partitioning Mode (GPM) Utilizing Template Matching (TM) Template matching is applied to GPM. When the GPM mode is enabled for a CU, a CU - level flag is signaled to indicate whether TM is applied to the two geometric partitions. TM is used to refine the motion information of each geometric partition. When TM is selected, templates are constructed using left - hand, above, or left - hand and above neighboring samples according to the partition angle, as shown in Table 5. Then the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern as the Merge mode with a disabled half - pixel interpolation filter. Table 5 Templates for the first and second geometric partitions, where A represents using above samples, L represents using left - hand samples, and L + A represents using both left - hand and above samples. The GPM candidate list is constructed as follows: 1. The interleaved list of 0MV candidates and the list of 1MV candidates are directly derived from the regular Merge candidate list, where the list of 0MV candidates has a higher priority than the list of 1MV candidates. A deduplication method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. 2. The interleaved list of 1MV candidates and the list of 0MV candidates are further directly derived from the regular Merge candidate list, where the list of 1MV candidates has a higher priority than the list of 0MV candidates. The same deduplication method with an adaptive threshold is also applied to remove redundant MV candidates. 3. Zero MV candidates are filled until the GPM candidate list is full. GPM-MMVD and GPM-TM are enabled specifically for one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to the two GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is presumed to be false. 2.1.22. GPM with Inter-Frame and Intra-Frame Prediction (GPM Inter-Frame - Intra-Frame) With GPM Inter-Frame - Intra-Frame, in addition to the Merge candidates for each non-rectangular partition region in the CU to which GPM is applied, a predetermined intra-frame prediction mode for the geometric split line can also be selected. In the proposed method, the intra-frame prediction mode or the inter-frame prediction mode is determined for each GPM separable region with a flag from the encoder. When in the inter-frame prediction mode, a unidirectional prediction signal is generated from the MVs in the Merge candidate list. On the other hand, when in the intra-frame prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra-frame prediction mode specified by an index from the encoder. The variation of possible intra-frame prediction modes is restricted by the geometry. Finally, the two unidirectional prediction signals are blended in the same way as ordinary GPM. 2.1.23. Adaptive Decoder-Side Motion Vector Refinement (Adaptive DMVR) The adaptive decoder-side motion vector refinement method consists of two new Merge modes, which are introduced to refine the MV only in one direction (L0 or L1) of the bi-directional prediction of the Merge candidates that satisfy the DMVR condition. A multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vector. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is zero. Similar to the conventional Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially neighboring coded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list is used by the two proposed Merge modes, and the Merge index is coded / decoded in the conventional Merge mode. 2.1.24. Bidirectional Matching AMVP-MERGE Mode (AMVP-MERGE) In the AMVP-Merge mode, the bidirectional prediction value consists of the AMVP prediction value in one direction and the Merge prediction value in the other direction. The AMVP part of the proposed mode is signaled as a conventional unidirectional AMVP, i.e., the reference index and MVD are signaled, and it has a derived MVP index (TM_AMVP) if template matching is used, or the MVP index is signaled when template matching is disabled. The Merge index is not signaled, and the Merge prediction value is selected from the candidate list with the minimum template or bidirectional matching cost. When the selected Merge prediction value and the AMVP prediction value satisfy the DMVR condition (which is at least one reference picture from the past and one reference picture from the future relative to the current picture) and the distances from the two reference pictures to the current picture are the same, bidirectional matching MV refinement is applied to the Merge MV candidate and the AMVP MVP as the starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the Merge prediction value or the AMVP prediction value with the higher template matching cost. The third pass of the 8×8 sub-PU BDOF refinement for multi-pass DMVR is enabled for the AMVP Merge mode coded / decoded blocks. 2.1.25. IBC Merge / AMVP List Construction The IBC Merge / AMVP list construction is modified as follows: ● An IBC Merge / AMVP candidate can be inserted into the IBC Merge / AMVP candidate list only if it is valid. ● The right upper spatial candidate, the lower left spatial candidate, and the upper left spatial candidate, as well as a paired average candidate, can be added to the IBC Merge / AMVP candidate list. ● Template-based Adaptive Reordering (ARMC-TM) is applied to the IBC Merge list. The size of the HMVP table for IBC is increased to 25. After deriving up to 20 IBC Merge candidates using full deduplication, they are reordered together. After reordering, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list. A set of BVP candidates located in the IBC reference region are used to replace the candidates for filling the zero vectors in the IBC Merge / AMVP list. The zero vectors are invalid as block vectors in the IBC Merge mode and, therefore, are discarded as BVPs in the IBC candidate list. Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of three sub-regions (A, B, and C), the coordinates of which are determined by the width and height of the current block and the ΔX parameter and ΔY parameter, as Figure 31 depicted. 2.1.26. IBC Using Template Matching Template matching is used in both the IBC Merge mode and the IBC AMVP mode in IBC. The IBC-TM Merge list is modified compared to the list used by the regular IBC Merge mode such that candidates are selected according to the deduplication method using the motion distance between candidates as in the regular TM Merge mode. The end-zero motion is fulfilled by motion vectors at the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU. In the IBC-TM Merge mode, the selected candidates are refined using the template matching method before the RDO or decoding process. The IBC-TM Merge mode has competed with the regular IBC Merge mode and is signaled by the TM-Merge flag. In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM Merge list. The template matching method is used to refine each of these 3 selected candidates, and they are sorted according to their resulting template matching cost. Then usually only the top two candidates are considered during the motion estimation process. The template matching refinement for both the IBC-TM Merge mode and the AMVP mode is very simple because the IBC motion vectors are constrained (i) to be integers and (ii) as Figure 32Within the reference region shown in []. Thus, in the IBC-TM Merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, it is performed with integer or 4-pixel precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the templates used in each refinement step must comply with the constraints of the reference region. 2.1.27. IBC Reference Region The reference region of IBC extends to two CTU rows above. Figure 33 The reference region for encoding / decoding CTU(m, n) is shown. Specifically, for CTU(m, n) to be encoded / decoded, the reference region includes CTUs with indices (m-2, n-2)…(W, n-2), (0, n-1), (W, n-1), (0, n)…(m, n), where W represents the maximum horizontal index within the current slice, strip, or picture. This setting ensures that for a CTU size of 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is horizontally limited to [-(C<<1), C>>2] and vertically limited to [-C, C>>2] to accommodate the reference region expansion, where C represents the CTU size. 2.1.28. MVD Symbol Prediction In this method, possible MVD symbol combinations are sorted according to the template matching cost, and the index corresponding to the true MVD symbol is derived and context decoded. On the decoder side, the MVD symbol is derived as follows: 1. Parse the size of the MVD component; 2. Parse the context decoded MVD symbol prediction index; 3. Construct MV candidates by creating combinations between possible symbols and absolute MVD values and adding them to the MV prediction value; 4. Derive the MVD symbol prediction cost for each derived MV based on the template matching cost and sorting; 5. Use the MVD symbol prediction index to select the true MVD symbol. MVD symbol prediction is applied to the inter-frame AMVP mode, affine AMVP mode, MMVD mode, and affine MMVD mode. 2.1.29. Enhanced Bidirectional Motion Compensation In bidirectional motion compensation, out-of-bounds (OOB) predicted samples are discarded, and only non-OOB predicted values are used to generate the final predicted value. Specifically, assume Pos_x i,j and Pos_y i,j represent the position of a predicted sample in a current block, and Indicates the MV of the current block; Pos LeftBdry 、Pos RightBdry 、Pos TopBdry and Pos BottomBdry are the positions of the four boundaries of the picture. When at least one of the following conditions is satisfied, a predicted sample is considered OOB: where half_pixel is equal to 8, which represents the half-pixel sample distance in 1 / 16 pixel sample precision. After checking the OOB condition for each sample, the final predicted sample of a bi-directional block is generated as follows: If is OOB and is non-OOB Otherwise, if is non-OOB and is OOB. Otherwise When BCW is enabled, the OOB check process also applies. 2.1.30. Block-level reference picture list reordering Use the block-level reference picture reordering method based on template matching. For the unidirectional prediction AMVP mode, the reference pictures in list 0 and list 1 are interleaved to generate a joint list. For each hypothesis of the reference pictures in the joint list, matching is performed to calculate the cost. The joint list is reordered based on the ascending order of the template matching cost. The index of the selected reference picture in the reordered joint list is signaled in the bitstream. For the bi-directional prediction AMVP mode, a list of pairs of reference pictures from list 0 and list 1 is generated and similarly reordered based on the template matching cost. The index of the selected pair is signaled. 2.1.31. Affine model inheritance based on historical parameters and non-adjacent affine modes Affine model inheritance based on historical parameters (HAMI) allows the affine model to inherit from previously affine-encoded blocks that may not be adjacent to the current block. Similar to the enhanced regular Merge mode, non-adjacent affine modes (NA-AFF) are introduced. Create the first Historical Parameter Table (HPT). Entries in the first HPT store sets of affine parameters: a, b, c, and d, each represented by a 16-bit signed integer. Entries in the HPT are classified by a reference list and a reference index. Five reference indices are supported for each reference list in the HPT. In a formulaic manner, the category of the HPT (denoted as HPTCat) is calculated as HPTCat(RefList, RefIdx) = 5 × RefList + min(RefIdx, 4), where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index, respectively. For each category, up to seven entries can be stored, resulting in a total of 70 entries in the HPT. At the start of each CTU row, the number of entries for each category is initialized to zero. After decoding a CU with affine encoding / decoding having a reference list RefList cur and RefIdx cur , the affine parameters are used to update the entry in the category HPTCat(RefList cur , RefIdx cur ) in a manner similar to the update of the HMVP table. A candidate based on historical affine parameters (HAPC) is derived from one of seven neighboring 4×4 blocks denoted as A0, A1, A2, B0, B1, B2, or B3 in Figure 4 and the set of affine parameters in the corresponding entry stored in the first HPT. The MV of the neighboring 4×4 block serves as the base MV. In a formulaic manner, the MV of the current block at position (x, y) is calculated as: where (mv h base , mv v base ) represents the MV of the neighboring 4×4 block, and (x base , y base ) represents the center position of the neighboring 4×4 block. (x, y) can be the upper-left, upper-right, and lower-left corners of the current block to obtain the corner position MV (CPMV) of the current block, or it can be the center of the current block to obtain the regular MV of the current block. A second Historical Parameter Table (HPT) with base MV information is also appended. There are nine entries in the second HPT, where the entries include the base MV, the reference index for each reference list, four affine parameters, and the base position. An additional Merge HAPC can be generated from the second HPT with base MV information, and the corresponding affine model is stored in the entry. The difference between the first HPT and the second HPT is shown in Figure 34 . In addition, paired affine Merge candidates are generated from two affine Merge candidates, which are either historically derived or non-historically derived. The paired affine Merge candidates are generated by averaging the CPMVs of the existing affine Merge candidates in the list. In response to the introduction of the new HAPC, the size of the Merge candidate list based on sub-blocks is increased from 5 to 15, all of which are involved in the ARMC process. In NA-AFF, the pattern of obtaining non-adjacent spatial neighbors is shown in Figure 6 Similar to the existing non-adjacent regular Merge candidates [8], the distance between the non-adjacent spatial neighbors and the current coding block in NA-AFF is also defined based on the width and height of the current CU. Utilize Figure 6 The motion information of the non-adjacent spatial neighbors in Figure 35 to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for the inherited candidates, except that the CPMV is inherited from the non-adjacent spatial neighbors, the same derivation process of the inherited affine Merge / AMVP candidates in VVC remains unchanged. The non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from near to far). At a specific distance, only the first available neighbors (coded with the affine mode) from each side of the current block (e.g., left and above) are included for the derivation of the inherited candidates. Figure 35 shows the spatial neighbors for deriving the affine Merge / AMVP candidates. In addition, Figure 35 sub-picture (a) of Figure 35 shows the spatial neighbors for deriving the inherited candidates, and sub-picture (b) of Figure 35 shows the spatial neighbors for deriving the first type of constructed candidates. As indicated by the dashed arrows in sub-picture (a) of Figure 35 Figure 35 , the checking order of the left and above neighbors is from bottom to top and from right to left, respectively. For the first type of constructed candidates, as shown in sub-picture (b) of Figure 35 Figure 35 , the positions of a left and an above non-adjacent spatial neighbor are first determined independently; afterwards, the position of the upper-left neighbor can be determined accordingly, which can enclose the rectangular virtual block with the left and above non-adjacent neighbors. Then, as shown in Figure 36 Figure 36 , the motion information of the three non-adjacent neighbors is used to form CPMVs at the upper-left (A), upper-right (B), and lower-left (C) of the virtual block, and finally project them onto the current CU to generate the corresponding constructed candidates. Insert the NA-AFF candidates into the existing affine Merge candidate list and affine AMVP candidate list according to the following order: Affine Merge Mode: 1. SbTMVP candidates, if available. 2. Inheritance from adjacent neighbors. 3. Inheritance from non - adjacent neighbors. 4. Construction from adjacent neighbors. 5. Construction affine candidates of the first type from non - adjacent neighbors. 6. Zero MV. Affine AMVP Mode: 1. Inheritance from adjacent neighbors. 2. Construction from adjacent neighbors. 3. Translational MV from adjacent neighbors. 4. Translational MV from temporal neighbors. 5. Inheritance from non - adjacent neighbors. 6. Construction affine candidates of the first type from non - adjacent neighbors. 7. Zero MV. Due to including additional candidates generated by NA - AFF, the size of the affine Merge candidate list increases from 5 to 15. The subgroup size of the ARMC for the affine Merge mode increases from 3 to 15. In NA - AFF: 1. Regions from non - adjacent neighbors are restricted within the current CTU (i.e., there is no additional storage requirement for line buffering). 2. The storage granularity of the affine motion information including CPMV and reference index is reduced from 8×8 to 16×16 (i.e., only the affine motion from the upper - left 8×8 block is saved). Additionally, the saved CPMV is projected onto each 16×16 block before storage, such that position and size information are not required. 3. Only the upper - left CPMV and the upper - right CPMV are stored (i.e., a 4 - parameter affine model for NA - AFF is always used). 2.1.32. Regression - based affine candidate derivation method A regression - based affine candidate derivation method is proposed. The sub - block motion fields from previously decoded affine CUs and the motion vectors from adjacent sub - blocks of the current CU are used as the inputs for the regression process. The predicted CPMV instead of the sub - block motion fields of the current block is derived as the output. The derived CPMV can be added to the sub - block Merge candidate list or the affine AMVP list. The scan pattern of the previously decoded affine CUs is the same as the non - adjacent scan pattern used in the construction of the regular Merge candidate list. 2.2. Transform and coefficient coding 2.2.1. Large - block - size transform with high - frequency zeroing In VVC, large block size transforms with sizes up to 64×64 are enabled, which are mainly used for higher resolution videos such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are zeroed out so that only the lower frequency coefficients are retained. For example, for an M×N transform block where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the first 32 rows of transform coefficients are retained. When the transform skip mode is used for large blocks, the entire block is used without zeroing out any values. Additionally, the transform shift is removed in the transform skip mode. VTM also supports a configurable maximum transform size in the SPS, giving the encoder the flexibility to select a transform size of up to 32 lengths or 64 lengths according to the needs of a particular implementation. 2.2.2. Multiple Transform Selection (MTS) for Kernel Transform In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of blocks coded inter and intra. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the basis functions of the selected DST / DCT. Table 6 - Transform Basis Functions of DCT-II / VIII and DSTVII for N-Point Input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS only applies to luminance. The MTS signaling is skipped when one of the following conditions is met: – The position of the last significant coefficient of the luminance TB is less than 1 (i.e., only DC). – The last significant coefficient of the luminance TB is within the MTS zeroing region. If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 7. A unified transform selection for ISP and implicit MTS is used by eliminating the intra mode and block shape dependencies. If the current block is in the ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for the horizontal and vertical transform kernels. In terms of the transform matrix precision, 8-bit primary transform kernels are used. Therefore, all the transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. Additionally, for other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8), 8-bit primary transform kernels are used. Table 7 - Transform and Signaling Mapping Table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16×16 low-frequency region are retained. Similar to HEVC, the residual of a block can be encoded and decoded using the transform skip mode. To avoid redundancy in syntax encoding and decoding, when the CU-level MTS_CU_flag is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Additionally, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled. 2.2.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, as Figure 37 shown, LFNST is applied between the forward primary transform and quantization (at the encoder) and between the de-quantization and inverse primary transform (at the decoder side). In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). The following uses the input as an example to describe the application of the non-separable transform used in LFNST. To apply 4×4 LFNST, the 4×4 input block X is first represented as a vector The non-separable transform is calculated as where indicates a transform coefficient vector, and T is a 16×16 transform matrix. Subsequently, the 16×1 coefficient vector is reorganized into 4×4 blocks using the scan order (horizontal, vertical, or diagonal) for the block. Coefficients with smaller indices are placed in the 4×4 coefficient block together with smaller scan indices. 2.2.3.1. Reduced non-separable transform The LFNST (Low Frequency Non-Separable Transform) applies the non-separable transform based on a direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix size to minimize the computational complexity and the memory spatial domain to store the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in the LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (where N is typically equal to 64 for 8×8 NSST) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows: ​Among them, the transformed R rows are the R bases of the N-dimensional space. The inverse transformation matrix for RT is the transpose of its forward transformation. For the 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix (which is the conventional 8×8 non-separable transformation matrix size) is reduced to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix is used on the decoder side to generate the kernel (primary) transformation coefficients in the upper left 8×8 region. When applying the 16×48 matrix instead of the 16×64 with the same transformation set configuration, each of them obtains 48 input data from three 4×4 blocks in the upper left 8×8 block except for the lower right 4×4 block. With the reduced size, the memory usage for storing all LFNST matrices is reduced from 10 KB to 8 KB, with a reasonable performance degradation. To reduce complexity, it is applicable to limit the LFNST only when all coefficients outside the first coefficient subgroup are not significant. Therefore, when applying the LFNST, all only the primary transformation coefficients must be zero. This allows adjusting the LFNST index signaling at the last valid position, and thus avoids the additional coefficient scanning in the current LFNST design, which requires checking the valid coefficients only at specific positions. The worst-case processing of the LFNST (in terms of multiplications per pixel) limits the non-separable transformations of the 4×4 block and the 8×8 block to 8×16 transformations and 8×48 transformations respectively. In these cases, when applying the LFNST, the last valid scan position must be less than 8, for other sizes less than 16. For blocks with shapes of 4×N and N×4 and N>8, the proposed limitation means that the LFNST is now applied only once and only to the upper left 4×4 region. Since all only the primary coefficients are zero when applying the LFNST, the number of operations required for the primary transformation is reduced in this case. From the perspective of the encoder, when testing the LFNST transformation, the quantization of the coefficients is significantly simplified. Rate-distortion optimized quantization must be maximally completed for the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.2.3.2. LFNST Transform Selection There are a total of 4 transformation sets and 2 non-separable transformation matrices (kernels) used in each transformation set in the LFNST. As shown in Table 8, the mapping from the intra prediction mode to the transformation set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transformation set 0 is selected for the current chrominance block. For each transformation set, the selected non-separable quadratic transformation candidate is further specified by the explicitly signaled LFNST index. The index is signaled once in the bitstream for each intra CU after the transformation coefficients. Table 8 - Transformation Selection Table IntraPredMode Transform Set Index IntraPredMode < 0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode <= 80 1 81 <= IntraPredMode <= 83 0 2.2.3.3. LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to apply only when all coefficients outside the first coefficient subgroup are not significant, the LFNST index encoding and decoding depends on the position of the last significant coefficient. Additionally, the LFNST index is context - decoded but does not depend on the intra - prediction mode, and only the first binary bit is context - decoded. Moreover, LFNST is applied to intra CUs in both intra - slices and inter - slices, and for both luminance and chrominance. If dual - tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter - slices (dual - tree is disabled), a single LFNST index is signaled and used for both luminance and chrominance. Considering that due to the existing maximum transform size limit (64×64), large CUs larger than 64×64 are implicitly partitioned (TU slicing), the LFNST index search can increase the data cache by up to four times the number of decoding pipeline stages. Therefore, the maximum size allowing LFNST is restricted to 64×64. Note that LFNST only enables DCT2. The LFNST index signaling is placed before the MTS index signaling. It is not obvious to use the scaling matrix for perceptual quantization, and the scaling matrix specified for the main matrix can be used for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For the single - tree splitting mode, chrominance LFNST is not applied. 2.2.4. Sub - block Transform (SBT) In VTM, a sub - block transform is introduced for CUs in inter - prediction. In this transform mode, for a CU, only a sub - part of the residual block is encoded and decoded. When the cu_cbf of an inter - predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub - part of the residual block is encoded and decoded. For the former case, the inter - frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by the presumed adaptive transform while the other part of the residual block is zeroed. When SBT is used for inter - decoded CUs, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 38As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in terms of luma samples), a 1:3 / 3:1 partition is not allowed along that dimension. A CU has at most 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal transform and the vertical transform for each SBT position are specified in Figure 38 For example, the horizontal transform and the vertical transform for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slicing, cbf, and the horizontal and vertical kernel transform types of the residual block. SBT is not applied to CUs coded using the combined inter-intra mode. 2.2.5. Maximum Transform Size and Zeroing of Transform Coefficients Both the CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the maximum intra-coded block can have a size of 128×128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, there is no standardized zeroing operation applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are standardizedly zeroed. 2.2.6. Enhanced MTS for Intra Coding In the current VVC design, for MTS, only the DST7 transform kernel and the DCT8 transform kernel are utilized, which are used for both intra and inter coding. Additional main transforms including DCT5, DST4, DST1, and the identity transform (IDT) are adopted. The MTS set also depends on the TU size and the intra mode information. 16 different TU sizes are considered, and for each TU size 5, different categories are considered according to the intra mode information. For each category, 1, 4, or 6 different transform pairs are considered. Multiple intra MTS candidates (among 1, 4, and 6 MTS candidates) are adaptively selected according to the sum of the absolute values of the transform coefficients. The sum is compared with two fixed thresholds to determine the total number of allowed MTS candidates: 1 candidate: sum <= th0. 4 candidates: th0 < sum <= th1. 6 candidates: sum > th1. Note that although 80 different categories are considered in total, some of these different categories usually share exactly the same set of transforms. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For the angular mode, joint symmetry on the TU shape and intra prediction is considered. Thus, a mode i (i > 34) with a TU shape A × B will be mapped to the same category corresponding to a mode j = (68 - i) with a TU shape B × A. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16 × 4 block with mode 18 (horizontal prediction) and a 4 × 16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For the wide-angle mode, the closest regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. 2.2.7. Quadratic Transform: LFNST Extension with Large Kernels The LFNST in VVC is extended as follows: ● The number of the LFNST set (S) and candidates (C) is extended to S = 35 and C = 3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: o For predModeIntra < 2, lfnstTrSetIdx is equal to 2; o lfnstTrSetIdx = predModeIntra, for predModeIntra in [0, 34]; o lfnstTrSetIdx = 68 - predModeIntra, for predModeIntra in [35, 66]. ● Three different kernel LFNSTs 4, LFNST 8, and LFNST 16 are defined to indicate the sets of LFNST kernels applied to 4 × N / N × 4 (N ≥ 4), 8 × N / N × 8 (N ≥ 8), and M × N (M, N ≥ 16), respectively. The kernel size is specified as follows: (LFSNT4, LLFNST8*, LFNST16*) = (16 × 16, 32 × 64, 32 × 96) The positive LFNST is applied to the upper-left low-frequency region called the region of interest (ROI). When applying the LFNST, the main transform coefficients existing in the region other than the ROI are cleared and are not changed from the VVC standard. The ROI of LFNST16 is in Figure 39 shown. It consists of six 4×4 sub-blocks, which are consecutive in scan order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be R×96. In this contribution, R is chosen as 32, and accordingly 32 coefficients (two 4×4 sub-blocks) are generated from the forward LFNST16, which are placed after the coefficient scan order. The ROI of LFNST8 is in Figure 40 shown. The forward LFNST8 matrix can be R×64, and R is chosen as 32. The generated coefficients are positioned in the same way as LFNST 16. The mapping from the intra prediction mode to these sets is shown in Table 9, Table 9. Mapping of Intra Prediction Mode to LFNST Set Index Intra Prediction Mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST Set Index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra Prediction Mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST Set Index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra Prediction Mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST Set Index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.2.8. Sign Prediction The basic idea of the coefficient sign prediction method is to calculate the reconstructed residuals for the negative and positive sign combinations for applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal sign, the cost function is defined as a measure of the discontinuity across Figure 41 the block boundaries shown above. All hypotheses are measured, and the hypothesis with the minimum cost is selected as the predicted value of the coefficient sign. The cost function is defined as the sum of the absolute second derivatives in the residual domain of the upper row and the left column as follows: where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. The term (-R-1 + 2R0 - P1) can be calculated only once per block and only the residual hypothesis is subtracted. The transform coefficients with the maximum K qIdx value in the upper-left 4×4 region are selected. After compensating for the effects of multiple quantizers in DQ, the qIdx value is at the transform system level. A larger qIdx value will result in a larger dequantized transform coefficient level. The qIdx is derived as follows: qIdx = (abs(level) << 1) - (state & 1); where level is the transform coefficient level parsed from the bitstream and state is a variable maintained by the encoder and decoder in DQ. The symbol prediction area is extended to a maximum of 32×32. The symbols of the upper-left M×N block are predicted. The values of M and N are calculated as follows: ○ M = min(w, maxW) ○ N = min(h, maxH) where w and h are the width and height of the transformed block. The maximum area for symbol prediction is not always set to 32×32. The encoder sets the maximum area (maxW, maxH) based on the configuration, sequence category, and QP, and signals the area in the SPS. The maximum number of predicted symbols remains unchanged. Symbol prediction is also applied to the LFNST block. And for the LFNST block, symbol prediction is allowed for the 4 largest coefficients in the upper-left 4×4 area. 2.2.9. Non-separable Primary Transform (NSPT) for Intra Coding and Decoding DCT-II + LFNST is replaced by NSPT for block sizes 4×4, 4×8, 8×4, and 8×8, as Figure 42 shown. Therefore, LFNST4 and LFNST8 will not be tested for these block sizes. However, they are still used for larger block sizes and are not removed. Thus, NSPT can be considered an extension of the DCT-II + LFNST design. The NSPT in this proposal follows the design of LFNST, i.e., 3 candidates and 35 sets, which are selected based on the intra mode. The kernel sizes are as follows: ● NSPT4×4: 16×16; ● NSPT4×8 / NSPT8×4: 32×20; ● NSPT8×8: 64×32. Therefore, 12 coefficients and 32 coefficients are zeroed for NSPT4×8 / NSPT8×4 and NSPT8×8, respectively. 2.3. Spatial GPM (SGPM) SGPM is an intra mode. In SGPM, a candidate list including split partitions and two intra prediction modes is established. No more than 11 MPMs of the intra prediction mode are used to form combinations, and the length of the candidate list is set to be equal to 16. The selected candidate index is signaled. Figure 43 Spatial GPM candidates are shown. Use Figure 44 the template shown to reorder the list. The GPM mixing process is not used in the template, and the SAD between the prediction and reconstruction of the template is used for sorting. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in the inter GPM. Figure 45 GPM mixing is shown. The following items are considered: ● Airspace GPM segmentation mode: 26 predefined modes; Adaptive derivation algorithm based on the ratio of horizontal gradient and vertical gradient. ● Intra prediction mode selection: IPM lists with and without TIMD: For each segmentation mode, the IPM list is derived for each part using the intra-inter GPM list. The IPM list size is 3. In the list, the TIMD derivation mode is replaced by 2 derivation modes (using the top template or the left template) with horizontal and vertical directions or the TIMD derivation mode is excluded. MPM list: A unified MPM list (not exceeding 11 elements) is used for all segmentation modes. ● Template size (left and above): 1 or 4. ● Extended block size: The airspace GPM is extended to be further applied to 4×8, 8×4, 4×16, and 16×4 blocks, which can be described as 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32. ● Adaptive mixing: Adaptive mixing is tested for the airspace GPM, where the mixing depth τ is derived as follows: ■ If min(width, height) == 4, then select 1 / 2τ; ■ Otherwise, if min(width, height) == 8, then select τ; ■ Otherwise, if min(width, height) == 16, then select 2τ; ■ Otherwise, if min(width, height) == 32, then select 4τ; ■ Otherwise, select 8τ. 3. Problems / Issues There are several problems with existing video coding and decoding technologies, which can be further improved for higher coding and decoding gains. 1. In ECM-6.0, affine candidates can be derived from adjacent affine-based candidates, history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. A similarity check is performed for affine candidate derivation. However, a differential similarity check rule is used for affine candidate derivation. This may not be optimal. 2. In ECM-6.0, hybrid modes such as CIIP and OBMC are applied to both videos captured by cameras and screen content videos, which may not be efficient. 3. In ECM-6.0, KLT is allowed to be used in the explicit inter-frame MTS mode. Specifically, if the TU for inter-frame coding / decoding is less than or equal to 16×16, two KLT options (i.e., KLT0 and KLT1) of the inter-frame MTS core are used to replace DST7 and DCT8. This design can be changed for higher coding / decoding efficiency. 4. In ECM-6.0, KLT is allowed to be used in the inter-frame MTS mode, but the use of KLT does not depend on which inter-frame prediction technique is used for the video unit, and this can be further improved. 5. In ECM-6.0, the following intra-frame mode derivation / mapping for intra-frame MTS and LFNST indices can be improved. 1) In the dual-tree case, LFNST is applied to the chrominance components. For the CCLM mode, the co-located luma mode is used for the LFNST transform set and the transpose flag index. 2) For the MIP mode, the planar mode is used for the LFNST transform set and the transpose flag index. 3) MIP is regarded as a special mode for the intra-frame MTS transform class and the intra-frame MTS transform pair index. 4) IntraTMP is allowed to use implicit MTS (e.g., DST7) and LFNST (regarded as the planar mode). 5) In hybrid modes such as the TIMD hybrid mode and the DIMD hybrid mode, only the first intra-frame mode is considered for the MTS / LFNST index. 6. For coding / decoding tools such as IntraTMP and IBC, there is no sample refinement process for motion-compensated prediction, and this can be designed in the future. 7. For sbTMVP coding / decoding, DMVR is not allowed in the current codec, and this can be changed for higher coding / decoding gain. 8. For coding / decoding tools such as LIC, OBMC, and LM, uniform parameters are derived for the samples to be processed in the block. However, a multi-model-based method can be used to perform non-uniform mixing. 9. In ECM7.0, the following transforms are designed, and the following transforms can be improved. a. The primary transform is based on a one-dimensional separable transform, such as DCT2 or an MTS core. i. Depending on the block size, the primary transform for the ISP block is DCT2 or implicit DST7. ii. Depending on the block size, the primary transform for the MIP block can be DCT2 or explicit intra-frame MTS. iii. Depending on the block size and the intra-frame mode, the primary transform for the SGPM block can be DCT2 or explicit intra-frame MTS. iv. The intra-mode within the first frame depending on the block size and the mixing mode. For the TIMD mixing mode and the DIMD mixing mode, the primary transform can be DCT2 or explicit intra MTS. v. The intra-mode within the first frame depending on the block size and the intra-fusion mode. For the intra-luma fusion mode, the primary transform can be DCT2 or explicit intra MTS. b. The secondary transform is based on a two-dimensional non-separable transform, such as LFNST. i. The secondary transform for the ISP block depends on the intra-mode of the entire ISP CU, which means that all ISP sub-blocks use the same LFNST kernel. ii. The secondary transform for the MIP block is based on the derived intra-mode. iii. The secondary transform for the SGPM block is based on the intra-mode co-located with the SGPM segmentation mode. iv. The secondary transform for the TIMD mixing mode and the DIMD mixing mode is based on the intra-mode within the first frame of the mixing mode. v. The secondary transform for the intra-luma fusion mode is based on the intra-mode within the first frame of the intra-fusion mode. c. The determination rule for the primary transform (e.g., DCT2, MTS) for the intra / inter / intra-TMP block is the same regardless of whether it is screen content or camera-captured content. d. The determination rule for the secondary transform (e.g., LFNST) for the intra / inter / intra-TMP block is the same regardless of whether it is screen content or camera-captured content. e. For the intra-inter mixing mode (such as CIIP, GPM intra-inter), it always follows the inter-block transform branch, i.e., uses the inter MTS other than the intra MTS. f. Transform skip is not allowed for the ISP block. 4. Embodiments of the present disclosure The following detailed embodiments should be considered as examples for explaining the general concept. These embodiments should not be interpreted in a narrow way. In addition, these embodiments can be combined in any way. The term "video unit" or "codec unit" or "block" may represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. The term "KLT" may refer to a type of transform. For example, it may refer to the Karhunen-Loeve transform. For example, it may refer to any transform type that is not DCT or DST or Hadmard. The coefficient matrix associated with a specific KLT can be trained online or predefined (e.g., offline training) based on some prior knowledge (e.g., the residuals / coefficients from already decoded neighboring blocks). In the present disclosure, regarding "blocks encoded / decoded using mode N", here "mode N" can be a specific prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a specific prediction technique (e.g., AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, GPM intra, MHP, OBMC, LIC, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, sub-block encoding / decoding, hypothesis encoding / decoding, etc.), or a specific transform process (IDTX, DCT-X, DST-Y, KLT-Z, where X / Y / Z are constants) or a specific filter process (deblocking, SAO, bilateral filter, adaptive loop filter, CCSAO, CC-ALF, etc.). Note that the terms mentioned below are not limited to the specific terms defined in the existing standards. Any changes to the encoding / decoding tools are also applicable. 4.1. Regarding the first problem of the derivation of affine candidates and other motion candidate deduplication processes, the following method is proposed: a. The same logic / rules / process for similarity / consistency / deduplication checking can be used for the derivation of all affine candidates. a. For example, it can refer to the derivation of affine Merge candidates. b. For example, it can refer to the derivation of affine AMVP candidates. c. For example, it can refer to both the derivation of affine Merge candidates and affine AMVP candidates. d. For example, it can refer to the derivation of history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. b. The logic / rules / process for similarity / consistency / deduplication checking can refer to comparing one or more of the following elements associated with the first affine candidate with these elements associated with the second affine candidate: a. Inter-frame direction (prediction direction). b. Affine type (e.g., 6-parameter affine or 4-parameter affine). c. Sub-block Merge type (e.g., sbTMVP or affine). d. Bcw index. e. LIC flag. f. Reference index. g. Motion vector (e.g., horizontal component and / or vertical component). h. Control point motion vector (CPMV). i. The first CPMV (e.g., the upper left CPMV) and / or the second CPMV (e.g., the upper right CPMV) and / or the third CPMV (e.g., the lower left CPMV). j. The horizontal displacement and / or vertical displacement between the first CPMV and the second CPMV (e.g., the absolute difference between the horizontal components and / or vertical components of the first CPMV and the second CPMV). k. The horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV (e.g., the absolute difference between the horizontal components and / or vertical components of the first CPMV and the third CPMV). c. A consistency check can be performed on the comparison of the elements listed in item b. a. For example, if the elements associated with the second candidate are the same as the elements associated with the first candidate, the second candidate is not added to the affine candidate list. d. A similarity check can be performed on the comparisons listed in item b. a. For example, if the elements associated with the second candidate are similar to the elements associated with the first candidate, the second candidate is not added to the affine candidate list. b. For example, "similar" can refer to a comparison based on a threshold. i. For example, the absolute difference is less than the threshold. ii. For example, the absolute difference is not greater than the threshold. c. For example, the threshold can depend on the block size, such as the width and / or height. i. For example, the threshold is determined adaptively according to the block width / height. ii. For example, a smaller threshold can be set for a smaller block size, while a larger threshold can be set for a larger block size. d. For example, the threshold can be a predefined fixed value (such as 0 or 1). e. For example, a similarity check can be performed separately for the horizontal and vertical components of the motion vector. f. For example, a similarity check can be performed separately for the horizontal and vertical components of the control point motion vector. e. Whether to apply a similarity / consistency / deduplication check to the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV can depend on the affine type (e.g., 6-parameter affine or 4-parameter affine). a. For example, only when the affine types of the first affine candidate and the second affine candidate are 6-parameter, a similarity / consistency / deduplication check can be performed on the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV. f. For example, for Merge list / candidate deduplication, coding / decoding information other than motion similarity can be checked (e.g., BCW index, LIC flag, etc.). a. For example, if the second candidate has the same / similar motion vector as the first candidate but a different BCW index value, the second candidate can be considered non-redundant / non-similar to the first candidate. b. For example, if the second candidate has the same / similar motion vector as the first candidate but a different LIC flag value, the second candidate can be considered non-redundant / non-similar to the first candidate. c. For example, Merge list / candidate deduplication can refer to a conventional Merge list. d. For example, Merge list / candidate deduplication can refer to an MMVD-based Merge list. e. For example, Merge list / candidate deduplication can refer to a TM-based Merge list. f. For example, Merge list / candidate deduplication can refer to a BM-based Merge list. g. For example, Merge list / candidate deduplication can refer to a DMVR-based Merge list. h. For example, Merge list / candidate deduplication can refer to an affine DMVR Merge list. i. For example, Merge list / candidate deduplication can refer to a CIIP Merge list (with or without using TM). j. For example, Merge list / candidate deduplication can refer to a GPM Merge list (with or without using TM). k. For example, Merge list / candidate deduplication can refer to a sbTMVP Merge list (with or without using TM). 4.2. Regarding the second problem of coding / decoding tools for screen content tools, the following methods are proposed: a. For at least one of the following coding / decoding tools, coding / decoding for video units can be disallowed or restricted or prohibited. a) CIIP and / or its variants (e.g., CIIP PDPC, CIIP TM, CIIP TIMD, etc.). b) OBMC and / or its variants (e.g., OBMC TM, etc.). c) TIMD and / or its variants. d) DIMD and / or its variants. e) MHP and / or its variants. f) DMVR and / or its variants. g) interTM and / or its variants. h) CCALF and / or its variants. i) CCSA0 and / or its variants. b. Whether the video unit does not allow or restricts or prohibits the codec tool may depend on the profile / level / layer. c. Whether the video unit does not allow or restricts or prohibits the codec tool may depend on whether the video unit belongs to a specific video type. a) The specific video type may refer to screen content video. d. In addition, the codec tool may be not allowed or restricted or prohibited for a video sequence or a group of pictures or a picture or a slice. a) Such a restriction or non - allowance may be reflected by a bitstream constraint. b) Such a restriction or non - allowance or allowance may be reflected by a syntax element (e.g., a flag) signaled in the bitstream. e. A syntax element (e.g., a flag) may be signaled in the bitstream to impose such a restriction or non - allowance or allowance on the specific codec tools listed in item a. a) The syntax element may be signaled at the sequence level / group of pictures level / picture level / slice level / tile group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header. f. A single syntax element (e.g., a flag) may be signaled to impose such a restriction or non - allowance or allowance on more than one codec tool listed in item a. a) The syntax element may be signaled at the sequence level / group of pictures level / picture level / slice level / tile group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header. g. Different weight values / factors / tables / sets may be allowed to mix multiple prediction hypotheses of video units coded with codec tool X. a) For example, which weight value / factor / table / set is allowed to mix multiple prediction hypotheses of a video unit may depend on the content type. b) For example, the first weight value / factor / table / set may be allowed to mix multiple prediction hypotheses of the first video unit, while the second weight value / factor / table / set may be allowed to mix multiple prediction hypotheses of the second video unit. c) For example, the first video unit may belong to a video sequence captured by a camera. d) For example, the second video unit may belong to a screen content video sequence. e) For example, assume that two prediction hypotheses are mixed into the final prediction of a video unit. i. For the first type of video unit, the final prediction block of the video unit encoded / decoded using codec tool X can be "exactly equal to the first prediction hypothesis" or "exactly equal to the second prediction hypothesis". 1. For example, the allowed weight value / factor / table / set for a video unit can be equal to 0 or 1. ii. For the second type of video unit, the final hybrid prediction block of the video unit encoded / decoded using codec tool X can be "a combination of both the first prediction hypothesis and the second prediction hypothesis". 1. For example, the allowed weight value / factor / table / set for a video unit can be a fraction between 0 and 1 (e.g., the fractional value of the weight of a specific prediction hypothesis can be quantized to an integer in the codec). f) For example, the final prediction (e.g., hybrid) of a video unit with codec tool X can be further mixed / weighted / combined with another video unit encoded / decoded using another codec tool. g) For example, a specific codec tool X can be CIIP and / or its variants. h) For example, a specific codec tool X can be GPM and / or its variants. i) For example, a specific codec tool X can be MHP and / or its variants. j) For example, a specific codec tool X can be OBMC and / or its variants. k) For example, a specific codec tool X can be TIMD hybrid mode and / or its variants. l) For example, a specific codec tool X can be DIMD hybrid mode and / or its variants. 4.3. Regarding the third issue of the general use of KLT for the transform process, the following method is proposed: a. More than two KLT kernels can be allowed in the codec. b. KLT kernels can be allowed for the primary transform and / or the secondary transform. c. KLT kernels can be allowed for the chrominance component. a. For example, different KLT kernels can be used for the luma component and the chrominance component of a video unit. i. Alternatively, all color components of a video unit can share the same KLT kernel. b. For example, different KLT kernels can be used for the chrominance Cb component and the chrominance Cr component of a video unit. i. Alternatively, the chrominance Cb component and the chrominance Cr component of a video unit can share the same KLT kernel. d. A pair {KLT, flipped - KLT} can be allowed / used for the {horizontal, vertical} transform or {vertical, horizontal} transform of a video unit. a. For example, {KLT, flipped-KLT} represents a pair of transform kernels that includes a first KLT and a second KLT, where the transform coefficient matrix of the second KLT (i.e., the flipped-KLT) can be the transpose matrix of the transform coefficient matrix of the first KLT. e. The horizontal or vertical transform type of a video unit can be selected from {KLT-X, flipped-KLT-X, DCT-Y, DST-Z}, where X / Y / Z are constants (e.g., Y = 8, Z = 7). a. For example, if more than one KLT kernel is defined in a codec, the more than one KLT kernel can be represented as KLT-X, such as X being an integer value, such as 1, 2, 3, …, n, i.e., KLT-1, KLT-2, KLT-3, …, KLT-n. b. For example, for a transform block, KLT can be used for horizontal (or vertical) transform, while flipped-KLT can be used for vertical (or horizontal) transform. c. For example, for a transform block, KLT can be used for horizontal (or vertical) transform, and non-KLT can be used for vertical (or horizontal) transform. d. For example, DCT2-KLT, KLT-DCT2 can be allowed for horizontal-vertical or vertical-horizontal transform of a block. e. For example, DST7-DCT2, DCT2-DST7, DCT8-DCT2, DCT2-DCT8 can be allowed for horizontal-vertical or vertical-horizontal transform of a block. f. For example, a video unit can be encoded / decoded using a specific prediction / transform / filter mode / technique. f. For a video unit encoded / decoded in a specific mode, KLT can be the only transform type. a. In one example, the specific mode can be SBT. i. In one example, for different SBT modes (SBT_horizontal_split, SBT_vertical_split, SBT_half_split, SBT_quad_split), KLT can be used for both the horizontal dimension and the vertical dimension. b. In one example, the specific mode can be MIP. c. In one example, the same KLT can be used for both the horizontal dimension and the vertical dimension of a video unit encoded / decoded in a specific mode. d. In one example, different KLTs can be used for the horizontal dimension and the vertical dimension of a video unit encoded / decoded in a specific mode. g. The transform type based on KLT can additionally be allowed for a video unit. a. For example, in addition to the existing MTS options (e.g., MTS indices from 0 to 5), the KLT-based transform type can be explicitly signaled. b. For example, a first syntax element (e.g., a flag) can be signaled to indicate whether KLT is used for a video unit. i. Additionally, alternatively, if KLT is used for a video unit, a second syntax element (e.g., a flag or an index) can be signaled to indicate whether and / or which KLT is used for the horizontal transform and / or the vertical transform. c. For example, a syntax element (e.g., an index) can be signaled to indicate which KLT is used for a video unit. i. For example, a syntax element can be signaled to indicate which KLT pair is used for the horizontal transform and the vertical transform for a video unit. ii. For example, a syntax element can be signaled to indicate which KLT is used for a specific dimension (e.g., width or height) of a video unit. iii. For example, syntax elements can be signaled for both non-KLT transforms and KLT transforms. 1. For example, indices 0 to 1 indicate a DCT2-DCT2 pair and transform skip; while indices 2 to N indicate non-DCT2-DCT2 pairs and non-transform skip pairs including combinations of non-KLT and KLT. iv. For example, syntax elements can be signaled when KLT is used for a video unit. 1. For example, an index (possible values starting from 0) can be signaled to indicate which KLT pair is used (i.e., KLT is being used in at least one of the horizontal or vertical directions). 2. For example, in this case, the intra (and / or inter) MTS index may not be signaled to the video unit (e.g., disabled). d. For example, which KLT-based transform type is used for a video unit can be implicitly determined. i. For example, the implicit determination can be based on the dimension / shape / size of the video unit. ii. For example, for a specific length of a block size (e.g., width or height), a specific KLT type is used in this direction without signaling. h. A KLT-based transform type can be applied to replace a specific existing transform type of a video unit. a. For example, the existing transform type to be replaced can be DST7, or DCT8, or DCT2. b. For example, a separable transform can be replaced by KLT. c. For example, the primary transform and / or the secondary transform can be replaced by KLT. i. For example, the KLT can be an inseparable KLT (e.g., a two-dimensional transform kernel). j. For example, the KLT can be a separable KLT (e.g., a one-dimensional transform kernel). k. For example, the video unit to which the KLT is applied can be intra-coded / decoded. l. For example, the video unit to which the KLT is applied can be inter-coded / decoded. m. For example, the KLT can be used as the primary transform. n. For example, the KLT can be used as the secondary transform. 4.4. Regarding the fourth issue of the mode-dependent KLT for the transform process, the following method is proposed: a. Which KLT kernel is used for a video unit can depend on a combination of at least one of the following coding / decoding information: a. The coding / decoding mode of the video unit. b. The size of the video unit. c. The motion vector of the video unit. d. The quantization parameter of the video unit. e. The temporal layer of the video unit. b. Which KLT kernel is used for a video unit can depend on the coding / decoding mode of the video unit. a. For example, it can be based on whether SBT is used for the video unit. b. For example, it can be based on whether implicit MTS is used for the video unit. c. For example, it can be based on whether explicit MTS is used for the video unit. d. For example, it can be based on whether intra MTS is used for the video unit. e. For example, it can be based on whether inter MTS is used for the video unit. f. For example, it can be based on whether LFNST is used for the video unit. g. For example, it can be based on whether IBC and / or its variant modes are used for the video unit. h. For example, it can be based on whether PLT and / or its variant modes are used for the video unit. i. For example, it can be based on whether the intra prediction mode is used for the video unit. j. For example, it can be based on whether ISP and / or its variant modes are used for the video unit. k. For example, it can be based on whether MIP and / or its variant modes are used for the video unit. l. For example, it can be based on whether DIMD and / or its variant modes are used for the video unit. m. For example, it can be based on whether TIMD and / or its variant modes are used for a video unit. n. For example, it can be based on whether LM / CCLM / CCCM / GLM and / or its variant modes are used for a video unit. o. For example, it can be based on whether an inter-frame prediction mode is used for a video unit. p. For example, it can be based on whether the AMVP mode is used for a video unit. q. For example, it can be based on whether the Merge mode is used for a video unit. r. For example, it can be based on whether inter-frame / intra-frame / IBC template matching and / or its variant modes are used for a video unit. s. For example, it can be based on whether DMVR and / or its variant modes are used for a video unit. t. For example, it can be based on whether a sub-block prediction mode is used for a video unit. u. For example, it can be based on whether affine and / or its variant modes are used for a video unit. v. For example, it can be based on whether sbTMVP and / or its variant modes are used for a video unit. w. For example, it can be based on whether a hybrid / fusion / multi-hypothesis mode and / or its variant modes are used for a video unit. i. In one example, it can be based on whether the hybrid / fusion / multi-hypothesis mode includes an intra-coding part, such as GPM inter-intra, GPM intra, CIIP, MHP using intra, partition GPM, etc. x. For example, it can be based on whether GPM and / or its variant modes are used for a video unit. y. For example, it can be based on whether CIIP and / or its variant modes are used for a video unit. z. For example, it can be based on whether MHP and / or its variant modes are used for a video unit. aa. For example, it can be based on whether OBMC and / or its variant modes are used for a video unit. bb. For example, it can be based on whether LIC and / or its variant modes are used for a video unit. c. Whether to use KLT and / or which KLT kernel is used for a video unit can depend on the size of the video unit. a. Different KLTs can be applied to blocks with different sizes. b. For example, it can be based on whether the width (W) and / or height (H) of a video unit satisfy predefined conditions, such as one or more combinations of the following: i. W < T1 or W <= T1, where T1 can be 8 or 16 or 32 or 64. ii. W > T2 or W >= T2, where T2 can be 2 or 4 or 8. iii. H < T3 or H <= T3, where T3 can be 8 or 16 or 32 or 64. iv. H > T4 or H >= T4, where T4 can be 2 or 4 or 8. v. W / H < T5 or W / H <= T5, where T5 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vi. W / H > T6 or W / H >= T6, where T6 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vii. H / W < T7 or H / W <= T7, where T7 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. viii. H / W > T8 or H / W >= T8, where T8 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. ix. W == T9, where T9 can be 8 or 16 or 32 or 64. x. H == T10, where T10 can be 8 or 16 or 32 or 64. d. Which KLT kernel is used for the video unit can depend on the motion vector of the video unit. a. In one example, it depends on the magnitude of the motion vector. e. Which KLT kernel is used for the video unit can depend on the quantization parameter of the video unit. a. In one example, it depends on the base QP derived / signaled at a syntax level higher than the slice level (e.g., PPS or SPS). b. In one example, it depends on the slice QP. c. In one example, it depends on the QP of the coding / decoding unit. f. Which KLT kernel is used for the video unit can depend on the temporal layer of the video unit. a. In one example, it depends on whether it is in temporal layer 0 or in temporal layer 1, or in temporal layer 2, or.... g. Depending on the coding / decoding mode and / or size of the video unit, different KLT kernels can be allowed for different video units. a. In one example, the first KLT set can be used for the first mode set, while the second KLT set can be used for the second mode set. i. For example, the first KLT set includes at least one type of KLT kernel. ii. For example, the second KLT set includes at least one type of KLT kernel. iii. For example, the first mode set includes at least one type of prediction / transformation / filter mode. iv. For example, the second mode set includes at least one prediction / transformation / filter mode. b. In one example, a video unit encoded / decoded using more than one different type of prediction / transformation / filter mode can use the same KLT kernel. c. Alternatively, video units encoded / decoded using different types of prediction / transformation / filter modes can use different KLT kernels. d. For example, one type of prediction / transformation / filter mode can be a sub-block based prediction mode (e.g., affine, sbTMVP, etc.). e. For example, one type of prediction / transformation / filter mode can be an affine based prediction mode (e.g., affine AMVP, affine Merge, etc.). f. For example, one type of prediction / transformation / filter mode can be a hybrid / fusion / multi-hypothesis based prediction mode (e.g., GPM inter-intra, GPM intra, CIIP, etc.). g. For example, one type of prediction / transformation / filter mode can be SBT and its variants. h. For example, one type of prediction / transformation / filter mode can be ISP and its variants. i. For example, one type of prediction / transformation / filter mode can be IBC and its variants. 4.5. Regarding the fifth problem of intra mode derivation / mapping for intra MTS and LFNST indices, the following method is proposed: a. The intra mode derived from the information of neighboring samples can be used for the chrominance transformation process. a. For example, in the case of single-tree and / or double-tree, the chrominance transformation process can refer to the primary transformation of the chrominance component (e.g., intra MTS, inter MTS, KLT, DCT2,...). b. For example, in the case of single-tree and / or double-tree, the chrominance transformation process can refer to the secondary transformation of the chrominance component (e.g., LFNST). c. For example, a template constructed from upper neighboring samples and / or left neighboring samples can be used to derive the intra mode. i. The gradient of the template samples can be used. ii. Construct the maximum histogram amplitude value from the gradient (e.g., histogram of gradients). iii. A DIMD-based method can be used. iv. A TIMD-based method can be used. v. The intra mode can be determined from template samples (e.g., according to measurements based on SAD / SATD cost) based on a preset intra mode candidate. vi. How many rows / columns and / or which neighboring samples can be based on a multi-reference row index. vii. How many rows / columns and / or which neighboring samples can be predefined (e.g., one row / column, or four rows / columns, etc.). viii. Neighboring samples in different rows / columns can have different weights for cost calculation. d. A derived intra mode can be generated for video units encoded / decoded using the following modes: i. CCLM mode and / or its variants. ii. CCCM mode and / or its variants. iii. GLM mode and / or its variants. iv. TIMD chrominance mode and / or its variants. v. DIMD chrominance mode and / or its variants. vi. Planar mode and / or its variants. vii. Chrominance intra modes with mode indices larger than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). e. The derived intra mode can be used for MTS transform class index derivation. f. The derived intra mode can be used for MTS transform pair index derivation. g. The derived intra mode can be used for MTS transform index derivation. h. The derived intra mode can be used for LFNST transform set index derivation. i. The derived intra mode can be used for LFNST transpose flag derivation. b. The intra mode derived from the information of neighboring samples can be used for the luminance transform process. a. For example, the luminance transform process can refer to the primary transform of the luminance component (e.g., intra MTS, inter MTS, KLT, DCT2,...). b. For example, the luminance transform process can refer to the secondary transform of the luminance component (e.g., LFNST). c. For example, a template constructed from upper neighboring samples and / or left neighboring samples can be used to derive the intra mode. i. The gradient of the template sample can be used. ii. The maximum histogram amplitude value is constructed from the gradient (e.g., gradient histogram). iii. The DIMD-based method can be used. iv. The TIMD-based method can be used. v. The intra mode can be determined from the template sample (e.g., according to the measurement based on SAD / SATD cost) based on the preset intra mode candidates. vi. How many rows / columns and / or which neighboring samples can be based on the multi-reference row index. vii. How many rows / columns and / or which neighboring samples can be predefined (e.g., one row / column, or four rows / columns, etc.). viii. The neighboring samples of different rows / columns can have different weights for cost calculation. d. The derived intra mode can be generated for video units encoded / decoded using the following modes: i. Intra template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Prediction obtained by mixing from at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content encoding / decoding tools with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. e. The derived intra mode can be used for MTS transform class index derivation. f. The derived intra mode can be used for MTS transform pair index derivation. g. The derived intra mode can be used for MTS transform index derivation. h. The derived intra mode can be used for LFNST transform set index derivation. i. The derived intra mode can be used for LFNST transpose flag derivation. c. The intra mode derived from the predicted samples of the current block can be used for the transform process. a. The predicted samples can refer to the prediction of the current video unit. b. The predicted samples can refer to the prediction of the template samples of the current video unit. c. The transform process can refer to the primary transform (e.g., intra MTS, inter MTS, KLT, DCT2,...). d. The transform process can refer to the secondary transform (e.g., LFNST). e. The derived intra mode can be generated for video units coded using the following modes: i. Intra template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Prediction obtained by mixing at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content coding / decoding tools with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. xv. CCLM mode and / or its variants. xvi. CCCM mode and / or its variants. xvii. GLM mode and / or its variants. xviii. Chrominance intra mode with a mode index larger than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). f. The derived intra mode can be used for MTS transform class index derivation. g. The derived intra mode can be used for MTS transform pair index derivation. h. The derived intra mode can be used for MTS transform index derivation. i. The derived intra mode can be used for LFNST transform set index derivation. j. The derived intra mode can be used for LFNST transpose flag derivation. d. For example, the derived intra mode can be used to index the intra MTS transform class and / or transform pair and / or transform set for the MIP-coded block. a. The derived intra mode can be based on the prediction samples of the MIP block before MIP prediction upsampling. b. The derived intra mode can be based on the prediction samples of the MIP block after MIP matrix vector multiplication. c. The derived intra mode can be based on the horizontal / vertical gradient of the prediction samples of the MIP block. d. The derived intra mode can be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the MIP block. e. The derived intra mode can be based on the neighboring sample values of the MIP block. f. Alternatively, the derived intra mode can be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. e. For example, the derived intra mode can be used to index the LFNST transform set and / or LFNST transpose flag for the block coded for intra template matching (e.g., intraTMP, intraTM, etc.). a. The derived intra mode can be based on the prediction samples of the block coded by intraTMP. b. The derived intra mode can be based on the horizontal / vertical gradient of the prediction samples of the block coded by intraTMP. c. The derived intra mode can be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block coded by intraTMP. d. The derived intra mode can be based on the neighboring sample values of the block coded by intraTMP. e. The derived intra mode can be based on the intra mode of the reference block of the block coded by intraTMP. f. The derived intra mode can be based on the gradient of the reference block of the block coded by intraTMP. g. The derived intra mode can be based on the angle of the block vector (or motion vector) of the block coded by intraTMP. h. The derived intra mode can be based on the intra mode information stored in the history-based intra mode cache for the block coded by intraTMP. i. Alternatively, the derived intra mode can be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. f. For example, the MTS index can be signaled for blocks that can be coded for intra-template matching (e.g., intraTMP, intraTM, etc.). a. For example, the intra MTS index can be signaled for blocks coded by intraTMP. b. For example, the inter MTS index can be signaled for blocks coded by intraTMP. c. The derived intra mode can be used to index the MTS transform set, MTS transform class, MTS transform pair for blocks coded by intraTMP. i. The derived intra mode can be based on the predicted samples of the block coded by intraTMP. ii. The derived intra mode can be based on the horizontal / vertical gradient of the predicted samples of the block coded by intraTMP. iii. The derived intra mode can be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block coded by intraTMP. iv. The derived intra mode can be based on the neighboring sample values of the block coded by intraTMP. v. The derived intra mode can be based on the intra mode of the reference block of the block coded by intraTMP. vi. The derived intra mode can be based on the gradient of the reference block of the block coded by intraTMP. vii. The derived intra mode can be based on the angle of the block vector (or motion vector) of the block coded by intraTMP. viii. The derived intra mode can be based on the intra mode information stored in the history-based intra mode cache for the block coded by intraTMP. ix. Alternatively, the derived intra mode can be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. d. Alternatively, an implicit MTS kernel can be applied to the blocks coded by intraTMP. i. For example, it can be determined based on the derived intra mode. ii. For example, it can be determined based on the block shape / size / dimension (e.g., width / height). iii. For example, it can be determined based on the values of the transform coefficients. iv. For example, no syntax element (e.g., index) is signaled for such an MTS kernel. e. In addition, KLT can be used for blocks coded by intraTMP. f. Alternatively, the main transform kernel other than DCT2 can be restricted for the blocks encoded / decoded by intraTMP. g. For example, the intra MTS transform class and / or the intra MTS transform pair and / or the LFNST transform set and / or the LFNST transpose flag of the TIMD hybrid mode can be derived based on the final predicted samples of the TIMD block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, TIMD information, DIMD information, etc.). b. Alternatively, it can be based on a fixed / predefined mode (e.g., other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. h. For example, the intra MTS transform class and / or the intra MTS transform pair and / or the LFNST transform set and / or the LFNST transpose flag of the DIMD hybrid mode can be derived based on the final predicted samples of the DIMD block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, etc.). b. Alternatively, it can be based on a fixed / predefined mode (e.g., other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. i. For example, the intra MTS transform class and / or the intra MTS transform pair and / or the LFNST transform set and / or the LFNST transpose flag of the planar (and / or planar horizontal and / or planar vertical) mode can be derived based on the final predicted samples of the block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, TIMD information, DIMD information, etc.). b. Alternatively, it can be based on a fixed / predefined mode (e.g., other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. 4.6. In one example, how to encode / decoderesidual blocks can depend on the selected transform. 4.7. In one example, the KLT can be separable or non-separable. a. For example, separable can refer to a one-dimensional transform kernel. b. For example, non-separable can refer to a two-dimensional transform kernel. 4.8. In one example, whether or how to apply the KLT can depend on the sum of the absolute values of the transform coefficients. c. For example, whether to allow the KLT for the blocks encoded / decoded intra (and / or inter) can be based on the sum of the absolute values of the transform coefficients of the block. d. For example, whether to allow the KLT for a specific block size / dimension / shape / orientation can be based on the sum of the absolute values of the transform coefficients of the block. e. For example, whether to apply a specific KLT-based transform kernel / set / pair to an intra-frame (and / or inter-frame) coded block can be based on the sum of the absolute values of the transform coefficients of the block. f. For example, the number of allowed KLT transform kernels / sets / pairs can be determined based on the sum of the absolute values of the transform coefficients of the block. g. For example, the sum of the absolute values of the transform coefficients of the block can be compared with at least one threshold. a. The threshold can be predefined. b. The threshold can be equal to a fixed value. c. The threshold can be adaptively determined based on predefined rules (e.g., block size / dimension, sequence resolution, prediction mode, whether it is screen content, etc.). 4.9. Regarding the sixth issue of prediction sample refinement for specific coding tools (e.g., intraTMP, IBC, etc.), the following methods are proposed: a. For example, the sample refinement process can be applied to the motion-compensated (or block-vector-compensated) prediction of a specific video block. a. For example, the specific video block can be coded by intraTMP. b. For example, the specific video block can be coded by IBC. c. For example, the block can be a screen content video block. d. For example, the block can be a camera-captured content video block. e. For example, the sample refinement process can refer to local illumination compensation (i.e., LIC). f. For example, the sample refinement process can refer to motion compensation based on overlapping sub-blocks (i.e., OBMC). g. For example, the sample refinement process can be applied to all samples within a specific block. h. For example, the sample refinement process can be applied to some samples within a specific block. i. For example, only the boundary samples at the left and / or top boundaries of a specific block can be refined. i. For example, uniform / consistent refinement parameters (e.g., weights, biases, scaling factors, etc.) can be used for all samples to be refined. j. For example, at least one sample can use refinement parameters different from those of another sample (e.g., weights, biases, scaling factors, etc.). k. For example, sample-based (e.g., position-dependent) refinement parameters (e.g., weights, biases, scaling factors, etc.) can be used. i. For example, refinement parameters can be assigned based on the position of each sample relative to the left and / or upper template (or block boundary). l. For example, refinement parameters (e.g., weights, biases, scaling factors, etc.) can be derived based on the difference / error / cost / SAD / MRSAD / SATD between samples adjacent to the current block and samples adjacent to the second block. i. For example, the second block can be pointed to by the motion vector (block vector) of the current block. ii. For example, a template-based method can be used. b. For example, a difference / error / cost measurement based on mean removal can be used as a criterion for searching for the motion vector (block vector) of a specific video block. a. For example, a specific video block can be encoded / decoded by intraTMP. b. For example, a specific video block can be encoded / decoded by inraTMP and LIC. c. For example, a specific video block can be encoded / decoded by intraTMP and OBMC. d. For example, a specific video block can be encoded / decoded by IBC. e. For example, a specific video block can be encoded / decoded by IBC and LIC. f. For example, a specific video block can be encoded / decoded by IBC and OBMC. g. For example, the block can be a screen content video block. h. For example, the block can be a camera-captured content video block. i. For example, a cost function based on mean removal can be used to measure the difference / error / cost / SAD / MRSAD / SATD between two templates / blocks to be compared. i. For example, the first template can be constructed from samples adjacent to a specific block, and the second template can be constructed from samples adjacent to the search candidate of the specific block. ii. For example, the mean removal-based method can first calculate the average / mean value between two templates. 1. For example, first accumulate all the sample differences between the first template and the second template, and then the accumulated value is further divided by the total number of samples in one template to obtain the mean (e.g., the total number of samples in the first template is equal to the total number of samples in the second template). iii. For example, when calculating the difference / error / cost / SAD / SATD, each sample difference in the sample differences can be further subtracted by the mean, and then the subtracted value is used to calculate the final difference / error / cost / SAD / MRSAD / SATD between the two templates. 4.10. Regarding the seventh question about DMVR for expanding codec tools, the following method is proposed: a. For example, DMVR refinement can be used for blocks in sub - block encoding and decoding (e.g., sbTMVP, affine, etc.). a. For example, DMVR can be a PU / CU - level DMVR process that outputs motion offsets based on PUs / CUs. b. For example, DMVR can be a sub - PU / sub - CU (e.g., 16×16, 8×8, 4×4, etc.) - level DMVR process that outputs motion offsets based on sub - PUs / sub - CUs (e.g., 16×16, 8×8, 4×4, etc.). c. For example, DMVR can be a multi - pass DMVR process that includes both a PU / CU - level DMVR process and a sub - PU / sub - CU - level DMVR process. b. For example, DMVR refinement can be applied to the motion of blocks encoded and decoded by sbTMVP. a. For example, the motion can be bi - directionally encoded and decoded. b. For example, the motion can satisfy DMVR conditions, such as pointing to a forward reference picture and a backward reference picture with the same POC distance. c. For example, the motion can be used to find a first CU - level reference block in a forward reference picture and a second CU - level reference block in a second reference picture. d. For example, the motion can be a sub - block - based motion vector to derive the prediction value of a block encoded and decoded by sub - TMVP. e. For example, to calculate the bilateral matching cost for a block encoded and decoded by sbTMVP, sub - block - based motion compensation can be performed to generate prediction values in two prediction directions. Then, the bilateral matching cost is calculated as the distortion between the two prediction values. f. For example, motion offsets can be added to each motion vector of each sub - block in a sub - block. i. For example, after a PU / CU - level reference block is retrieved, then sub - block motion vectors are derived, and it is assumed that motion offsets are added to each motion vector of each sub - block in a sub - block to perform sub - block - based motion compensation. ii. For example, the same motion offset (e.g., delta) can be added to the motion vectors of all sub - blocks in one prediction direction. iii. For example, opposite motion offsets (e.g., - delta) can be added to the motion vectors of all sub - blocks in other prediction directions. iv. For example, the motion offset can be based on the steps of DMVR refinement. g. For example, motion offsets can be added to motion candidates for finding two CU - level reference blocks. i. For example, before the PU / CU-level reference block is retrieved, a motion offset (e.g., delta) can be added to one direction of the motion candidate, and the opposite offset (e.g., -delta) can be added to the other direction of the motion candidate, and then the PU / CU-level reference block is retrieved based on the refined motion candidate. ii. For example, the motion offset can be based on the DMVR refinement step. 4.11. Regarding the eighth question of the coding / decoding tools for multi-model / hypothesis mixing, the following methods are proposed: a. For example, the prediction can be generated based on mixing the prediction of the LM-T mode and the prediction of the LM-L mode. b. For example, the prediction of the LM-TL mode can be generated by mixing the first prediction derived based on the upper neighbor and the second prediction derived based on the left neighbor. c. For example, the prediction of the LIC mode can be generated by mixing the first LIC prediction derived based on the upper neighbor / template and the second LIC prediction derived based on the left neighbor / template. d. For example, it can determine which direction of the neighbor / template (e.g., upper or left) has a major impact on the final prediction. a. For example, it can determine whether the final prediction mainly comes from the upper neighbor / template or the left neighbor / template. b. For example, if the difference / error / cost / SAD / MRSAD / SATD of the neighbor / template in one direction (e.g., upper, left) is less than (or greater than) that of the other direction, the first direction can be considered the main contributor. e. For example, the mixing weights of the two predictions can be uniform. a. For example, weight A can be assigned to all samples of the first prediction, and weight B can be assigned to all samples of the second prediction. b. For example, the value of weight A can be equal to the value of weight B. c. For example, the value of weight A can be not equal to the value of weight B. i. For example, the weight can be determined based on which prediction makes more contributions. ii. For example, a larger weight can be assigned to the prediction that makes more contributions. f. For example, the mixing weights of the two predictions can be sample-based. a. For example, for different samples, the mixing weights can be different. b. For example, weights A0, weights A1, weights A2, … can be assigned to the first predicted samples 0, sample 1, sample 2, …, and weights B0, weights B1, weights B2, … can be assigned to the second predicted samples 0, sample 1, sample 2, … c. For example, the value of the mixing weight can depend on the position. i. For example, if the position of the sample is closer to the direction determined as the main contributor (e.g., above or to the left), the mixing weight of the sample can be larger. 4.12. Regarding the ninth issue of transform design and related problems, the following method is proposed: a. NSPT can be related to block size and / or intra mode. a. For example, different block sizes can use different NSPT kernels / NSPT sets / NSPT classes / NSPT types. i. For example, each block size smaller than M×M can have its own NSPT kernel. 1. For example, if M = 8, each transform block in the 4×4, 4×8, 8×4, 8×8 transform blocks can have its own NSPT kernel. 2. For example, if M = 16, each transform block in the 4×4, 4×8, 8×4, 8×8, 4×16, 8×16, 16×4, 16×8, 16×16 transform blocks can have its own NSPT kernel. 3. For example, for block sizes larger than M×M, at least two block sizes can share an NSPT kernel. b. For example, different intra modes can use different NSPT kernels / NSPT sets / NSPT classes / NSPT types. i. For example, at least two different NSPT kernels / NSPT sets / NSPT classes / NSPT types can be used for two different intra mode encoding / decoding blocks. ii. For example, each intra mode index can refer to a specific NSPT kernel. iii. For example, at least two intra modes can share an NSPT kernel. iv. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks encoded / decoded by MIP. v. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks encoded / decoded by ISP. vi. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks encoded / decoded by intra-intra fusion modes (e.g., DIMD mixing, TIMD mixing, SGPM, intra luminance fusion, etc.). vii. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks that can be coded for intra - inter fusion modes (e.g., GPM inter - intra, CIIP, etc.). viii. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks that can be coded for intra TMP. ix. For example, special NSPT kernels / NSPT sets / NSPT classes / NSPT types can be defined for blocks that can be coded for IBC. b. ISP blocks / ISP sub - partitions / ISP sub - blocks can be allowed to use multiple transform sets (e.g., MTS), and / or NSPT, and / or non - separable transforms. a. For example, different sub - partitions / sub - blocks of an ISP block can have different intra modes. i. For example, a derived intra mode can be obtained for an ISP sub - partition / ISP sub - block. 1. For example, a derived intra mode can be used to obtain a specific kernel of MTS / NSPT / LFNST. 2. For example, a derived intra mode can be obtained based on the prediction of an ISP block. 3. For example, a derived intra mode can be obtained based on the prediction (and / or reconstruction) of neighboring samples of an ISP block. 4. For example, a derived intra mode can be obtained based on gradients and / or gradient histograms, etc. ii. For example, an intra mode transmitted by signal can be obtained for an ISP sub - partition / ISP sub - block. b. For example, different sub - partitions / sub - blocks of an ISP block can use different MTS types. c. For example, different sub - partitions / sub - blocks of an ISP block can use different LFNST types. d. For example, different sub - partitions / sub - blocks of an ISP block can use different NSPT types. e. For example, which transform type / transform pair / transform class / transform set / transform kernel of MTS is used can depend on the intra mode and / or block dimension and / or color component of the ISP block / ISP sub - partition / ISP sub - block. f. For example, which transform type / transform pair / transform class / transform set / transform kernel of NSPT is used can depend on the intra mode and / or block dimension and / or color component of the ISP block / ISP sub - partition / ISP sub - block. g. For example, which transform type / transform pair / transform class / transform set / transform kernel of the LFNST to be used can depend on the intra mode and / or block dimension and / or color component of the ISP block / ISP sub - partition / ISP sub - block. h. For example, special NSPT / LFNST / MTS / transform class / transform set can be defined specifically for the blocks encoded / decoded by ISP. i. The transform type / transform pair / transform class / transform set / transform kernel / transform rule for the ISP block / ISP sub - partition / ISP sub - block can be different from that for the blocks encoded / decoded in another mode (other than ISP). ii. For example, a video unit encoded / decoded in another intra mode / technique (different from ISP) may not use the same transform type / transform pair / transform class / transform set / transform kernel / transform rule as the blocks / sub - partitions / sub - blocks encoded / decoded by ISP. i. A video unit can refer to a luminance block / luminance sub - partition / luminance sub - block. j. A video unit can refer to a chrominance block / chrominance sub - partition / chrominance sub - block. c. The blocks / sub - partitions / sub - blocks encoded / decoded by ISP can use transform skip. a. For example, transform skip can be allowed for the blocks / sub - partitions / sub - blocks encoded / decoded by ISP. b. For example, a CU - level transform skip flag can be signaled for the entire ISP block, indicating whether all sub - blocks / sub - partitions use transform skip. i. For example, in this case, if transform skip is used, each ISP sub - block / sub - partition (if having non - zero coefficients) should use transform skip. c. For example, a sub - block / sub - partition - level transform skip flag can be signaled for each individual ISP sub - block / sub - partition, indicating whether each specific sub - block / sub - partition uses transform skip. i. For example, in this case, one ISP sub - block / sub - partition (if having non - zero coefficients) uses transform skip, and another ISP sub - block / sub - partition can use a different transform type (e.g., DCT2, DST7, etc.). d. The blocks encoded / decoded by SGPM can use NSPT. a. For example, special NSPT classes / NSPT sets can be defined specifically for the blocks encoded / decoded by SGPM. b. For example, which NSPT to be used can depend on the intra mode and / or block dimension and / or color component of the SGPM mode. i. For example, the intra mode can be derived based on decoding information (e.g., the gradient of neighboring samples / HOG of neighboring samples, the gradient of predicted samples / HOG of predicted samples, partitioning method, partitioning angle, distance from the partitioning line, etc.). c. For example, the luma component of a video unit coded / decoded by SGPM can be coded / decoded by NSPT. d. For example, the chroma component of a video unit coded / decoded by SGPM can be coded / decoded by DCT-2. e. Blocks coded / decoded in the intra blending / fusion mode (e.g., TIMD blending mode, and / or DIMD blending mode, and / or intra luma fusion mode) can use NSPT. a. For example, special NSPT types / NSPT pairs / NSPT classes / NSPT sets / NSPT kernels / NSPT rules can be defined specifically for blocks coded / decoded in the TIMD blending mode. b. For example, special NSPT types / NSPT pairs / NSPT classes NSPT sets / NSPT kernels / NSPT rules can be defined specifically for blocks coded / decoded in the DIMD blending mode. c. For example, special NSPT types / NSPT pairs / NSPT classes NSPT sets / NSPT kernels / NSPT rules can be defined specifically for blocks coded / decoded in the intra luma fusion blending mode. d. The NSPT types / NSPT pairs / NSPT classes NSPT sets / NSPT kernels / NSPT rules for blocks coded / decoded in a specific mode can be different from those for blocks coded / decoded in another mode (other than the specific mode). f. Blocks coded / decoded in the intra-inter blending / fusion mode (such as CIIP, GPM intra-inter) can follow the transformation rules of intra blocks for transformation. a. For example, intra MTS can be used for the intra-inter blending / fusion mode. b. For example, LFNST can be used for the intra-inter blending / fusion mode. c. For example, NSPT can be used for the intra-inter blending / fusion mode. d. For example, the above transformation can be used for the luma component. g. The derived intra mode can be obtained for the intra blending / fusion mode and / or the intra-inter blending / fusion mode, and then be used to derive specific MTS / LFNST / NSPT types / pairs / classes / sets / kernels / rules. a. For example, the intra mode can be derived based on decoding information (e.g., the gradient of neighboring samples / HOG of neighboring samples, the gradient of predicted samples / HOG of predicted samples, etc.). h. Allow at least two different transforms for a specific mode. a. Which transform is used / allowed for a specific mode may depend on whether the video unit belongs to screen content or content captured by a camera. i. For example, if the video unit belongs to screen content, an inseparable transform kernel type can be used as the primary transform for encoding / decoding blocks for a specific mode, while if the video unit belongs to content captured by a camera, a separable transform kernel type can be used as the primary transform for the blocks. ii. For example, it can be controlled by syntax elements signaled at a level higher than the block level, such as slice level / picture level / sub-picture level / SH level / PH level / PPS level / SPS level. iii. For example, it can be implicitly determined based on decoding information (e.g., based on content analysis such as sample gradients). b. For example, a special transform can be defined for screen content blocks. c. For example, the transform can refer to the primary transform (e.g., DCT2, MTS, NSPT, intra MTS, inter MTS, etc.). d. For example, the transform can refer to the secondary transform (e.g., LFNST). e. For example, the transform can refer to a separable transform. f. For example, the transform can refer to an inseparable transform. g. For example, the transform can refer to the intra MTS transform type / intra MTS transform pair / intra MTS transform class / intra MTS transform set / intra MTS transform kernel. h. For example, the transform can refer to the inter MTS transform type / inter MTS transform pair / inter MTS transform class / inter MTS transform set / inter MTS transform kernel. i. For example, the transform can refer to the LFNST type / LFNST pair / LFNST class / LFNST set / LFNST kernel. j. For example, the transform can refer to the NSPT type / NSPT pair / NSPT class / NSPT set / NSPT kernel. k. For example, the allowed transforms can be predefined. l. For example, the specific mode can be ISP. m. For example, the specific mode can be MIP. n. For example, the specific mode can be SGPM (e.g., GPM intra-intra). o. For example, the specific mode can be DIMD and / or DIMD hybrid / fusion mode. p. For example, the specific mode can be TIMD and / or TIMD hybrid / fusion mode. q. For example, a specific mode may be an intra - luminance fusion mode. r. For example, a specific mode may be an intra - TMP mode. s. For example, a specific mode may be an IBC mode. t. For example, a specific mode may be an inter - intra / IBC hybrid mode (e.g., GPM inter - intra, CIIP, MHP, etc.). 4.13. Whether and / or how to apply the methods disclosed above can be signaled at the sequence level / group - of - pictures level / picture level / strip level / slice group level, e.g., in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.14. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub - picture / other types of regions containing more than one sample or pixel. 4.15. The binary bits of the syntax elements signaled in the disclosed methods can be context - decoded or bypass - decoded. 4.16. Whether and / or how to apply the methods disclosed above can depend on the decoded information, such as block size, color format, single / double - tree splitting, color component, strip / picture type.

[0104] More details of embodiments of the present disclosure related to transforms in image / video coding, screen content coding (SCC), affine, deduplication checking, prediction refinement with motion compensation (MC), and weighted - based schemes will be described below. The embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow way. In addition, these embodiments can be applied individually or in combination in any way.

[0105] As used herein, the term "block" may represent a color component, sub - picture, picture, strip, slice, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), sub - blocks of a video block, sub - regions within a video block, video processing units including multiple samples / pixels, etc. The block can be rectangular or non - rectangular.

[0106] Figure 46 A flowchart of a method 4600 for video processing according to some embodiments of the present disclosure is shown. Method 4600 can be implemented during the conversion between the current block of a video and the bit - stream of the video. As Figure 46As shown, method 4600 starts at 4602, where information related to the non-separable primary transform (NSPT) applied to the current block is obtained. This information depends on the block size of the current block and / or the intra mode for the current block. In some embodiments, the current block can be a transform block or the like.

[0107] In some embodiments, the information can include the transform kernel of the NSPT, the transform set of the NSPT, the transform class of the NSPT, the transform type of the NSPT, the transform pair of the NSPT, the transform rule of the NSPT, etc. It should be understood that the above examples are only described for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0108] In some embodiments, the information can be determined at the encoder and signaled to the decoder. Alternatively, the information can be determined at both the encoder and the decoder.

[0109] At 4604, a transformation is performed based on this information. In some embodiments, the transformation can include encoding the current block into the bitstream. Alternatively or additionally, the transformation can include decoding the current block from the bitstream.

[0110] In view of the above, the information related to the NSPT depends on the block size of the current block and / or the intra mode of the current block. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency and the encoding and decoding quality.

[0111] In some embodiments, the information can be different for different block sizes. For example, for multiple block sizes smaller than M×M, the corresponding transform kernels of the NSPT can be different, and M is an integer. In one example, if M is equal to 8, the corresponding transform kernels of the NSPT for the block sizes of 4×4, 4×8, 8×4, and 8×8 are different. If M is equal to 16, the corresponding transform kernels of the NSPT for the block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 8×16, 16×4, 16×8, and 16×16 are different. Additionally or alternatively, at least two block sizes larger than M×M can share the same transform kernel of the NSPT.

[0112] In some additional or alternative embodiments, the information can be different for different intra modes. In one example, at least two different transform kernels of the NSPT can be used for two different intra mode encoding and decoding blocks. In another example, at least two different transform sets of the NSPT are used for two different intra mode encoding and decoding blocks. In another example, at least two different transform classes of the NSPT are used for two different intra mode encoding and decoding blocks. In yet another example, at least two different transform types of the NSPT are used for two different intra mode encoding and decoding blocks.

[0113] In some embodiments, the intra mode index may be associated with the transform kernel of the NSPT. For example, each intra mode index may refer to a specific NSPT kernel. In some additional embodiments, at least two intra modes may share the same transform kernel of the NSPT.

[0114] In some embodiments, the current block may be coded / decoded using a first mode, and at least one of the following may be allowed for the current block: a predetermined transform kernel of the NSPT, a predetermined transform set of the NSPT, a predetermined transform class of the NSPT, or a predetermined transform type of the NSPT. For example, a special NSPT kernel, NSPT set, NSPT class, and / or NSPT type may be defined for the first mode.

[0115] As an example, the first mode may be a matrix weighted intra prediction (MIP) mode, an intra sub - partition (ISP) mode, an intra - intra fusion mode, a decoder - side intra mode derivation (DIMD) hybrid mode, a template - based intra mode derivation (TIMD) hybrid mode, a spatial geometry partition mode (SGPM), an intra - luminance fusion mode, an intra - inter fusion mode, a GPM inter - intra mode, an intra - inter joint prediction (CIIP) mode, an intra TMP mode, an intra block copy (IBC) mode, etc. As used herein, the intra TMP mode may refer to a template - matching - based prediction mode for intra - coded / decoded blocks.

[0116] In some embodiments, the current block may be coded / decoded using an intra hybrid mode or an intra - inter hybrid mode. Additionally, a first intra mode may be determined for the current block and the first intra mode is used to determine the transform type of the multiple transform selection (MTS), the transform pair of the MTS, the transform class of the MTS, the transform set of the MTS, the transform kernel of the MTS, the transform rule of the MTS, the transform type of the low - frequency non - separable transform (LFNST), the transform pair of the LFNST, the transform class of the LFNST, the transform set of the LFNST, the transform kernel of the LFNST, the transform rule of the LFNST, the transform type of the NSPT, the transform pair of the NSPT, the transform class of the NSPT, the transform set of the NSPT, the transform kernel of the NSPT, the transform rule of the NSPT, etc. In some embodiments, the first intra mode may be different from the coding / decoding mode used to determine the prediction of the current block.

[0117] In some embodiments, the first intra mode may be determined based on the coding / decoding information associated with the current block. By way of example and not limitation, the coding / decoding information may include the gradient of neighboring samples of the current block, the histogram of oriented gradients (HOG) of neighboring samples, the gradient of predicted samples of the current block, and / or the HOG of predicted samples.

[0118] In some embodiments, a current block may be encoded or decoded using an intra-frame blending mode. For example, the intra-frame blending mode may be a TIMD blending mode, a DIMD blending mode, an intra-frame luminance fusion mode, etc. Additionally, a predetermined transform type of NSPT, a predetermined transform pair of NSPT, a predetermined transform class of NSPT, a predetermined transform kernel of NSPT, a predetermined transform set of NSPT, and / or a predetermined transform rule of NSPT may be used for the current block. In one example, a special NSPT type, NSPT pair, NSPT class, NSPT set, NSPT kernel, or NSPT rule may be defined for a block that can be encoded or decoded specifically for the TIMD blending mode. In another example, a special NSPT type, NSPT pair, NSPT class, NSPT set, NSPT kernel, or NSPT rule may be defined for a block that can be encoded or decoded specifically for the DIMD blending mode. In yet another example, a special NSPT type, NSPT pair, NSPT class, NSPT set, NSPT kernel, or NSPT rule may be defined for a code block that can be decoded specifically for the intra-frame luminance fusion blending mode.

[0119] In some embodiments, a current block may be encoded or decoded using a first mode. Additionally, information related to the NSPT applied to the current block may be different from information related to the NSPT applied to another block that is encoded or decoded using a second mode different from the first mode. For example, the NSPT type, NSPT pair, NSPT class, NSPT set, NSPT kernel, or NSPT rule for encoding or decoding a block for a specific mode may be different from the NSPT type, NSPT pair, NSPT class, NSPT set, NSPT kernel, or NSPT rule for encoding or decoding a block for another mode (other than the specific mode).

[0120] In some embodiments, a one-dimensional transform kernel may be used for a separable Karhunen-Loeve transform (KLT) for the current block. Additionally or alternatively, a two-dimensional transform kernel may be used for an inseparable KLT for the current block.

[0121] In some embodiments, the current block may include an ISP block or an ISP sub-division. In such a case, a multi-transform set (e.g., MTS), NSPT, and / or an inseparable transform may be allowed for the current block.

[0122] In some embodiments, different intra modes may be used for different sub - partitions of the ISP block. For example, a second intra mode may be determined for the ISP sub - partition. The second intra mode may be used to determine the transform kernel of MTS, the transform kernel of NSPT, the transform kernel of LFNST, etc. By way of example and not limitation, the second intra mode may be determined based on the prediction of the ISP block, the prediction of neighboring samples of the ISP block, the reconstruction of neighboring samples, the gradient of neighboring samples, the HOG of neighboring samples, the gradient of predicted samples of the ISP block, and / or the HOG of predicted samples. In some alternative embodiments, the second intra mode of the ISP sub - partition may be indicated in the bitstream and obtained from the bitstream.

[0123] In some embodiments, different MTS types may be used for different sub - partitions of the ISP block. Additionally or alternatively, different LFNST types may be used for different sub - partitions of the ISP block. In some additional or alternative embodiments, different NSPT types may be used for different sub - partitions of the ISP block.

[0124] In some embodiments, the transform type, transform pair, transform class, transform set, and / or transform kernel of MTS used for the ISP block may depend on the intra mode, block dimension, and / or color component of the ISP block. Additionally or alternatively, the transform type, transform pair, transform class, transform set, and / or transform kernel of MTS used for the ISP sub - partition may depend on the intra mode, block dimension, and / or color component of the ISP sub - partition.

[0125] In some embodiments, the transform type, transform pair, transform class, transform set, and / or transform kernel of NSPT used for the ISP block may depend on the intra mode, block dimension, and / or color component of the ISP block. Additionally or alternatively, the transform type, transform pair, transform class, transform set, and / or transform kernel of NSPT used for the ISP sub - partition may depend on the intra mode, block dimension, and / or color component of the ISP sub - partition.

[0126] In some embodiments, the transform type, transform pair, transform class, transform set, and / or transform kernel of LFNST used for the ISP block may depend on the intra mode, block dimension, and / or color component of the ISP block. Additionally or alternatively, the transform type, transform pair, transform class, transform set, and / or transform kernel of LFNST used for the ISP sub - partition may depend on the intra mode, block dimension, and / or color component of the ISP sub - partition.

[0127] In some embodiments, the current block may be coded and decoded using the ISP. In this case, a predetermined NSPT, a predetermined LFNST, a predetermined MTS, a predetermined transform class, and / or a predetermined transform set may be allowed for the current block. For example, special NSPT, LFNST, MTS, transform classes, and transform sets may be defined specifically for blocks coded and decoded using the ISP. Additionally, the transform type, transform pair, transform class, transform set, transform kernel, and / or transform rule used for the current block may be different from that of another block coded and decoded using a mode different from the ISP. By way of example and not limitation, the other block may include a luminance block, a luminance sub - partition, a luminance sub - block, a chrominance block, a chrominance sub - partition, and / or a chrominance sub - block.

[0128] In some embodiments, the current block may be an ISP - coded block, an ISP - coded sub - block, or an ISP - coded sub - partition. In this case, the transform skip mode may be allowed to be applied to the current block. For example, in the transform skip mode, the vertical transform and / or the horizontal transform may be skipped.

[0129] In some embodiments, a coding unit (CU) - level transform skip flag may be indicated in the bitstream for the current block. The transform skip flag indicates whether the transform skip mode is applied to all sub - blocks or all sub - partitions of the current block. For example, if the transform skip mode is used for the current block, the transform skip mode may be applied to each ISP sub - block or each ISP sub - partition of the current block. Alternatively, if the transform skip mode is used for the current block, the transform skip mode may be applied to each ISP sub - block or each ISP sub - partition of the current block having non - zero coefficients.

[0130] In some embodiments, a sub - block - level transform skip flag may be indicated in the bitstream for each individual ISP sub - block. The sub - block - level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub - block. Alternatively, a sub - partition - level transform skip flag may be indicated in the bitstream for each individual ISP sub - partition. The sub - partition - level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub - partition.

[0131] In some embodiments, the transform skip mode may be applied to an ISP sub - block, and different transform types (e.g., DCT2, DST7, etc.) may be applied to another ISP sub - block. In some additional embodiments, the transform skip mode may be applied to an ISP sub - partition, and different transform types may be applied to another ISP sub - partition.

[0132] In some embodiments, the current block may be encoded and decoded using the SGPM mode. In this case, a predetermined transform set of the NSPT and / or a predetermined transform class of the NSPT may be allowed for the current block. For example, a special NSPT class or NSPT set may be defined specifically for blocks encoded and decoded in the SGPM mode.

[0133] In some embodiments, the NSPT for the current block may depend on the intra mode for the current block, the block dimension of the current block, and / or the color component of the current block. For example, the intra mode may be determined based on the encoding and decoding information associated with the current block. By way of example, the encoding and decoding information may include the gradient of neighboring samples of the current block, the HOG of neighboring samples, the gradient of predicted samples of the current block, the HOG of predicted samples, the partitioning method for the current block, the partitioning angle for the current block, and / or the distance from the partitioning line for the current block. It should be understood that the above examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0134] In some embodiments, the luminance component of the current block may be encoded and decoded using the NSPT. Additionally or alternatively, the chrominance component of the current block may be encoded and decoded using the discrete cosine transform type 2 (DCT-2).

[0135] In some embodiments, the current block may be encoded and decoded using an intra-inter frame hybrid mode. In this case, the transform may be applied to the current block based on the same transform rules used for intra-encoded blocks. For example, blocks encoded and decoded in the intra-inter frame hybrid mode may follow the transform rules of intra blocks for transformation. As an example, the intra-inter frame hybrid mode may include CIIP, GPM intra-inter, etc.

[0136] In some embodiments, intra MTS, LFNST, and / or NSPT may be applied to the current block encoded and decoded in the intra-inter frame hybrid mode. Alternatively, intra MTS, LFNST, and / or NSPT may be applied to the luminance component of the current block encoded and decoded in the intra-inter frame hybrid mode.

[0137] In some embodiments, the current block may be encoded and decoded using a first mode. Additionally, at least two different transforms may be allowed for the current block. As an example, at least two different transforms may be predetermined.

[0138] For example, the information related to the transform used for the current block and / or the transforms allowed for the current block may depend on whether the current block belongs to screen content or camera-captured content.

[0139] For example, if the current block belongs to screen content, an in-separable transform kernel type can be used as the primary transform for the current block. Otherwise, if the current block belongs to content captured by a camera, a separable transform kernel type can be used as the primary transform for the current block.

[0140] In some embodiments, information related to the transform used for the current block and / or the transform allowed for the current block can depend on syntax elements signaled at a first level higher than the block level. By way of example and not limitation, the first level can be a slice level, a picture level, a sub-picture level, a slice header (SH) level, a picture header (PH) level, a picture parameter set (PPS) level, a sequence parameter set (SPS) level, or any other suitable level higher than the block level.

[0141] In some embodiments, information related to the transform used for the current block and / or the transform allowed for the current block can be determined based on the information being coded and decoded. For example, the information can be determined based on content analysis (such as sample gradients).

[0142] In some embodiments, a predetermined transform can be allowed for screen content blocks. For example, special transforms can be defined for screen content blocks.

[0143] In some embodiments, the transform can include a primary transform, DCT2, MTS, NSPT, intra MTS, inter MTS, secondary transform, LFNST, separable transform, in-separable transform, intra MTS transform type, intra MTS transform pair, intra MTS transform class, intra MTS transform set, intra MTS transform kernel, inter MTS transform type, inter MTS transform pair, MTS transform class, MTS transform set, MTS transform kernel, LFNST transform type, LFNST transform pair, LFNST transform class, LFNST transform set, LFNST transform kernel, NSPT transform type, NSPT transform pair, NSPT transform class, NSPT transform set, NSPT transform kernel, etc.

[0144] In some embodiments, the first mode can include ISP, MIP, SGPM, GPM intra-intra mode, DIMD, DIMD hybrid mode, TIMD, TIMD hybrid mode, intra luminance fusion mode, intra TMP mode, IBC mode, inter-intra / IBC hybrid mode, GPM inter-inter mode, CIIP mode, multiple hypothesis prediction (MHP) mode, etc.

[0145] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by an apparatus for video processing. In the method, information related to an in-separable primary transform (NSPT) applied to a current block of a video is obtained. The information depends on at least one of a block size of the current block or an intra mode for the current block. Further, the bitstream is generated based on the information.

[0146] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, information related to an in-separable primary transform (NSPT) applied to a current block of a video is obtained. The information depends on at least one of a block size of the current block or an intra mode for the current block. Further, the bitstream is generated based on the information, and the bitstream is stored in a non-transitory computer-readable recording medium.

[0147] Embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.

[0148] Item 1. A method for video processing, comprising: obtaining information related to an in-separable primary transform (NSPT) applied to a current block of a video for conversion between the current block of the video and a bitstream of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and performing the conversion based on the information.

[0149] Item 2. The method according to Item 1, wherein the information includes at least one of the following: a transform kernel of the NSPT, a transform set of the NSPT, a transform class of the NSPT, a transform type of the NSPT, a transform pair of the NSPT, or a transform rule of the NSPT.

[0150] Item 3. The method according to any one of Items 1 to 2, wherein the information is different for different block sizes.

[0151] Item 4. The method according to any one of Items 2 to 3, wherein for a plurality of block sizes smaller than M×M, corresponding transform kernels of the NSPT are different, and M is an integer.

[0152] Item 5. The method according to Item 4, wherein if M is equal to 8, corresponding transform kernels of the NSPT are different for block sizes of 4×4, 4×8, 8×4, and 8×8, or if M is equal to 16, corresponding transform kernels of the NSPT are different for block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 8×16, 16×4, 16×8, and 16×16.

[0153] Item 6. The method according to any one of Items 4 to 5, wherein at least two block sizes greater than M×M share the same transform kernel of the NSPT.

[0154] Item 7. The method according to any one of Items 1 to 6, wherein the information is different for different intra modes.

[0155] Item 8. The method according to Item 7, wherein at least two different transform kernels of the NSPT are used for two different intra mode coded blocks, or wherein at least two different transform sets of the NSPT are used for two different intra mode coded blocks, or wherein at least two different transform classes of the NSPT are used for two different intra mode coded blocks, or wherein at least two different transform types of the NSPT are used for two different intra mode coded blocks.

[0156] Item 9. The method according to any one of Items 7 to 8, wherein the intra mode index is associated with the transform kernel of the NSPT.

[0157] Item 10. The method according to any one of Items 7 to 9, wherein at least two intra modes share the same transform kernel of the NSPT.

[0158] Item 11. The method according to any one of Items 1 to 10, wherein the current block is coded using a first mode, and at least one of the following is allowed for the current block: a predetermined transform kernel of the NSPT, a predetermined transform set of the NSPT, a predetermined transform class of the NSPT, or a predetermined transform type of the NSPT.

[0159] Item 12. The method according to Item 11, wherein the first mode includes one of the following: matrix weighted intra prediction (MIP) mode, intra sub - partition (ISP) mode, intra - intra fusion mode, decoder - side intra mode derivation (DIMD) hybrid mode, template - based intra mode derivation (TIMD) hybrid mode, spatial domain geometry partition mode (SGPM), intra luminance fusion mode, intra - inter fusion mode, GPM inter - intra mode, combined intra - inter prediction (CIIP) mode, intra TMP mode, or intra block copy (IBC) mode.

[0160] Item 13. The method according to any one of Items 1 to 12, wherein the current block is encoded and decoded using an intra-frame hybrid mode or an intra-frame / inter-frame hybrid mode, and a first intra-frame mode is determined for the current block, and the first intra-frame mode is used to determine at least one of the following: the transform type of multi-transform selection (MTS), the transform pair of the MTS, the transform class of the MTS, the transform set of the MTS, the transform kernel of the MTS, the transform rule of the MTS; the transform type of low-frequency non-separable transform (LFNST), the transform pair of the LFNST, the transform class of the LFNST, the transform set of the LFNST, the transform kernel of the LFNST, the transform rule of the LFNST, the transform type of the NSPT, the transform pair of the NSPT, the transform class of the NSPT, the transform set of the NSPT, the transform kernel of the NSPT, or the transform rule of the NSPT.

[0161] Item 14. The method according to Item 13, wherein the first intra-frame mode is determined based on the encoding and decoding information associated with the current block.

[0162] Item 15. The method according to Item 14, wherein the encoding and decoding information includes at least one of the following: the gradient of neighboring samples of the current block, the histogram of oriented gradients (HOG) of the neighboring samples, the gradient of predicted samples of the current block, or the HOG of the predicted samples.

[0163] Item 16. The method according to any one of Items 1 to 15, wherein the current block is encoded and decoded using an intra-frame hybrid mode.

[0164] Item 17. The method according to Item 16, wherein the intra-frame hybrid mode includes one of the following: TIMD hybrid mode, DIMD hybrid mode, or intra-frame luminance fusion mode.

[0165] Item 18. The method according to any one of Items 16 to 17, wherein at least one of the following is used for the current block: the predetermined transform type of the NSPT, the predetermined transform pair of the NSPT, the predetermined transform class of the NSPT, the predetermined transform kernel of the NSPT, the predetermined transform set of the NSPT, or the predetermined transform rule of the NSPT.

[0166] Item 19. The method according to any one of Items 1 to 18, wherein the current block is encoded and decoded using a first mode, and the information related to the NSPT applied to the current block is different from the information related to the NSPT applied to another block encoded and decoded using a second mode different from the first mode.

[0167] Item 20. The method according to any one of Items 1 to 19, wherein a one-dimensional transform kernel is used for the separable Karhunen-Loeve transform (KLT) of the current block, or a two-dimensional transform kernel is used for the non-separable KLT of the current block.

[0168] Item 21. The method according to any one of Items 1 to 12, wherein the current block includes an ISP block or an ISP sub-division, and at least one of the following is allowed for the current block: a multi-transform set, the NSPT, or a non-separable transform.

[0169] Item 22. The method according to Item 21, wherein different intra modes are used for different sub-divisions of the ISP block.

[0170] Item 23. The method according to any one of Items 21 to 22, wherein a second intra mode is determined for the ISP sub-division.

[0171] Item 24. The method according to Item 23, wherein the second intra mode is used to determine one of the following: the transform kernel of the MTS, the transform kernel of the NSPT, or the transform kernel of the LFNST.

[0172] Item 25. The method according to any one of Items 23 to 24, wherein the second intra mode is determined based on at least one of the following: the prediction of the ISP block, the prediction of the neighboring samples of the ISP block, the reconstruction of the neighboring samples, the gradient of the neighboring samples, the HOG of the neighboring samples, the gradient of the predicted samples of the ISP block, or the HOG of the predicted samples.

[0173] Item 26. The method according to any one of Items 21 to 22, wherein the second intra mode for the ISP sub-division is indicated in the bitstream and obtained from the bitstream.

[0174] Item 27. The method according to any one of Items 21 to 26, wherein different MTS types are used for different sub-divisions of the ISP block, or different LFNST types are used for different sub-divisions of the ISP block, or different NSPT types are used for different sub-divisions of the ISP block.

[0175] Item 28. The method according to any one of Items 21 to 27, wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the MTS used for the ISP block depends on at least one of an intra mode, a block dimension, or a color component of the ISP block, or wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the MTS used for the ISP sub - division depends on at least one of an intra mode, a block dimension, or a color component of the ISP sub - division.

[0176] Item 29. The method according to any one of Items 21 to 28, wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the NSPT used for the ISP block depends on at least one of an intra mode, a block dimension, or a color component of the ISP block, or wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the NSPT used for the ISP sub - division depends on at least one of an intra mode, a block dimension, or a color component of the ISP sub - division.

[0177] Item 30. The method according to any one of Items 21 to 29, wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the LFNST used for the ISP block depends on at least one of an intra mode, a block dimension, or a color component of the ISP block, or wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the LFNST used for the ISP sub - division depends on at least one of an intra mode, a block dimension, or a color component of the ISP sub - division.

[0178] Item 31. The method according to any one of Items 1 to 12, wherein the current block is encoded and decoded using the ISP, and at least one of the following is allowed for the current block: a predetermined NSPT, a predetermined LFNST, a predetermined MTS, a predetermined transform class, or a predetermined transform set.

[0179] Item 32. The method according to Item 31, wherein at least one of the following used for the current block is different from that of another block encoded and decoded using a mode different from the ISP: a transform type, a transform pair, a transform class, a transform set, a transform kernel, or a transform rule.

[0180] Item 33. The method according to Item 31, wherein the other block includes one of the following: a luma block, a luma sub - division, a luma sub - block, a chroma block, a chroma sub - division, or a chroma sub - block.

[0181] Item 34. The method according to any one of Items 1 to 12, wherein a transform skip mode is allowed to be applied to the current block, and the current block includes one of the following: an ISP encoded / decoded block, an ISP encoded / decoded sub-block, or an ISP encoded / decoded sub-partition.

[0182] Item 35. The method according to Item 34, wherein in the transform skip mode, at least one of the following transforms is skipped: a vertical transform, or a horizontal transform.

[0183] Item 36. The method according to any one of Items 34 to 35, wherein a coding unit (CU) level transform skip flag is indicated in the bitstream for the current block, and the transform skip flag indicates whether the transform skip mode is applied to all sub-blocks or all sub-partitions of the current block.

[0184] Item 37. The method according to Item 36, wherein if the transform skip mode is used for the current block, the transform skip mode is applied to each ISP sub-block or each ISP sub-partition of the current block, or wherein if the transform skip mode is used for the current block, the transform skip mode is applied to each ISP sub-block or each ISP sub-partition of the current block having non-zero coefficients.

[0185] Item 38. The method according to any one of Items 34 to 35, wherein a sub-block level transform skip flag is indicated in the bitstream for each individual ISP sub-block, and the sub-block level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub-block, or wherein a sub-partition level transform skip flag is indicated in the bitstream for each individual ISP sub-partition, and the sub-partition level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub-partition.

[0186] Item 39. The method according to Item 38, wherein the transform skip mode is applied to an ISP sub-block, and a different transform type is applied to another ISP sub-block, or wherein the transform skip mode is applied to an ISP sub-partition, and a different transform type is applied to another ISP sub-partition.

[0187] Item 40. The method according to any one of Items 1 to 12, wherein the current block is encoded / decoded using the SGPM mode.

[0188] Item 41. The method according to Item 40, wherein at least one of the following is allowed for the current block: a predetermined transform set of the NSPT, or a predetermined transform class of the NSPT.

[0189] Item 42. The method according to any one of Items 40 to 41, wherein the NSPT for the current block depends on at least one of the following: the intra mode for the current block, the block dimension of the current block, or the color component of the current block.

[0190] Item 43. The method according to Item 42, wherein the intra mode is determined based on the codec information associated with the current block.

[0191] Item 44. The method according to Item 43, wherein the codec information includes at least one of the following: the gradient of the neighboring samples of the current block, the HOG of the neighboring samples, the gradient of the predicted samples of the current block, the HOG of the predicted samples, the partitioning method for the current block, the partitioning angle for the current block, or the distance from the partitioning line for the current block.

[0192] Item 45. The method according to any one of Items 40 to 44, wherein the luminance component of the current block is coded and decoded using NSPT.

[0193] Item 46. The method according to any one of Items 40 to 45, wherein the chrominance component of the current block is coded and decoded using discrete cosine transform type 2 (DCT-2).

[0194] Item 47. The method according to any one of Items 1 to 12, wherein the current block is coded and decoded using an intra-inter intermixing mode, and the transform is applied to the current block based on the same transform rules for intra-coded blocks.

[0195] Item 48. The method according to Item 47, wherein the intra-inter intermixing mode includes at least one of CIIP or GPM intra-inter.

[0196] Item 49. The method according to any one of Items 47 to 48, wherein at least one of intra MTS, LFNST, or the NSPT is applied to the current block, or wherein at least one of the intra MTS, the LFNST, or the NSPT is applied to the luminance component of the current block.

[0197] Item 50. The method according to any one of Items 1 to 10, wherein the current block is coded and decoded using a first mode, and at least two different transforms are allowed for the current block.

[0198] Item 51. The method according to Item 50, wherein at least one of the following depends on whether the current block belongs to screen content or camera-captured content: the transform for the current block, or the transforms allowed for the current block.

[0199] Item 52. The method according to Item 51, wherein if the current block belongs to the screen content, a non-separable transform kernel type is used as the primary transform for the current block, or if the current block belongs to the content captured by the camera, a separable transform kernel type is used as the primary transform for the current block.

[0200] Item 53. The method according to Item 50, wherein at least one of the following depends on a syntax element signaled at a first level higher than the block level: the transform for the current block, or the transforms allowed for the current block.

[0201] Item 54. The method according to Item 53, wherein the first level includes one of the following: slice level, picture level, sub-picture level, slice header (SH) level, picture header (PH) level, picture parameter set (PPS) level, or sequence parameter set (SPS) level.

[0202] Item 55. The method according to Item 50, wherein at least one of the following is determined based on the information being encoded or decoded: the transform for the current block, or the transforms allowed for the current block.

[0203] Item 56. The method according to any one of Items 50 to 55, wherein a predetermined transform is allowed for screen content blocks.

[0204] Item 57. The method according to any one of Items 50 to 56, wherein the transform includes at least one of the following: primary transform, DCT2, MTS, NSPT, intra MTS, inter MTS, secondary transform, LFNST, separable transform, non-separable transform, intra MTS transform type, intra MTS transform pair, intra MTS transform class, intra MTS transform set, intra MTS transform kernel, inter MTS transform type, inter MTS transform pair, inter MTS transform class, inter MTS transform set, inter MTS transform kernel, LFNST transform type, LFNST transform pair, LFNST transform class, LFNST transform set, LFNST transform kernel, NSPT transform type, NSPT transform pair, NSPT transform class, NSPT transform set, or NSPT transform kernel.

[0205] Item 58. The method according to any one of Items 50 to 57, wherein the at least two different transforms are predetermined.

[0206] Item 59. The method according to any one of Items 50 to 58, wherein the first mode includes one of the following: ISP, MIP, SGPM, GPM Intra-Intra mode, DIMD, DIMD Hybrid mode, TIMD, TIMD Hybrid mode, Intra Luma Blending mode, Intra TMP mode, IBC mode, Inter-Intra / IBC Hybrid mode, GPM Inter-Inter mode, CIIP mode, or Multiple Hypothesis Prediction (MHP) mode.

[0207] Item 60. The method according to any one of Items 1 to 59, wherein the current block is a transform block.

[0208] Item 61. The method according to any one of Items 1 to 60, wherein the transformation includes encoding the current block into the bitstream.

[0209] Item 62. The method according to any one of Items 1 to 60, wherein the transformation includes decoding the current block from the bitstream.

[0210] Item 63. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 62.

[0211] Item 64. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 62.

[0212] Item 65. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method includes: obtaining information related to an Inseparable Primary Transform (NSPT) applied to a current block of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and generating the bitstream based on the information.

[0213] Item 66. A method for storing a bitstream of a video, including: obtaining information related to an Inseparable Primary Transform (NSPT) applied to a current block of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0214] Figure 47FIG. shows a block diagram of a computing device 4700 in which various embodiments of the present disclosure may be implemented. The computing device 4700 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).

[0215] It should be understood that Figure 47 the computing device 4700 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.

[0216] As Figure 47 shown, the computing device 4700 includes a general computing device 4700. The computing device 4700 may include at least one or more processors or processing units 4710, a memory 4720, a storage unit 4730, one or more communication units 4740, one or more input devices 4750, and one or more output devices 4760.

[0217] In some embodiments, the computing device 4700 may be implemented as any user terminal or server terminal with computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is conceivable that the computing device 4700 may support any type of interface to the user (such as "wearable" circuitry, etc.).

[0218] The processing unit 4710 may be a physical processor or a virtual processor, and may implement various processes based on the programs stored in the memory 4720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4700. The processing unit 4710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0219] The computing device 4700 generally includes various computer storage media. Such media can be any media accessible by the computing device 4700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4720 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 4730 can be any removable or non-removable media and can include machine-readable media, such as a memory, flash drive, magnetic disk, or other media that can be used to store information and / or data and can be accessed in the computing device 4700.

[0220] The computing device 4700 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 47 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0221] The communication unit 4740 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4700 may be implemented by a single computing cluster or multiple computer machines, which may communicate via a communication connection. Thus, the computing device 4700 may operate in a networked environment using a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0222] The input device 4750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 4760 can be one or more of various output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 4740, the computing device 4700 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 4700 can also communicate with one or more devices that enable a user to interact with the computing device 4700, or if needed, the computing device 4700 can also communicate with any device that enables the computing device 4700 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be carried out via an input / output (I / O) interface (not shown).

[0223] In some embodiments, some or all components of computing device 4700 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. Software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at the location of a remote data center. Cloud computing infrastructure may provide services through a shared data center, although to a user they appear as a single access point. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0224] In an embodiment of the present disclosure, computing device 4700 may be used to implement video encoding / decoding. Memory 4720 may include one or more video codec modules 4725 having one or more program instructions. These modules are accessible and executable by processing unit 4710 to perform the functions of the various embodiments described herein.

[0225] In an example embodiment of performing video encoding, input device 4750 may receive video data as input 4770 to be encoded. The video data may be processed, for example, by video codec module 4725 to generate an encoded bitstream. The encoded bitstream may be provided as output 4780 via output device 4760.

[0226] In an example embodiment of performing video decoding, input device 4750 may receive the encoded bitstream as input 4770. The encoded bitstream may be processed, for example, by video codec module 4725 to generate decoded video data. The decoded video data may be provided as output 4780 via output device 4760.

[0227] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: Obtaining information related to an Inseparable Main Transform (NSPT) applied to a current block of a video for conversion between the current block of the video and a bitstream of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and Performing the conversion based on the information.

2. The method according to claim 1, wherein the information includes at least one of the following: A transform kernel of the NSPT, A transform set of the NSPT, A transform class of the NSPT, A transform type of the NSPT, A transform pair of the NSPT, or A transform rule of the NSPT.

3. The method according to any one of claims 1 to 2, wherein the information is different for different block sizes.

4. The method according to any one of claims 2 to 3, wherein for a plurality of block sizes smaller than M×M, corresponding transform kernels of the NSPT are different, and M is an integer.

5. The method according to claim 4, wherein if M is equal to 8, corresponding transform kernels of the NSPT are different for block sizes of 4×4, 4×8, 8×4, and 8×8, or if M is equal to 16, corresponding transform kernels of the NSPT are different for block sizes of 4×4, 4×8, 8×4, 8×8, 4×16, 8×16, 16×4, 16×8, and 16×16.

6. The method according to any one of claims 4 to 5, wherein at least two block sizes larger than M×M share the same transform kernel of the NSPT.

7. The method according to any one of claims 1 to 6, wherein the information is different for different intra modes.

8. The method according to claim 7, wherein at least two different transform kernels of the NSPT are used for two different intra mode coding / decoding blocks, or wherein at least two different transform sets of the NSPT are used for two different intra mode coding / decoding blocks, or wherein at least two different transform classes of the NSPT are used for two different intra mode coding / decoding blocks, or wherein at least two different transform types of the NSPT are used for two different intra mode coding / decoding blocks.

9. The method according to any one of claims 7 to 8, wherein an intra mode index is associated with a transform kernel of the NSPT.

10. The method according to any one of claims 7 to 9, wherein at least two intra modes share the same transform kernel of the NSPT.

11. The method according to any one of claims 1 to 10, wherein the current block is coded / decoded using a first mode, and at least one of the following is allowed for the current block: A predetermined transform kernel of the NSPT, A predetermined transform set of the NSPT, A predetermined transform class of the NSPT, or A predetermined transform type of the NSPT.

12. The method according to claim 11, wherein the first mode includes one of the following: Matrix-Weighted Intra Prediction (MIP) mode, Intra Sub-Partition (ISP) mode, Intra-Intra Fusion mode, Decoder-side Intra Mode Derivation (DIMD) Hybrid Mode Template-based Intra Mode Derivation (TIMD) Hybrid Mode Spatial Geometry Partitioning Mode (SGPM) Intra Luminance Fusion Mode Intra-Inter Frame Fusion Mode GPM Inter-Intra Mode Concurrent Intra-Inter Prediction (CIIP) Mode Intra TMP Mode, or Intra Block Copy (IBC) Mode 13. The method according to any one of claims 1 to 12, wherein the current block is encoded or decoded using an intra hybrid mode or an intra-inter hybrid mode, and a first intra mode is determined for the current block, and the first intra mode is used to determine at least one of the following: The transform type of Multiple Transform Selection (MTS) The transform pair of the MTS The transform class of the MTS The transform set of the MTS The transform kernel of the MTS The transform rule of the MTS The transform type of Low Frequency Non-Separable Transform (LFNST) The transform pair of the LFNST The transform class of the LFNST The transform set of the LFNST The transform kernel of the LFNST The transform rule of the LFNST The transform type of the NSPT The transform pair of the NSPT The transform class of the NSPT The transform set of the NSPT The transform kernel of the NSPT, or The transform rule of the NSPT 14. The method according to claim 13, wherein the first intra mode is determined based on encoding and decoding information associated with the current block.

15. The method according to claim 14, wherein the encoding and decoding information includes at least one of the following: The gradient of neighboring samples of the current block The Histogram of Oriented Gradients (HOG) of the neighboring samples The gradient of predicted samples of the current block, or The HOG of the predicted samples 16. The method according to any one of claims 1 to 15, wherein the current block is encoded or decoded using an intra hybrid mode.

17. The method according to claim 16, wherein the intra hybrid mode includes one of the following: TIMD Hybrid Mode DIMD Hybrid Mode, or Intra Luminance Fusion Mode 18. The method according to any one of claims 16 to 17, wherein at least one of the following is used for the current block: The predetermined transform type of the NSPT The predetermined transform pair of the NSPT The predetermined transform class of the NSPT The predetermined transform kernel of the NSPT The predetermined transform set of the NSPT, or The predetermined transform rule of the NSPT 19. The method according to any one of claims 1 to 18, wherein the current block is encoded or decoded using a first mode, and the information related to the NSPT applied to the current block is different from the information related to the NSPT applied to another block encoded or decoded using a second mode different from the first mode.

20. The method according to any one of claims 1 to 19, wherein a one-dimensional transform kernel is used for the separable Karhunen-Loeve Transform (KLT) of the current block, or The two-dimensional transform kernel is used for the non-separable KLT of the current block.

21. The method according to any one of claims 1 to 12, wherein the current block includes an ISP block or an ISP sub-partition, and at least one of the following is allowed for the current block: A multi-transform set, The NSPT, or A non-separable transform.

22. The method according to claim 21, wherein different intra modes are used for different sub-partitions of the ISP block.

23. The method according to any one of claims 21 to 22, wherein a second intra mode is determined for the ISP sub-partition.

24. The method according to claim 23, wherein the second intra mode is used to determine one of the following: The transform kernel of the MTS, The transform kernel of the NSPT, or The transform kernel of the LFNST.

25. The method according to any one of claims 23 to 24, wherein the second intra mode is determined based on at least one of the following: The prediction of the ISP block, The prediction of neighboring samples of the ISP block, The reconstruction of the neighboring samples, The gradient of the neighboring samples, The HOG of the neighboring samples, The gradient of the predicted samples of the ISP block, or The HOG of the predicted samples.

26. The method according to any one of claims 21 to 22, wherein the second intra mode for the ISP sub-partition is indicated in the bitstream and obtained from the bitstream.

27. The method according to any one of claims 21 to 26, wherein different MTS types are used for different sub-partitions of the ISP block, or wherein different LFNST types are used for different sub-partitions of the ISP block, or wherein different NSPT types are used for different sub-partitions of the ISP block.

28. The method according to any one of claims 21 to 27, wherein at least one of the transform type, transform pair, transform class, transform set, or transform kernel of the MTS used for the ISP block depends on at least one of the intra mode, block dimension, or color component of the ISP block, or wherein at least one of the transform type, transform pair, transform class, transform set, or transform kernel of the MTS used for the ISP sub-partition depends on at least one of the intra mode, block dimension, or color component of the ISP sub-partition.

29. The method according to any one of claims 21 to 28, wherein at least one of the transform type, transform pair, transform class, transform set, or transform kernel of the NSPT used for the ISP block depends on at least one of the intra mode, block dimension, or color component of the ISP block, or wherein at least one of the transform type, transform pair, transform class, transform set, or transform kernel of the NSPT used for the ISP sub-partition depends on at least one of the intra mode, block dimension, or color component of the ISP sub-partition.

30. The method according to any one of claims 21 to 29, wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the LFNST used for the ISP block depends on at least one of an intra mode, a block dimension, or a color component of the ISP block, or wherein at least one of a transform type, a transform pair, a transform class, a transform set, or a transform kernel of the LFNST used for the ISP sub - partition depends on at least one of an intra mode, a block dimension, or a color component of the ISP sub - partition.

31. The method according to any one of claims 1 to 12, wherein the current block is encoded and decoded using the ISP, and at least one of the following is allowed for the current block: a predetermined NSPT, a predetermined LFNST, a predetermined MTS, a predetermined transform class, or a predetermined transform set.

32. The method according to claim 31, wherein at least one of the following used for the current block is different from that of another block encoded and decoded using a mode different from the ISP: transform type, transform pair, transform class, transform set, transform kernel, or transform rule.

33. The method according to claim 31, wherein the other block includes one of the following: a luminance block, a luminance sub - partition, a luminance sub - block, a chrominance block, a chrominance sub - partition, or a chrominance sub - block.

34. The method according to any one of claims 1 to 12, wherein a transform skip mode is allowed to be applied to the current block, and the current block includes one of the following: an ISP - encoded and decoded block, an ISP - encoded and decoded sub - block, or an ISP - encoded and decoded sub - partition.

35. The method according to claim 34, wherein in the transform skip mode, at least one of the following transforms is skipped: a vertical transform, or a horizontal transform.

36. The method according to any one of claims 34 to 35, wherein a coding unit (CU) - level transform skip flag is indicated in the bitstream for the current block, and the transform skip flag indicates whether the transform skip mode is applied to all sub - blocks or all sub - partitions of the current block.

37. The method according to claim 36, wherein if the transform skip mode is used for the current block, the transform skip mode is applied to each ISP sub - block or each ISP sub - partition of the current block, or wherein if the transform skip mode is used for the current block, the transform skip mode is applied to each ISP sub - block or each ISP sub - partition of the current block having non - zero coefficients.

38. The method according to any one of claims 34 to 35, wherein a sub - block - level transform skip flag is indicated in the bitstream for each individual ISP sub - block, and the sub - block - level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub - block, or wherein a sub - partition - level transform skip flag is indicated in the bitstream for each individual ISP sub - partition, and the sub - partition - level transform skip flag indicates whether the transform skip mode is applied to the individual ISP sub - partition.

39. The method according to claim 38, wherein the transform skip mode is applied to an ISP sub-block, and different transform types are applied to another ISP sub-block, or wherein the transform skip mode is applied to an ISP sub-division, and different transform types are applied to another ISP sub-division.

40. The method according to any one of claims 1 to 12, wherein the current block is encoded and decoded using the SGPM mode.

41. The method according to claim 40, wherein at least one of the following is allowed for the current block: a predetermined transform set of the NSPT, or a predetermined transform class of the NSPT.

42. The method according to any one of claims 40 to 41, wherein the NSPT for the current block depends on at least one of the following: the intra mode for the current block, the block dimension of the current block, or the color component of the current block.

43. The method according to claim 42, wherein the intra mode is determined based on the encoding and decoding information associated with the current block.

44. The method according to claim 43, wherein the encoding and decoding information includes at least one of the following: the gradient of the neighboring samples of the current block, the HOG of the neighboring samples, the gradient of the predicted samples of the current block, the HOG of the predicted samples, the partitioning method for the current block, the partitioning angle for the current block, or the distance from the partitioning line for the current block.

45. The method according to any one of claims 40 to 44, wherein the luminance component of the current block is encoded and decoded using the NSPT.

46. The method according to any one of claims 40 to 45, wherein the chrominance component of the current block is encoded and decoded using the discrete cosine transform type 2 (DCT-2).

47. The method according to any one of claims 1 to 12, wherein the current block is encoded and decoded using an intra-inter hybrid mode, and the transform is applied to the current block based on the same transform rules for the intra-encoded block.

48. The method according to claim 47, wherein the intra-inter hybrid mode includes at least one of CIIP or GPM intra-inter.

49. The method according to any one of claims 47 to 48, wherein at least one of intra MTS, LFNST or the NSPT is applied to the current block, or wherein at least one of the intra MTS, the LFNST or the NSPT is applied to the luminance component of the current block.

50. The method according to any one of claims 1 to 10, wherein the current block is encoded and decoded using a first mode, and at least two different transforms are allowed for the current block.

51. The method according to claim 50, wherein at least one of the following depends on whether the current block belongs to screen content or camera-captured content: the transform for the current block, or the transforms allowed for the current block.

52. The method according to claim 51, wherein if the current block belongs to the screen content, an inseparable transform kernel type is used as the primary transform for the current block, or if the current block belongs to the content captured by the camera, a separable transform kernel type is used as the primary transform for the current block.

53. The method according to claim 50, wherein at least one of the following depends on a syntax element signaled at a first level higher than the block level: the transform for the current block, or the transforms allowed for the current block.

54. The method according to claim 53, wherein the first level comprises one of the following: slice level, picture level, sub - picture level, slice header (SH) level, picture header (PH) level, picture parameter set (PPS) level, or sequence parameter set (SPS) level.

55. The method according to claim 50, wherein at least one of the following is determined based on the information being encoded or decoded: the transform for the current block, or the transforms allowed for the current block.

56. The method according to any one of claims 50 to 55, wherein a predetermined transform is allowed for screen content blocks.

57. The method according to any one of claims 50 to 56, wherein the transform comprises at least one of the following: primary transform, DCT2, MTS, NSPT, intra MTS, inter MTS, secondary transform, LFNST, separable transform, inseparable transform, intra MTS transform type, intra MTS transform pair, intra MTS transform class, intra MTS transform set, intra MTS transform kernel, inter MTS transform type, inter MTS transform pair, inter MTS transform class, inter MTS transform set, inter MTS transform kernel, LFNST transform type, LFNST transform pair, LFNST transform class, LFNST transform set, LFNST transform kernel, NSPT transform type, NSPT transform pair, NSPT transform class, NSPT transform set, or NSPT transform kernel.

58. The method according to any one of claims 50 to 57, wherein the at least two different transforms are predetermined.

59. The method according to any one of claims 50 to 58, wherein the first mode comprises one of the following: ISP, MIP, SGPM, GPM intra - intra mode, DIMD, DIMD hybrid mode, TIMD, TIMD hybrid mode, intra - luminance fusion mode, intra - TMP mode, IBC mode, inter - intra / IBC hybrid mode, GPM inter - inter mode, CIIP mode, or multiple hypothesis prediction (MHP) mode.

60. The method according to any one of claims 1 to 59, wherein the current block is a transform block.

61. The method according to any one of claims 1 to 60, wherein the transformation comprises encoding the current block into the bitstream.

62. The method according to any one of claims 1 to 60, wherein the transformation comprises decoding the current block from the bitstream.

63. A device for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 62.

64. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1 to 62.

65. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by a device for video processing for a video, wherein the method comprises: obtaining information related to an in-separable primary transform (NSPT) applied to a current block of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; and generating the bitstream based on the information.

66. A method for storing a bitstream of a video, comprising: obtaining information related to an in-separable primary transform (NSPT) applied to a current block of the video, the information depending on at least one of a block size of the current block or an intra mode for the current block; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.