Method and device for video processing and medium

By using the adjacent sample point information and prediction information of the video block in video encoding and decoding, and applying the transformation process or KLT transformation, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and more efficient video processing is achieved.

CN120077648APending Publication Date: 2025-05-30DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073482.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has shortcomings in improving the encoding and decoding efficiency, especially when processing the intra-frame mode and transformation process of video blocks.

Method used

Intra mode is determined by taking into account the adjacent sample point information and prediction information of the video block, and applying the transformation process based on the mode, or using the Karl Huning-Love transform (KLT) to optimize the transformation coefficient of the video block.

Benefits of technology

Improve the efficiency of video encoding and decoding, and reduce the size and codec time of bitstreams through more accurate intra-mode determination and transformation process optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077648A_ABST
    Figure CN120077648A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a transition between a current video block of the video and a bitstream of the video, obtaining an intra mode for the current video block, the intra mode being determined based on at least one of: information associated with neighboring samples of the current video block, or prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and performing the transition based on the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to video encoding and decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is generally a desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a current video block of a video and the bitstream of the video, obtaining an intra mode for the current video block, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and performing the conversion based on the application.

[0005] According to the method of the first aspect of the present disclosure, the intra mode for the transform process is determined by considering information associated with neighboring samples of the video block and / or prediction associated with the video block. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency.

[0006] In a second aspect, another method for video processing is proposed. The method includes: for the conversion between a current video block of a video and the bitstream of the video, obtaining information regarding applying a Karhunen-Loeve transform (KLT) to the current video block, the information depending on the transform coefficients of the current video block; and performing the conversion based on the information.

[0007] According to the method of the second aspect of the present disclosure, the information regarding applying KLT to the video block depends on the transform coefficients of the video block. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency.

[0008] In a third aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0009] In a fourth aspect, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0010] In a fifth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and generating a bitstream based on the application.

[0011] In a sixth aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.

[0012] In a seventh aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: obtaining information about applying a Karhunen-Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; and generating a bitstream based on the information.

[0013] In an eighth aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining information about applying a Karhunen-Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; generating a bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] The present invention content is provided to introduce a selection of concepts further described below in the detailed description in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the example embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0016] Figure 1 A block diagram showing an example video coding and decoding system according to some embodiments of the present disclosure is shown;

[0017] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure is shown;

[0018] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure is shown;

[0019] Figure 4 The positions of spatial Merge candidates are shown;

[0020] Figure 5 Candidate pairs considered for redundancy check of spatial Merge candidates are shown;

[0021] Figure 6 Motion vector scaling for temporal Merge candidates is shown;

[0022] Figure 7 Candidate positions C 0 and C 1 ;

[0023] Figure 8 MMVD search points are shown;

[0024] Figure 9 The extended CU regions used in BDOF are shown;

[0025] Figure 10 Symmetric MVD modes are shown;

[0026] Figure 11 An affine motion model based on control points is shown;

[0027] Figure 12 The affine MVF for each sub-block is shown;

[0028] Figure 13 The positions of inherited affine motion prediction values are shown;

[0029] Figure 14 Control point motion vector inheritance is shown;

[0030] Figure 15 The positions of candidate positions for constructing affine Merge modes are shown;

[0031] Figure 16 It is a diagram used for the motion vectors of the proposed combination method;

[0032] Figure 17 It shows the sub-block MV VSB and the pixel Δv(i, j);

[0033] Figure 18A It shows the spatial neighboring blocks used by ATVMP;

[0034] Figure 18B It shows the derivation of the sub-CU motion field by applying the motion displacement from the spatial neighbors and scaling the motion information from the corresponding co-located sub-CU;

[0035] Figure 19 It shows the extended CU region used in BDOF;

[0036] Figure 20 It shows the motion vector refinement on the decoding side;

[0037] Figure 21 It shows the top neighboring block and the left neighboring block used in the CIIP weight derivation;

[0038] Figure 22 It shows an example of the GPM division grouped at the same angle;

[0039] Figure 23 It shows the unidirectional prediction MV selection for the geometric segmentation mode;

[0040] Figure 24 It shows the use of the bending weight w for the geometric segmentation mode 0 exemplary generation;

[0041] Figure 25 It shows the spatial neighboring blocks for deriving the spatial Merge candidates;

[0042] Figure 26 It shows the execution of template matching on the search region around the initial MV;

[0043] Figure 27 It shows the diamond region in the search region;

[0044] Figure 28 It shows the frequency responses of the interpolation filter and the VVC interpolation filter at the half-pixel phase;

[0045] Figure 29 It shows the template and the reference sample points of the template in the reference picture;

[0046] Figure 30Shows templates and reference sample points for blocks with sub - block motion using motion information of sub - blocks of the current block;

[0047] Figure 31 Shows filling candidates for replacing zero vectors in the IBC list.

[0048] Figure 32 Shows the IBC reference region depending on the current CU position;

[0049] Figure 33 Shows the reference region for IBC when CTU(m, n) is encoded / decoded. The blue block represents the current CTU; the green block represents the reference region; and the white block represents the invalid reference region;

[0050] Figure 34 Shows the first HPT and the second HPT;

[0051] Figure 35 Shows the spatial neighbors for deriving affine Merge candidates / AMVP candidates;

[0052] Figure 36 Shows the constructed affine Merge candidates / AMVP candidates from non - adjacent neighbors to the first type;

[0053] Figure 37 Shows the low - frequency non - separable transform (LFNST) process;

[0054] Figure 38 Shows the SBT position, type, and transform type;

[0055] Figure 39 Shows the ROI for LFNST 16;

[0056] Figure 40 Shows the ROI for LFNST 8;

[0057] Figure 41 Shows the discontinuity measurement;

[0058] Figure 42 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure;

[0059] Figure 43 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and

[0060] Figure 44 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.

[0061] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed implementation manners

[0062] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is only for the purpose of illustration and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the manner described below.

[0063] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.

[0064] As used herein, the terms "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0065] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0066] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein indicate the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example environment

[0067] Figure 1FIG. 0 is a block diagram showing an example video coding and decoding system 100 that may utilize the techniques of the present disclosure. As shown, video coding and decoding system 100 may include a source device 110 and a destination device 120. Source device 110 may also be referred to as a video encoding device, and destination device 120 may also be referred to as a video decoding device. In operation, source device 110 may be configured to generate encoded video data, and destination device 120 may be configured to decode the encoded video data generated by source device 110. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0068] Video source 112 may include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0069] The video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are the coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to destination device 120 via I / O interface 116 through network 130A. The encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0070] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 may include a receiver and / or a modulator. I / O interface 126 may obtain the encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, which is configured to interface with an external display device.

[0071] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.

[0072] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.

[0073] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0074] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.

[0075] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0076] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are shown separately in Figure 2 an example.

[0077] The segmentation unit 201 may divide a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0078] The mode selection unit 203 may select one of a plurality of coding / decoding modes (intra coding / decoding or inter coding / decoding), for example, based on error results, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0079] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0080] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to portions of a picture composed of macroblocks independent of macroblocks in the same picture.

[0081] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0082] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find one reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, where the plurality of reference indices indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0083] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0084] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, which value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0085] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0086] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0087] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0088] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0089] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0090] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0091] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0092] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0093] After the reconstruction unit 212 reconstructs the video block, a loop filter operation can be performed to reduce block effect artifacts in the video block.

[0094] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0095] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.

[0096] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.

[0097] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.

[0098] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and the motion compensation unit 302 may determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the "Merge mode" may refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.

[0099] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.

[0100] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to received syntax information, and the motion compensation unit 302 may use the interpolation filter to generate a prediction block.

[0101] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.

[0102] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0103] The reconstruction unit 306 can obtain the decoded block, for example, by adding a residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307, and the cache 307 provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also generates the decoded video for presentation on a display device.

[0104] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video coding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview Embodiments of the present disclosure relate to video coding techniques. Specifically, it is about coding and decoding techniques for transformation, screen content coding and decoding, and local illumination compensation in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, ECM, etc. It is also applicable to future video coding and decoding standards or video codecs. 2. Introduction Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly simultaneously, and in the April 2018 JVET meeting, the new video coding standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at this time. Then the VVC working draft and the test model VTM are updated after each meeting. The VVC project achieved Feature Complete (FDIS) at the July 2020 meeting. In January 2021, JVET established the Exploration Experiments (EE) aiming at enhanced compression efficiency beyond the capabilities of VVC using novel conventional algorithms. Soon after, ECM was built as a common software library for the long-term exploration work towards the next-generation video coding standard. 2.1. Inter-frame Prediction Coding Tools For each inter-frame prediction CU, the motion parameters include the motion vector, the reference picture index and the reference picture list use index, and additional information required for the new decoding features of VVC that will be used for inter-frame prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When the CU is coded using the skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded motion vector difference or reference picture index. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs (including spatial candidates and temporal candidates, as well as additional lists introduced in VVC). The Merge mode can be applied to any inter-frame prediction CU, not just the skip mode. An alternative to the Merge mode is the explicit transmission of the motion parameters, where the motion vector, the corresponding reference picture index and the reference picture list use flag for each reference picture list, and other required information are explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC includes multiple new and refined inter-frame prediction coding tools listed as follows: - Extended Merge Prediction; - Merge Mode with MVD (MMVD); - Symmetric MVD (SMVD) signaling; - Affine motion compensation prediction; - Sub-block based temporal motion vector prediction (SbTMVP); - Adaptive motion vector resolution (AMVR); - Motion field storage: 1 / 16 luma sample MV storage and 8×8 motion field compression; - Bi-directional prediction with CU-level weights (BCW); - Bi-directional optical flow (BDOF); - Decoder-side motion vector refinement (DMVR); - Geometric partitioning mode (GPM); - Combined inter and intra prediction (CIIP). The following text provides details on those inter prediction methods specified in VVC. 2.1.1. Extended Merge prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVP from spatially neighboring CUs; 2) Temporal MVP from co-located CUs; 3) History-based MVP from the FIFO table; 4) Pairwise-averaged MVP; 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in the Merge mode, rounding-unary binary coding (TU) is used to code the index of the best Merge candidate. The first binary bit of the Merge index is coded using context, and bypass coding is used for the other binary bits. The derivation process for each type of Merge candidate is provided in this session. Similar to what was done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain sized region. 2.1.1.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except for swapping the positions of the first two Merge candidates. Among the candidates at the positions depicted in Figure 4 select up to four Merge candidates. The derivation order is B 0 , A 0 , B 1 , A 1 and B 2 . Only when the position B 0, A 0 , B 1 , A 1 When one or more CUs of A (for example, because it belongs to another strip or slice) are unavailable or are intra-coded, position B 2 is considered. After a candidate at position A 1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, so as to improve the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrows in Figure 5 are considered, and a candidate is added to the list only if the corresponding candidate for the redundancy check does not have the same motion information. 2.1.1.2. Temporal Candidate Derivation In this step, only one candidate is added to the list. Specifically, when deriving the temporal Merge candidate at this time, the scaled motion vector is derived based on the co-located CUs belonging to the co-located reference picture. The reference picture list to be used for deriving the co-located CUs is explicitly signaled in the slice header. As shown by the dashed line in Figure 6 , the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. As depicted in Figure 7 , position C 0 for the temporal candidate is selected between the candidates 1 and C. If the CU at position C 0 is unavailable, intra-coded, or outside the current row of the CTU, position C 1 is used. Otherwise, position C 0 is used to derive the temporal Merge candidate. Merge Candidate Derivation Based on History After the spatial MVP and TMVP, the Merge candidates of the Motion Vector Prediction based on History (HMVP) are added to the Merge list. In this method, the motion information of the previously coded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a CU that is non-sub-block inter-coded, the associated motion information is added to the last entry of the table as a new HMVP candidate. The size S of the HMVP table is set to 6, which indicates that at most 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if the same HMVP exists in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are shifted forward. HMVP candidates can be used in the Merge candidate list construction process. Check the latest several HMVP candidates in the table in order and insert them into the candidate list after the TMVP candidates. Apply a redundancy check to the HMVP candidates for spatial or temporal Merge candidates. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list construction process from HMVP is terminated. 2.1.1.3. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices into the Merge candidate list. The average motion vectors are calculated separately for each reference list. If both motion vectors are available in a list, the two motion vectors are averaged even if they point to different reference images; if only one motion vector is available, that one vector is used directly; if no motion vector is available, the list is kept invalid. When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached. 2.1.1.4. Merge estimation region The Merge Estimation Region (MER) allows for the independent derivation of the Merge candidate list for a CU within the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for the generation of the Merge candidate list for the current CU. Additionally, the update process for the history-based motion vector prediction value candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.1.2. Merge Mode with MVD (MMVD) In addition to the Merge mode, in the case where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge Mode with Motion Vector Difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, an index for specifying the motion size, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The Merge candidate flag is signaled to specify which one is used. The distance index specifies the motion size information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1. Table 1 - Relationship between Distance Index and Predefined Offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate the four directions shown in Table 2. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a uni-directional prediction MV or a bi-directional prediction MV where two of the lists point to the same side of the current picture (i.e., both of the two referenced POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), Error! No symbol found in the reference source. Specify the symbol of the MV offset added to the starting MV. When the starting MV is a bi-directional prediction MV with two MVs pointing to different sides of the current picture (i.e., one referenced POC is greater than the POC of the current picture and the other referenced POC is less than the POC of the current picture), the symbol in Table 2 specifies the symbol of the MV offset added to the list 0 MV component of the starting MV, and the symbol of the list 1 MV has the opposite value. Table 2 - Symbols of MV Offsets Specified by the Direction Index Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.1. Bi-Directional Prediction with CU-Level Weights (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 1) Five weights are allowed in weighted average bi-directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) for non-Merge CUs, the weight index is signaled after the motion vector difference; 2) for Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used. - At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-delay picture, unequal weights are only conditionally checked for 1-pixel and 4-pixel motion vector precisions. - When combined with affine, if and only if the affine mode is selected as the current best mode, affine ME will be performed for unequal weights. - When the two reference pictures in bidirectional prediction are the same, unequal weights are only conditionally checked. - When specific conditions are met, unequal weights are not searched for, which depends on the POC distance between the current picture and its reference pictures, the coding / decoding QP, and the temporal level. The BCW weight index is coded using a context - coded bit followed by a bypass - coded bit. The first context - coded bit indicates whether equal weights are used; and if unequal weights are used, the bypass - coded bit signals additional bits to indicate the use of unequal weights. Weighted Prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding / decoding of video content in fading situations. Support for WP has also been added in the VVC standard. WP allows signaling of weighted parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.2.2. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously called BIO, is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub - block level. BDOF is applied to a CU if the CU meets all of the following conditions: - The CU is coded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order. - The distances from two reference pictures to the current picture (i.e., the POC differences) are the same. - Both of the two reference pictures are short-term reference pictures. - The CU is coded or decoded without using the affine mode or the ATMVP Merge mode. - The CU has more than 64 luma samples. - Both the CU height and the CU width are greater than or equal to 8 luma samples. - The BCW weight index indicates equal weights. - WP is not enabled for the current CU. - The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of the object is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bi-predicted sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two predicted signals are calculated by directly computing the differences between two neighboring samples, and k = 0, 1, i.e., where I (k) (i, j) is the sample value at the coordinate (i, j) of the predicted signal in the list k (k = 0, 1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the auto-correlations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as: where where Ω is a 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then the motion refinement (v x , v y): where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4×4 sub-block: Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points as follows: pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o offset ) >> shift (2 - 7) These values are selected such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some predicted sample points I (k) (i, j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 9 depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating the predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma sample points, it is divided into sub-blocks with a width and / or height equal to 16 luma sample points, and the sub-block boundaries are considered as CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 predicted sample points and the L1 predicted sample points is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8 * W * (H >> 1)), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated during the DVMR process between the initial L0 predicted sample points and the L1 predicted sample points is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. When a CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, BDOF is also disabled. 2.1.2.3. Symmetric MVD Coding (SMVD) In VVC, in addition to normal uni-directional prediction and bi-directional prediction mode MVD signaling, a symmetric MVD mode for bi-directional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indices for both list 0 and list 1 and the MVD for list 1 is not signaled, but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. - Otherwise, if the nearest reference picture in list -0 and the nearest reference picture in list -1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, then BiDirPredFlag is set to 1, and both the list -0 reference picture and the list -1 reference picture are short-term reference pictures. Otherwise, BiDirPred- Flag is set to 0. 2) At the CU level, if the CU is bi-directionally predicted and encoded and BiDirPredFlag is equal to 1, then the symmetric mode flag indicating whether to use the symmetric mode is signaled explicitly. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indices for list 0 and list 1 are set to be equal to the reference picture pair respectively. MVD1 is set to be equal to (-MVD0). The final motion vector is as follows. Figure 10 The symmetric MVD mode is shown. In the encoder, the symmetric MVD motion estimation starts with an initial MV evaluation. A set of initial MV candidates includes the MVs obtained from uni-directional prediction search, the MVs obtained from bi-directional prediction search, and the MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.3. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensated prediction (MCP). However, in the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensated prediction is applied. As Figure 11 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: where (mv 0x , mv 0y ) is the motion vector of the upper left control point, (mv 1x , mv 1y ) is the motion vector of the upper right control point, and (mv 2x , mv 2y ) is the motion vector of the lower left control point. To simplify motion compensated prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4×4 luma sub-block, the motion vector of the central sample of each sub-block is calculated according to the above equations (as Figure 12 shown), and rounded to 1 / 16 fractional precision. Then a motion compensated interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of a 4×4 chrominance sub-block is calculated as the average of the MVs of the upper left luma sub-block and the lower right luma sub-block in the co-located 8×8 luma region. Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.1.3.1. Affine Merge Prediction The AF_MERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates are used to form the affine Merge candidate list: - Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; - Constructed affine Merge candidate CPMVP derived using the translational MVs of neighboring CUs; - Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 13 shown. For the prediction value on the left side, the scanning order is A0 -> A1, and for the prediction value on the upper side, the scanning order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When the neighboring affine CU is identified, its control point motion vectors are used to derive the CPMV candidates in the affine Merge list of the current CU. As Figure 30 shown, if the neighboring lower left block A is encoded / decoded using the affine mode, the motion vectors v 2 , v 3 and v 4 of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. When block A is encoded / decoded using the 4-parameter affine model, two CPMVs of the current CU are calculated according to v 2 and v 3 . In the case where block A is encoded / decoded using the 6-parameter affine model, three CPMVs of the current CU are calculated according to v 2 , v 3 and v 4 . Figure 14 The control point motion vector inheritance is shown. The constructed affine candidates refer to constructing candidates by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the Figure 15 specified spatial neighbors and temporal neighbors shown, and CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , the B2 -> B3 -> A2 block is checked, and the MV of the first available block is used. For CPMV 2 , the B1 -> B0 block is checked, and for CPMV 3 , the A1 -> A0 block is checked. TMVP is used as CPMV 4 (if available). After obtaining the MVs of the four control points, affine Merge candidates are constructed based on those motion information. The following combinations of control point MVs are used to construct in sequence: {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1, CPMV 3 , CPMV 4}, {CPMV 2 , CPMV 3 , CPMV 4}, {CPMV 1 , CPMV 2}, {CPMV 1 , CPMV 3}。 Combinations of 3 CPMVs construct 6 - parameter affine Merge candidates, and combinations of 2 CPMVs construct 4 - parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded. After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list. 2.1.3.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with both width and height greater than or equal to 16. In the bitstream, a CU - level affine flag is signaled to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4 - parameter affine or 6 - parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The size of the affine AVMP candidate list is 2, and it is generated by sequentially using the following four types of CPVM candidates: - Inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs; - Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs; - Translational MVs from neighboring CUs. - Zero MVs. The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 15 . The same checking order as in the construction of affine Merge candidates is used. Additionally, the reference picture indices of neighboring blocks are checked. The first block in the checking order that is inter - frame coded and has the same reference picture as the current CU is used. There is only one. When the current CU is coded using the 4 - parameter affine mode, and mv 0 and mv 1When all are available, they are added as a candidate in the affine AMVP list. When the current CU is coded / decoded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable. If the affine AMVP list candidates are still less than 2 after inserting valid inherited affine AMVP candidates and the constructed AMVP candidates, then mv 0 , mv 1 and mv 2 will be added in sequence as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.1.3.3. Affine Motion Information Storage In VVC, the CPMVs of affine CUs are stored in a separate cache. The stored CPMVs are only used to generate the inherited CPMVs in the affine Merge mode and the inherited CPMVs in the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation of the Merge / AMVP list of translational MVs, and deblocking. To avoid picture line caches for additional CPMVs, the inheritance from the affine motion data of the CU above the CTU is processed differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the left-bottom sub-block MV and the right-bottom sub-block MV in the line cache are used for affine MVP derivation instead of the CPMV. In this way, the CPMV is only stored in the local cache. If the candidate CU is 6-parameter affine coded / decoded, the affine model is degraded to a 4-parameter model. As Figure 16 shown, along the top boundary of the CTU, the left-bottom sub-block motion vector and the right-bottom sub-block motion vector of the CU are used for affine inheritance of the CU in the bottom of the CTU. 2.1.3.4. Prediction Refinement using Optical Flow (PROF) for Affine Mode Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equation. PROF is described as the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate sub-block prediction I(i, j). Step 2) Using a 3-tap filter [-1, 0, 1], the spatial gradient g of sub-block prediction x (i, j) and g y (i, j) are calculated at each sample position. The gradient calculation is exactly the same as the gradient calculation in BDOF. g x (i, j) = (I(i + 1, j) >> shift1) - (I(i - 1, j) << shift1) (2-11) g v (i, j) = (I(i, j + 1) << shift1) - (I(i, j - 1) << shift1) (2-12) shift1 is used to control the accuracy of the gradient. The sub-block (i.e., 4×4) prediction extends one sample on each side of the gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Step 3) Luminance prediction refinement is calculated through the following optical flow equation. ΔI(i, j) = g x (i, j) * Δv x (i, j) + g y (i, j) * Δv y (i, j) (2-13) where as Figure 17 shown, Δv(i, j) is the sample MV calculated for the sample position (i, j), denoted as v(i, j), and the difference from the sub-block MV of the sub-block to which the sample (i, j) belongs. Δv(i, j) is quantized in units of 1 / 32 luminance sample accuracy. Since the affine model parameters and the sample position relative to the sub-block center do not change from sub-block to sub-block, Δv(i, j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i, j) and dy(i, j) be the horizontal and vertical offsets of the sample position (i, j) to the sub-block center (x SB , y SB ), and Δv(x, y) can be derived through the following equation. To maintain accuracy, the input of the sub-block (x SB , y SB ) is calculated as ((W SB - 1) / 2, (H SB - 1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For the 4-parameter affine model, For the 6-parameter affine model, where (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the motion vectors of the control points in the upper left, upper right, and lower left, and w and h are the width and height of the CU. Step 4) Finally, the luminance prediction refinement ΔI(i, j) is added to the sub-block prediction I(i, j). The final prediction I’ is generated by the following equation. I′(i, j) = I(i, j) + ΔI(i, j) PROF is not applicable to affine-coded CUs in two cases: 1) all control point MVs are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limit because the sub-block-based affine MC is degraded to CU-based MC to avoid large memory access bandwidth requirements. A fast coding method is applied to reduce the coding complexity of affine motion estimation using PROF. In the following two cases, PROF is not applied to the affine motion estimation stage: a) if the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low-latency picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated. 2.1.4. Sub-block-based Temporal Motion Vector Prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and Merge mode of the CU in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: - TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; - While TMVP prefetches the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the central block relative to the current CU), SbTMVP applies a motion displacement before prefetching the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in Figure 18. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 18(a) is examined. If A1 has a motion vector using the co-located picture as its reference picture, then that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located picture as shown in Figure 18(b). The example in Figure 18(b) assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) in the co-located picture is used to derive the motion information of the sub-CU. After the motion information of the co-located sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a similar way to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used to signal the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8×8, and like the affine Merge mode, the SbTMVP mode only applies to CUs with both width and height greater than or equal to 8. The encoding and decoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.1.5. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in quarter-luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded and decoded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows: - Normal AMVP mode: quarter-luminance samples, half-luminance samples, integer-luminance samples, or quarter-luminance samples. - Affine AMVP mode: quarter-luminance samples, integer-luminance samples, or 1 / 16-luminance samples. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both the horizontal and vertical MVDs of reference list L0 and reference list L1) are zero, the quarter-luminance sample MVD resolution is assumed. For a CU with at least one non-zero MVD component, the first flag is signaled to indicate whether quarter-luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luminance sample MVD precision is used for the current CU. Otherwise, the second flag is signaled to indicate whether half-luminance sample or other MVD precision (integer or quarter-luminance samples) is used for normal AMVP CUs. In the case of half-luminance samples, the half-luminance sample position uses a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, the third flag is signaled to indicate whether integer-luminance sample or quarter-luminance sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer-luminance sample MVD precision or 1 / 16-luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luminance samples, half-luminance samples, integer-luminance samples, or quarter-luminance samples), the motion vector prediction value of the CU is rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., a negative motion vector prediction value is rounded to positive infinity, and a positive motion vector prediction value is rounded to negative infinity). The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM13, the RD check for MVD accuracy outside the quarter-luminance samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for the MVD accuracy of the quarter-luminance samples and the RD cost for the MV accuracy of the full-luminance samples are calculated. Then, the RD cost for the MVD accuracy of the full-luminance samples is compared with the RD cost for the MVD accuracy of the quarter-luminance samples to decide whether it is necessary to further check the RD cost for the MVD accuracy of the four-luminance samples. When the RD cost for the MVD accuracy of the quarter-luminance samples is much smaller than the RD cost for the MVD accuracy of the full-luminance samples, the RD check for the MVD accuracy of the four-luminance samples is skipped. Then, if the RD cost for the MVD accuracy of the full-luminance samples is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for the MVD accuracy of the half-luminance samples is skipped. For the affine AMVP mode, if the rate-distortion cost of the affine inter mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, the Merge / skip mode, the normal AMVP mode with the MVD accuracy of the quarter-luminance samples, and the affine AMVP mode with the MVD accuracy of the quarter-luminance samples, the MV accuracy of the 1 / 16-luminance samples and the affine inter mode with the 1-pixel MV accuracy are not checked. Additionally, in the affine inter modes with the 1 / 16-luminance samples and the quarter-luminance samples MV accuracy, the affine parameters obtained in the affine inter mode with the quarter-luminance samples MV accuracy are used as the starting search points. 2.1.6. Bi-directional prediction with CU-level weights (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 18) Five weights are allowed in weighted-average bi-directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luminance samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights are used (w ∈ {3, 4, 5}). - At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-delay picture, unequal weights for 1-pixel and 4-pixel motion vector precisions are only conditionally checked. - When combined with affine, affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode. - When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked. - Unequal weights are not searched when certain conditions are met, depending on the POC distance between the current picture and its reference picture, the coding / decoding QP, and the temporal level. The BCW weight index is decoded using a context-coded binary bit followed by a bypass-coded binary bit. The first context-coded binary bit indicates whether equal weights are used; and if unequal weights are used, the bypass-coded additional binary bits are signaled to indicate which unequal weight is used. Weighted prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding / decoding of video content in fading situations. Support for WP is also added in the VVC standard. WP allows signaling of weighted parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded / decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.7. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. The BDOF, previously known as BIO, was included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version, which requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: - The CU is encoded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order. - The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same. - Both of the two reference pictures are short-term reference pictures. - The CU is not encoded and decoded using the affine mode or the SbTMVP Merge mode. - The CU has more than 64 luma samples. - Both the CU height and the CU width are greater than or equal to 8 luma samples. - The BCW weight index indicates equal weights. - WP is not enabled for the current CU. - The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. Then the motion refinement is used to adjust the bidirectional prediction sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the differences between two neighboring samples, and k = 0, 1, i.e., where I (k) (i, j) is the sample value at the coordinate (i, j) of the prediction signal in list k (k = 0, 1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the gradients S 1 , S 2 , S 3 , S 5 and S6 The autocorrelation and cross-correlation are calculated as follows: where where Ω is a 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, the motion refinement (v x , v y ) is derived using the cross-correlation term and the autocorrelation term with the following equation: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on the motion refinement and gradient, the following adjustment is calculated for each sample point in the 4×4 sub-block: Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points as follows: pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o offset ) >> shift (2 - 24) These values are selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some predicted sample points I (k) (i, j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 9 depicted, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating the predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it is divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are considered CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1)), where W indicates the sub-block width and H indicates the sub-block height. To avoid additional complexity in SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated during the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, the BDOF is also disabled. When the CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.8. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the Merge mode MV, decoder-side motion vector refinement based on bilateral matching is applied in VVC. During the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. As Figure 20 shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs encoded or decoded using the following modes and features: - CU-level Merge mode with bidirectional prediction MV. - For the current picture, one reference picture is past and the other reference picture is future. - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same. - Both reference pictures are short-term reference pictures. - The CU has more than 64 luma samples. - Both the CU height and the CU width are greater than or equal to 8 luma samples. - The BCW weight index indicates equal weights. - WP is not enabled for the current block. - The CIIP mode is not used for the current block. The refined MVs derived through the DMVR process are used to generate inter-prediction samples and are also used for temporal motion vector prediction in future picture coding. The original MVs are used for the deblocking process and are also used for spatial motion vector prediction in future CU coding. The additional functions of DMVR are mentioned in the following sub-articles. 2.1.8.1. Search Scheme In DVMR, the search points are centered around the initial MV, and the MV offsets follow the MV difference mirroring rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2-25) MV1′ = MV1 - MV_offset (2-26) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luminance samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample phase of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks referred to by the initial MV candidates reduces the SAD value by 1 / 4. The integer sample search is followed by fractional sample refinement. To save computational complexity, the fractional sample refinement is derived using the parametric error surface equation instead of performing an additional search using SAD comparison. The fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. When the integer sample search phase ends at the center with the minimum SAD in the first or second iteration search, the fractional sample refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x, y) = A(x - x min ) 2 + B(y - y min ) 2 + C (2-27) where (x min , y min) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (2-28) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) (2-29) x min and y min values are automatically restricted between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer-distance refined MV to obtain a sub-pixel accurate refined delta MV. 2.1.8.2. Bilinear Interpolation and Sample Padding In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate the samples at the fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples during the DMVR search process. Another important effect is that by using the bilinear filter, within a 2-sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples of the normal MC process, samples will be padded from those available samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV. 2.1.8.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16×16. 2.1.9. Combined Inter - and Intra - Prediction (CIIP) In VVC, when a CU is encoded or decoded using the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal and the intra prediction signal. The inter prediction signal P in the CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived after the regular intra prediction process with the planar mode. Then, a weighted average is used to combine the intra prediction signal and the inter prediction signal, where the weight value depends on the coding modes of the top neighboring block and the left neighboring block (depicted in Figure 21 ) and is calculated as follows: - If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; - If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; - If (isIntraLeft + isIntraTop) equals 2, set wt to 3; - Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2; - Otherwise, set wt to 1. The CIIP prediction is formed as follows: P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (2 - 30) 2.1.10. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signaled using a CU-level flag as a type of Merge mode, where other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. Among a total of 64 partitions, for each possible CU size w × h = 2 m × 2 n , m, n ∈ {3...6}, partitioning is supported by the geometric partitioning mode. When using this mode, the CU is divided into two parts by a geometrically positioned line ( Figure 22)。The position of the dividing line is derived mathematically from the angular parameter and the offset parameter of a specific partition. Each part of the geometric partition in the CU is inter-frame predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure that, as in the case of regular bidirectional prediction, only two motion-compensated predictions are required for each CU. If the current CU uses the geometric partition mode, the geometric partition index indicating the partition mode of the geometric partition (angle and offset) and two Merge indices (one Merge index for each partition) are further signaled. The number of candidates of the maximum GPM size is explicitly signaled in the SPS and the syntax binarization of the GPM Merge index is specified. After predicting each part in the geometric partition, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partition edge. This is the prediction signal for the entire CU, and the transform and quantization processes will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored. 2.1.10.1. Unidirectional Prediction Candidate List Construction The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X is equal to the parity of n) is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 23 . In the case where there is no corresponding LX motion vector of the nth extended Merge candidate, the L(1 - X) motion vector of the same candidate is used instead of the unidirectional prediction motion vector for the geometric partition mode. 2.1.10.2. Hybrid Along the Geometric Partition Edge After predicting each part of the geometric partition using its own motion, a hybrid is applied to the two prediction signals to derive the samples around the geometric partition edge. The hybrid weight for each position of the CU is derived based on the distance between the respective position and the partition edge. The distance of the position (x, y) to the partition edge is derived as: where i, j are the indices of the angle and offset of the geometric partition, which depend on the signaled geometric partition index. The symbols ρ x,j and ρ y,j depend on the angle index. The weight of each part of the geometric partition is derived as follows: wIdxL(x, y) = partIdx? 32 + d(x, y) : 32 - d(x, y) w 1 (x, y) = 1 - w 0 (x, y) partIdx depends on the angular index i. The weight w 0 An example of Figure 24 is shown in 2.1.10.3. Motion Vector Field Storage for Geometric Partitioning Patterns Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combination Mv of Mv1 and Mv2 are stored in the motion vector field of the CU encoded / decoded in the geometric partitioning pattern. The type of motion vector stored for each individual position in the motion vector field is determined as follows: sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 - partIdx) : partIdx) where motionIdx is equal to d(4x + 2, 4y + 2). partIdx depends on the angular index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion vector field, otherwise if sType is equal to 2, then the combination Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then simply combine Mv1 and Mv2 to form a bi - directional predicted motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, then only the unidirectional predicted motion Mv2 is stored. 2.1.11. Local Illumination Compensation (LIC) LIC is an inter - frame prediction technique that models the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template and the reference block template. The parameters of the function can be represented by a scale ɑ and an offset β, which form a linear equation, i.e., ɑ * p[x] + β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. Since ɑ and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for the AMVP mode to indicate the use of LIC. Local illumination compensation is used for unidirectional predicted inter - frame CUs with the following modifications. · Intra - frame neighboring samples can be used for LIC parameter derivation; · LIC is disabled for blocks with fewer than 32 luma samples; ·For both non-sub-block and affine modes, LIC parameter derivation is performed based on the modulo-block samples corresponding to the current CU rather than the partial modulo-block samples corresponding to the first top-left 16×16 unit. ·Samples of the reference block template are generated by using MC with block MVs without rounding them to integer pixel precision. 2.1.12. Non-adjacent spatial candidates Insert non-adjacent spatial Merge candidates after TMVP in the regular Merge candidate list. The pattern of the spatial Merge candidates is shown in Figure 25 . The distance between the non-adjacent spatial candidates and the current coding block is based on the width and height of the current coding block. The line buffer limit is not applied. 2.1.13. Template matching I Template matching I (TM) is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top neighboring block and / or the left neighboring block of the current CU) and a block in the reference picture (i.e., the same size as the template). As shown in Figure 26 , within the [-8, +8] pixel search range, search for a better MV around the initial motion of the current CU. The template matching method is used for the following modifications: the search step size is determined based on the AMVR mode, and TM can be cascaded with the bilateral matching process in the Merge mode. In the AMVP mode, the MVP candidates are determined based on the template matching error to select the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM is performed only for this specific MVP candidate for MV refinement. TM refines this MVP candidate by starting from the full pixel MVD precision (or 4 pixels for the 4-pixel AMVR mode) within the [-8, +8] pixel search range using iterative diamond search. The AMVP candidates can be further refined by using a cross-shaped search with full pixel MVD precision (or 4 pixels for the 4-pixel AMVR mode), and then followed by sequential half-pixel and quarter-pixel searches depending on the AMVR mode specified in Table 3. This search process ensures that the MVP candidates still maintain the same MV precision as the MV precision indicated by the AMVR mode after the TM process. Table 3. Search patterns of AMVR and Merge modes with AMVR. In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 3, depending on the merged motion information and the alternative interpolation filter (the interpolation filter used when AMVR is in the half-pixel mode), the TM can be performed in all ways up to 1 / 8 pixel MVD accuracy or skipped for accuracies beyond the half-pixel MVD accuracy. In addition, when the TM mode is enabled, template matching can be used as an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether the BM can be enabled according to its enable condition check. 2.1.14. Multi-pass decoder-side motion vector refinement (mpDMVR) Multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coded / decoded blocks. In the second pass, BM is applied to each 16×16 sub-block within the coded / decoded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction. 2.1.14.1. First pass - Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), in the bidirectional prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer sample accuracy intDeltaMV. The local search applies a 3×3 square search pattern to loop within the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. Further apply the existing fractional sample refinement to derive the final deltaMV. Then, the refined MV after the first pass is derived as: · MV0_pass1 = MV0 + deltaMV; · MV1_pass1 = MV1 - deltaMV. 2.1.14.2. Second Pass - Sub - block - based Bilateral Matching MV Refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub - blocks. For each sub - block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub - blocks in L0 and L1. For each sub - block, BM performs a full search to derive the integer - sample accuracy intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub - blocks, as: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond - shaped search areas, as Figure 27 shown. Each search area is assigned a cost factor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster - scan order from the upper - left corner of the area to the lower - right corner. When the minimum bilCost within the current search area is less than a threshold equal to sbW * sbH, the integer - pixel full search is terminated; otherwise, the integer - pixel full search continues to the next search area until all search points are checked. Further apply the existing VVC DMVR fractional sample refinement to derive the final deltaMV(sbIdx2). Then, the refined MV in the second pass is derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2); ·MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2). 2.1.14.3. Third Pass - Sub - block - based Bidirectional Optical Flow MV Refinement In the third pass, refined MVs are derived by applying BDOF to 8×8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MVs of the parent - child blocks in the second pass, BDOF refinement is applied to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between - 32 and 32. The refined MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows: ·MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2)+bioMv; ·MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2)-bioMv. 2.1.15. OBMC When applying OBMC, the motion information of neighboring blocks with weighted prediction is used to refine the top - boundary pixels and left - boundary pixels of the CU. The conditions for not applying OBMC are as follows: ·When OBMC is disabled at the SPS level. ·When the current block has an intra mode or IBC mode. ·When LIC is applied to the current block. ·When the current luma block area is less than or equal to 32. Sub - block boundary OBMC is performed by applying the same blend to the top - sub - block boundary pixels, left - sub - block boundary pixels, bottom - sub - block boundary pixels, and right - sub - block boundary pixels using neighboring sub - blocks. It enables sub - block - based coding / decoding tools: ·Affine AMVP mode; ·Affine Merge mode and sub - block - based temporal motion vector prediction (SbTMVP); ·Sub - block - based bilateral matching. 2.1.16. Sample - based BDOF In sample - based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The coding block is divided into 8×8 sub - blocks. For each sub - block, it is determined whether to apply BDOF by checking the SAD between two reference sub - blocks against a threshold. If it is decided to apply BDOF to the sub - block, for each sample point in the sub - block, a sliding 5×5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample point of the window. 2.1.17. Interpolation The 8 - tap interpolation filter used in VVC is replaced by a 12 - tap filter. The interpolation filter is derived from a sine function, where the frequency response is truncated at the Nyquist frequency and tapered by a cosine window function. Table 4 gives the filter coefficients for all 16 phases. Figure 28 The frequency response of the interpolation filter is compared with the VVC interpolation filter, all at the half - pixel phases. Table 4. Filter coefficients of the 12 - tap interpolation filter 2.1.18. Multiple Hypothesis Prediction (MHP) In the multiple hypothesis inter - prediction mode, in addition to the regular bi - prediction signal, one or more additional motion - compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted sample - by - sample superposition. Using the bi - prediction signal p bi and the first additional inter - prediction signal / hypothesis h 3 , the resulting prediction signal p 3 is obtained as follows. p 3 =(1 - ɑ)P bi +ɑh 3 According to the following mapping, the weighting factor ɑ is specified by the new syntax element add_hyp_weight_idx. add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 Similarly, more than one additional prediction signal can be used. The resulting overall prediction signal is cumulatively obtained iteratively using each additional prediction signal. p n+1 =(1 - ɑ n+1 )p n +ɑ n+1 h n+1 The resulting overall prediction signal is obtained as the last p n (i.e., with the largest index). Within this EE, up to two additional prediction signals can be used (i.e., n is limited to 2). The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference or implicitly by specifying a Merge index. A separate multi-hypothesis Merge flag differentiates between these two signaling modes. For the inter-frame AMVP mode, if unequal weights in the BCW are selected in the bi-prediction mode, only MHP is applied. The combination of MHP and BDOF is possible. However, BDOF is only applied to the bi-prediction signal part of the prediction signal (i.e., the normal first two hypotheses). 2.1.19. Adaptive Reordering of Merge Candidates Using Template Matching (ARMC-TM) Merge candidates are adaptively reordered using template matching (TM). The reordering method is applied to the regular Merge mode, the template matching (TM) Merge mode, and the affine Merge mode (excluding SbTMVP candidates). For the TM Merge mode, the Merge candidates are reordered before the refinement process. After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5 for the regular Merge mode and the TM Merge mode. The subgroup size is set to 3 for the affine Merge mode. The Merge candidates in each subgroup are reordered in ascending order according to the template matching-based cost values. For simplicity, the Merge candidates in the last rather than the first subgroup are not reordered. The template matching cost of a Merge candidate is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate. When a Merge candidate uses bi-prediction, the reference samples of the template of the Merge candidate are also generated by bi-prediction as shown in Figure 29 below. For sub-block-based Merge candidates with a sub-block size equal to Wsub×Hsub, the above template includes several sub-templates of size Wsub×1, and the left template includes several sub-templates of size 1×Hsub. As shown in Figure 30 below, the motion information of the sub-blocks in the first row and the first column of the current block is used to derive the reference samples of each sub-template. 2.1.20. Geometric Partitioning Mode (GPM) with Merged Motion Vector Difference (MMVD) The GPM in VVC is extended by applying motion vector refinement to the top of the existing GPM unidirectional MVs. First, the flag of the GPM CU is signaled to specify whether this mode is used. If the mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD. If the MVD is signaled for a geometric partition, after selecting the GPMMerge candidate, the motion of the partition is further refined by the signaled MVD information. All other processes remain the same as GPM. The MVD is signaled as a pair of distance and direction, similar to in MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) involved in GPM-MMVD in GPM. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD is left-shifted by 2 in MMVD. 2.1.21. Geometric Partitioning Mode (GPM) Using Template Matching (TM) Template matching is applied to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to the two geometric partitions. TM is used to refine the motion information of each geometric partition. When TM is selected, depending on the partition angle, left, above, or both left and above neighboring samples are used to construct the template, as shown in Table 5. Then the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern as the Merge mode with the half-pixel interpolation filter disabled. Table 5 Templates for the first geometric partition and the second geometric partition, where A indicates using the above sample, L indicates using the left sample, and L+A indicates using both the left sample and the above sample. Segmentation Angle 0 2 3 4 5 8 11 12 13 14 First Segmentation A A A A L+A L+A L+A L+A A A Second Segmentation L+A L+A L+A L L L L L+A L+A L+A Segmentation Angle 16 18 19 20 21 24 27 28 29 30 First Segmentation A A A A L+A L+A L+A L+A A A Second Segmentation L+A L+A L+A L L L L L+A L+A L+A The GPM candidate list is constructed as follows: 1. The interleaved list 0 MV candidates and list 1 MV candidates are directly derived from the regular Merge candidate list, where the list 0 MV candidates have a higher priority than the list 1 MV candidates. A deduplication method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. 2. The interleaved list 1 MV candidates and list 0 MV candidates are further directly derived from the regular Merge candidate list, where the list 1 MV candidates have a higher priority than the list 0 MV candidates. The same deduplication method with an adaptive threshold is also applied to remove redundant MV candidates. 3. Zero MV candidates are filled until the GPM candidate list is full. GPM-MMVD and GPM-TM are specifically enabled for one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is presumed to be false. 2.1.22. GPM with Inter-Frame and Intra-Frame Prediction (GPM inter-intra) With GPM inter-intra, in addition to the Merge candidates for each non-rectangular partition region in the CU to which GPM is applied, a predetermined intra-frame prediction mode for the geometric split line can also be selected. In the proposed method, an intra-frame prediction mode or an inter-frame prediction mode is determined for each GPM separable region with a flag from the encoder. When in the inter-frame prediction mode, a unidirectional prediction signal is generated from the MVs in the Merge candidate list. On the other hand, when in the intra-frame prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra-frame prediction mode specified by the index from the encoder. The variation of possible intra-frame prediction modes is restricted by the geometry. Finally, the two unidirectional prediction signals are blended in the same way as in ordinary GPM. 2.1.23. Adaptive Decoder-Side Motion Vector Refinement (Adaptive DMVR) The Adaptive Decoder-Side Motion Vector Refinement method consists of two new Merge modes that are introduced to refine the MV only in one direction (L0 or L1) of the bi-directional prediction of the Merge candidates that satisfy the DMVR condition. The selected Merge candidates are applied with a multi-pass DMVR process to refine the motion vector. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is zero. Similar to the regular Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially neighboring coded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list is used by the two proposed Merge modes, and the Merge index is coded / decoded in the regular Merge mode. 2.1.24. Bilateral Matching AMVP-MERGE Mode (AMVP-MERGE) In the AMVP Merge mode, the bi-directional prediction value consists of the AMVP prediction value in one direction and the Merge prediction value in the other direction. The AMVP part of the proposed mode is signaled as regular unidirectional AMVP, i.e., the reference index and MVD are signaled, and it has a derived MVP index (TM_AMVP) if template matching is used, or the MVP index is signaled when template matching is disabled. The Merge index is not signaled, and the Merge prediction value is selected from the candidate list with the minimum template or bilateral matching cost. When the selected Merge prediction value and AMVP prediction value meet the DMVR condition (which is at least one reference picture from the past and one reference picture from the future relative to the current picture) and the distance from the two reference pictures to the current picture is the same, bilateral matching MV refinement is applied to the Merge MV candidate and AMVP MVP as a starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the Merge prediction value or AMVP prediction value with a higher template matching cost. The third pass of 8×8 sub-PU BDOF refinement as multi-pass DMVR is enabled as AMVP Merge mode codec block. 2.1.25.IBC Merge / AMVP List Construction IBC Merge / AMVP list construction is modified as follows: An IBC Merge / AMVP candidate can be inserted into the IBC Merge / AMVP candidate list only if it is valid. The upper right spatial domain candidate, the lower left spatial domain candidate, the upper left spatial domain candidate and a pairwise average candidate may be added to the IBC Merge / AMVP candidate list. Adaptive Reordering Based on Template (ARMC-TM) is applied to the IBC Merge list. The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC Merge candidates with full deduplication, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list. Candidates for zero vectors used to populate the IBC Merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in IBC Merge mode, and therefore, are discarded as BVPs in the IBC candidate list. Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Figure 31 Depicted in. 2.1.26. IBC Using Template Matching Template matching is used in both the IBC Merge mode and the IBC AMVP mode in IBC. The IBC-TM Merge list is modified compared to the list used by the regular IBC Merge mode such that candidates are selected according to a deduplication method using the motion distance between candidates as in the regular TM Merge mode. End-zero motion fulfillment is replaced by motion vectors of the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU. In the IBC-TM Merge mode, the selected candidates are refined using a template matching method before the RDO or decoding process. The IBC-TM Merge mode has competed with the regular IBC Merge mode and is signaled by the TM-Merge flag. In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM Merge list. The template matching method is used to refine each of these 3 selected candidates and they are sorted according to their resulting template matching cost. Then usually only the top two candidates are considered during the motion estimation process. The template matching refinement for both the IBC-TM Merge mode and the AMVP mode is very simple because the IBC motion vectors are constrained (i) to be integers and (ii) within the reference region as shown in Figure 32 Therefore, in the IBC-TM Merge mode, all refinements are performed with integer precision and in the IBC-TM AMVP mode, it is performed with integer or 4-pixel precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the templates used in each refinement step must comply with the constraints of the reference region. 2.1.27. IBC Reference Region The reference region of IBC extends to two CTU rows above. Figure 33Shows the reference region for encoding and decoding CTU(m, n). Specifically, for the CTU(m, n) to be encoded and decoded, the reference region includes CTUs with indices (m - 2, n - 2)...(W, n - 2), (0, n - 1), (W, n - 1), (0, n)...(m, n), where W represents the maximum horizontal index within the current slice, strip, or picture. This setting ensures that for a CTU size of 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is horizontally limited to [-(C << 1), C >> 2] and vertically limited to [-C, C >> 2] to accommodate the reference region expansion, where C represents the CTU size. 2.1.28. MVD Symbol Prediction In this method, possible MVD symbol combinations are sorted according to the template matching cost, and the index corresponding to the true MVD symbol is derived and context decoded. On the decoder side, the MVD symbol is derived as follows: 1. Parse the size of the MVD component; 2. Parse the context decoded MVD symbol prediction index; 3. Construct MV candidates by creating combinations between possible symbols and absolute MVD values and adding them to the MV prediction value; 4. Derive the MVD symbol prediction cost for each derived MV based on the template matching cost and sorting; 5. Use the MVD symbol prediction index to pick the true MVD symbol. MVD symbol prediction is applied to the inter-frame AMVP mode, affine AMVP mode, MMVD mode, and affine MMVD mode. 2.1.29. Enhanced Bidirectional Motion Compensation In bidirectional motion compensation, out-of-bounds (OOB) predicted samples are discarded, and only non-OOB predicted values are used to generate the final predicted value. Specifically, assume Pos_x i,j and Pos_y i,j represent the position of a predicted sample in a current block, and represent the MV of the current block; Pos LeftBdry 、Pos RightBdry 、Pos TopBdry and Pos BottomBdry are the positions of the four boundaries of the picture. A predicted sample is considered OOB when at least one of the following conditions is met: where half_pixel is equal to 8, which represents the half-pixel sample distance in 1 / 16 pixel sample precision. After checking the OOB condition for each sample point, a final predicted sample point of a bi - directional block is generated as follows: If is OOB and is non - OOB Otherwise, if is non - OOB and is OOB. Otherwise When BCW is enabled, the OOB checking process also applies. 2.1.30. Block - level reference picture list re - ordering Use the block - level reference picture re - ordering method based on template matching. For the unidirectional prediction AMVP mode, the reference pictures in list 0 and list 1 are interleaved to generate a combined list. For each hypothesis of the reference pictures in the combined list, a match is performed to calculate the cost. The combined list is re - ordered in ascending order of the template - matching cost. The index of the selected reference picture in the re - ordered combined list is signaled in the bitstream. For the bi - directional prediction AMVP mode, a list of reference picture pairs from list 0 and list 1 is generated and similarly re - ordered based on the template - matching cost. The index of the selected pair is signaled. 2.1.31. Affine model inheritance based on historical parameters and non - adjacent affine modes Affine model inheritance based on historical parameters (HAMI) allows an affine model to inherit from a previously affine - encoded block that may not be adjacent to the current block. Similar to the enhanced regular Merge mode, non - adjacent affine mode (NA - AFF) is introduced. Establish a first historical parameter table (HPT). The entries of the first HPT store sets of affine parameters: a, b, c, and d, each of which is represented by a 16 - bit signed integer. The entries in the HPT are classified by reference list and reference index. Five reference indices are supported for each reference list in the HPT. In a formulaic way, the category of the HPT (denoted as HPTCat) is calculated as HPTCat(RefList, RefIdx)=5×RefList + min(RefIdx, 4), where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index respectively. For each category, up to seven entries can be stored, resulting in a total of 70 entries in the HPT. At the start of each CTU row, the number of entries for each category is initialized to zero. When decoding a block with reference list RefList curand RefIdx cur After the CU of affine encoding and decoding, the affine parameters are utilized to update the entries in the category HPTCat(RefList cur , RefIdx cur ) in a manner similar to the update of the HMVP table. Candidate based on historical affine parameters (HAPC) is derived from one of the seven neighboring 4×4 blocks represented as A0, Al, A2, B0, B1, B2, or B3 in Figure 4 and the set of affine parameters in the corresponding entry stored in the first HPT. The MV of the neighboring 4×4 block serves as the base MV. In a formulated way, the MV of the current block at position (x, y) is calculated as: where (mv h base , mv v base ) represents the MV of the neighboring 4×4 block, and (x base , y base ) represents the center position of the neighboring 4×4 block. (x, y) can be the upper left, upper right, and lower left corners of the current block to obtain the corner position MV (CPMV) of the current block, or it can be the center of the current block to obtain the regular MV of the current block. A second historical parameter table (HPT) with base MV information is also appended. There are nine entries in the second HPT, where the entries include the base MV, the reference index for each reference list, four affine parameters, and the base position. Additional Merge HAPC can be generated from the second HPT with base MV information, and the corresponding affine model is stored in the entry. The difference between the first HPT and the second HPT is shown in Figure 34 . In addition, paired affine Merge candidates are generated from two affine Merge candidates, which are either historically derived or non-historically derived. The paired affine Merge candidates are generated by averaging the CPMVs of the existing affine Merge candidates in the list. In response to the introduction of new HAPC, the size of the Merge candidate list based on sub-blocks is increased from 5 to 15, all of which are involved in the ARMC

[10] process. In NA-AFF, the pattern of obtaining non-adjacent spatial domain neighbors is shown in Figure 6 . Similar to the existing non-adjacent regular Merge candidates [8], the distance between the non-adjacent spatial domain neighbors and the current coding block in NA-AFF is also defined based on the width and height of the current CU. Utilize Figure 6Motion information of non - adjacent spatial neighbors in it is used to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for inherited candidates, except that CPMV is inherited from non - adjacent spatial neighbors, the same derivation process of inherited affine Merge / AMVP candidates in VVC remains unchanged. Non - adjacent spatial neighbors are checked based on their distance from the current block (i.e., from near to far). At a specific distance, only the first available neighbors (coded / decoded using affine mode) from each side of the current block (e.g., left and above) are included for inherited candidate derivation. Figure 35 shows the spatial neighbors for deriving affine Merge / AMVP candidates. In addition, Figure 35 sub - picture (a) of shows the spatial neighbors for deriving inherited candidates, and Figure 35 sub - picture (b) of shows the spatial neighbors for deriving the first type of constructed candidates. As Figure 35 indicated by the dashed arrows in sub - picture (a) of, the checking order of the left and above neighbors is from bottom to top and from right to left, respectively. For the first type of constructed candidates, as shown in Figure 35 sub - picture (b) of, the positions of a left and an above non - adjacent spatial neighbor are first determined independently; afterwards, the position of the upper - left neighbor can be determined accordingly, which can enclose a rectangular virtual block with the left and above non - adjacent neighbors. Then, as shown in Figure 36 , the motion information of the three non - adjacent neighbors is used to form CPMVs at the upper - left (A), upper - right (B), and lower - left (C) of the virtual block, and finally project them to the current CU to generate the corresponding constructed candidates. NA - AFF candidates are inserted into the existing affine Merge candidate list and affine AMVP candidate list according to the following order: Affine Merge Mode: 1. SbTMVP candidates, if available. 2. Inherited from adjacent neighbors. 3. Inherited from non - adjacent neighbors. 4. Constructed from adjacent neighbors. 5. The first type of constructed affine candidates from non - adjacent neighbors. 6. Zero MV. Affine AMVP Mode: 1. Inherited from adjacent neighbors. 2. Constructed from adjacent neighbors. 3. Translational MV from adjacent neighbors. 4. Translational MV from temporal neighbors. 5. Inherited from non - adjacent neighbors. 6. First type of constructing affine candidates from non - adjacent neighbors. 7. Zero MV. Due to including additional candidates generated by NA - AFF, the size of the affine Merge candidate list increases from 5 to 15. The subgroup size of ARMC for the affine Merge mode increases from 3 to 15. In NA - AFF: 1. Regions from non - adjacent neighbors are restricted within the current CTU (i.e., there is no additional storage requirement for line buffers). 2. The storage granularity of affine motion information including CPMV and reference indices is reduced from 8×8 to 16×16 (i.e., only the affine motion from the top - left 8×8 block is saved). Additionally, the saved CPMV is projected onto each 16×16 block before storage, such that position and size information is not required. 3. Only the top - left CPMV and top - right CPMV are stored (i.e., a 4 - parameter affine model for NA - AFF is always used). 2.1.32. Regression - based affine candidate derivation method A regression - based affine candidate derivation method is proposed. The sub - block motion fields from previously decoded affine CUs and the motion vectors from adjacent sub - blocks of the current CU are used as inputs to the regression process. The predicted CPMV instead of the sub - block motion field of the current block is derived as the output. The derived CPMV can be added to the sub - block Merge candidate list or the affine AMVP list. The scan pattern of the previously decoded affine CUs is the same as the non - adjacent scan pattern used in the construction of the regular Merge candidate list. 2.2. Transform and coefficient coding 2.2.1. Large - block - size transform with high - frequency zeroing In VVC, large - block - size transforms with sizes up to 64×64 are enabled, which are mainly used for higher - resolution videos such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high - frequency transform coefficients are zeroed out, such that only the lower - frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 - column transform coefficients are retained. Similarly, when N equals 64, only the first 32 - row transform coefficients are retained. When the transform skip mode is used for large blocks, the entire block is used without zeroing out any values. Additionally, the transform displacement is removed in the transform skip mode. VTM also supports a configurable maximum transform size in the SPS, such that the encoder has the flexibility to select a transform size of up to 32 lengths or 64 lengths according to the needs of a particular implementation. 2.2.2. Multiple transform selection (MTS) for kernel transforms In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of inter- and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the basis functions of the selected DST / DCT. Table 6 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that of the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified for intra and inter at the SPS level. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applicable to luma. The MTS signaling is skipped when one of the following conditions is met: - The position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC). - The last significant coefficient of the luma TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 7. A unified transform selection for ISP and implicit MTS is used by eliminating the intra-mode and block-shape dependencies. If the current block is in the ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for horizontal and vertical transform kernels. In terms of transform matrix precision, 8-bit primary transform kernels are used. Thus, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, for other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8), 8-bit primary transform kernels are used. Table 7 - Transform and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are zeroed. Only the coefficients within the 16×16 low-frequency region are retained. Similar to HEVC, the residual of a block can be coded and decoded using the transform skip mode. To avoid redundancy in syntax coding and decoding, when the MTS_CU_flag at the CU level is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled. 2.2.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, as Figure 37 shown, LFNST is applied between the forward main transform and quantization (at the encoder) and between de-quantization and the inverse main transform (at the decoder side). In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). The following uses the input as an example to describe the application of the non-separable transform used in LFNST. To apply 4×4 LFNST, the 4×4 input block X is first represented as a vector The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. Subsequently, the 16×1 coefficient vector is reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) for the block. Coefficients with smaller indices will be placed in the 4×4 coefficient block together with smaller scan indices. 2.2.3.1. Reduced Non-Separable Transform LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method, enabling it to be implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix size to minimize the computational complexity and the memory spatial domain for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N (N is usually equal to 64 for 8×8 NSST) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows: The transformed R rows are the R bases of the N-dimensional space. The inverse transformation matrix for RT is the transpose of its forward transformation. For the 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix (which is the size of the conventional 8×8 non-separable transformation matrix) is reduced to a 16×48 direct matrix. Thus, a 48×16 inverse RST matrix is used on the decoder side to generate the kernel (primary) transformation coefficients in the upper left 8×8 region. When applying the 16×48 matrix instead of the 16×64 with the same transformation set configuration, each of them takes 48 input data from three 4×4 blocks in the upper left 8×8 block except for the lower right 4×4 block. With the reduced size, the memory usage for storing all LFNST matrices is reduced from 10 KB to 8 KB with a reasonable performance degradation. To reduce complexity, it is applicable to limit the LFNST only when all coefficients outside the first coefficient subgroup are not significant. Thus, when applying the LFNST, all only the primary transformation coefficients must be zero. This allows adjusting the LFNST index signaling at the last valid position and thus avoids the additional coefficient scanning in the current LFNST design, which requires checking valid coefficients only at specific positions. The worst-case processing of the LFNST (in terms of multiplications per pixel) limits the non-separable transformation of the 4×4 block and the 8×8 block to 8×16 transformation and 8×48 transformation respectively. In these cases, when applying the LFNST, the last valid scan position must be less than 8 for other sizes less than 16. For blocks with shapes of 4×N and N×4 and N>8, the proposed limitation means that the LFNST is now applied only once and only to the upper left 4×4 region. Since all only the primary coefficients are zero when applying the LFNST, the number of operations required for the primary transformation is reduced in this case. From the encoder's perspective, when testing the LFNST transformation, the quantization of the coefficients is significantly simplified. Rate-distortion optimized quantization must be maximally done for the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.2.3.2. LFNST Transformation Selection There are a total of 4 transformation sets and 2 non-separable transformation matrices (kernels) used in the LFNST. As shown in Table 8, the mapping from the intra prediction mode to the transformation set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transformation set 0 is selected for the current chroma block. For each transformation set, the selected non-separable quadratic transformation candidate is further specified by the explicitly signaled LFNST index. The index is signaled once in the bitstream for each intra CU after the transformation coefficients. Table 8 - Transformation Selection Table IntraPredMode Transform Set Index IntraPredMode < 0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode <= 80 1 81 <= IntraPredMode <= 83 0 2.2.3.3. LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are not significant, the LFNST index encoding and decoding depends on the position of the last significant coefficient. Additionally, the LFNST index is context - decoded, but does not depend on the intra - prediction mode, and only the first binary bit is context - decoded. Moreover, LFNST is applied to intra CUs in both intra - slices and inter - slices as well as for both luminance and chrominance. If dual - tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter - slices (dual - tree is disabled), a single LFNST index is signaled and used for both luminance and chrominance. Considering that due to the existing maximum transform size limit (64×64), large CUs larger than 64×64 are implicitly partitioned (TU slicing), the LFNST index search can increase the data buffer up to four times the number of decoding pipeline stages. Therefore, the maximum size allowing LFNST is restricted to 64×64. Note that LFNST only enables DCT2. The LFNST index signaling is placed before the MTS index signaling. It is not obvious to use the scaling matrix for perceptual quantization, and the scaling matrix specified for the main matrix can be used for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For the single - tree partition mode, chrominance LFNST is not applied. 2.2.4. Sub - block Transform (SBT) In VTM, the sub - block transform is introduced for CUs in inter - prediction. In this transform mode, for a CU, only a sub - part of the residual block is encoded and decoded. When the cu_cbf of an inter - predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub - part of the residual block is encoded and decoded. For the former case, the inter - frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by the presumptive adaptive transform while the other part of the residual block is zeroed. When SBT is used for inter - decoded CUs, the SBT type and SBT position information are signaled in the bit - stream. There are two SBT types and two SBT positions, as Figure 38As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in terms of luma samples), a 1:3 / 3:1 partition along that dimension is not allowed. A CU can have at most 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal transform and the vertical transform for each SBT position are specified in Figure 38 . For example, the horizontal transform and the vertical transform for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slicing, cbf, and the horizontal and vertical kernel transform types for the residual blocks. SBT is not applied to CUs coded using a combined inter-intra mode. 2.2.5. Maximum Transform Size and Zeroing of Transform Coefficients Both the CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the maximum intra-coded block can have a size of 128×128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the primary transform process, there is no standardized zeroing operation applied to the transform coefficients. However, if LFNST is applied, the primary transform coefficients outside the LFNST region are standardized to zero. 2.2.6. Enhanced MTS for Intra Coding In the current VVC design, for MTS, only the DST7 transform kernel and the DCT8 transform kernel are utilized, which are used for both intra and inter coding. Additional primary transforms including DCT5, DST4, DST1, and the identity transform (IDT) are adopted. The MTS set also depends on the TU size and the intra mode information. 16 different TU sizes are considered, and for each TU size 5, different categories are considered according to the intra mode information. For each category, 1, 4, or 6 different transform pairs are considered. Multiple intra MTS candidates (among 1, 4, and 6 MTS candidates) are adaptively selected according to the sum of the absolute values of the transform coefficients. The sum is compared with two fixed thresholds to determine the total number of allowed MTS candidates: 1 candidate: sum <= th0. 4 candidates: th0 < sum <= th1. 6 candidates: sum > th1. Note that although 80 different categories are considered in total, some of these different categories usually share exactly the same set of transforms. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For the angular mode, joint symmetry on the TU shape and intra prediction is considered. Thus, a mode i (i > 34) with a TU shape A×B will be mapped to the same category corresponding to a mode j = (68 - i) with a TU shape B×A. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For the wide-angle mode, the closest regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. 2.2.7. Quadratic Transform: LFNST Extension with Large Kernels The LFNST in VVC is extended as follows: · The number of the LFNST set (S) and candidates (C) is extended to S = 35 and C = 3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: ο For predModeIntra < 2, lfnstTrSetIdx is equal to 2; ο lfnstTrSetIdx = predModeIntra, for predModeIntra in [0, 34]; ο lfnstTrSetIdx = 68 - predModeIntra, for predModeIntra in [35, 66]. · Three different kernels LFNST 4, LFNST 8, and LFNST 16 are defined to indicate the sets of LFNST kernels applied to 4×N / N×4 (N≥4), 8×N / N×8 (N≥8), and M×N (M, N≥16), respectively. The kernel size is specified as follows: (LFSNT4, LLFNST8*, LFNST16*) = (16×16, 32×64, 32×96) The forward LFNST is applied to the upper-left low-frequency region called the region of interest (ROI). When applying the LFNST, the main transform coefficients existing in the region other than the ROI are cleared and are not changed from the VVC standard. The ROI of LFNST16 is in Figure 39 shown. It consists of six 4×4 sub-blocks which are consecutive in the scan order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be R×96. In this contribution, R is chosen to be 32, and accordingly 32 coefficients (two 4×4 sub-blocks) are generated from the forward LFNST16, which are placed after the coefficient scan order. The ROI of LFNST8 is in Figure 40 shown. The forward LFNST8 matrix can be R×64, and R is chosen to be 32. The generated coefficients are positioned in the same way as LFNST 16. The mapping from the intra prediction mode to these sets is shown in Table 9, Table 9. Mapping of Intra Prediction Mode to LFNST Set Index Intra Prediction Mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST Set Index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra Prediction Mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST Set Index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra Prediction Mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST Set Index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.2.8. Sign Prediction The basic idea of the coefficient sign prediction method is to calculate the reconstruction residuals for the negative and positive sign combinations for the applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal sign, the cost function is defined as the measurement of discontinuity across Figure 41 the block boundaries shown above. All hypotheses are measured, and the hypothesis with the minimum cost is selected as the predicted value of the coefficient sign. The cost function is defined as the sum of the absolute second derivatives in the residual domain of the upper row and the left column as follows: where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. The term (-R -1 +2R 0 -P 1 ) can be calculated only once per block and only the residual hypothesis is subtracted. The transform coefficients with the maximum K qIdx value in the upper-left 4×4 region are selected. After compensating for the effects of multiple quantizers in DQ, the qIdx value is at the transform system level. A larger qIdx value will result in a larger dequantized transform coefficient level. The qIdx is derived as follows: qIdx = (abs(level) << 1) - (state & 1); where level is the transform coefficient level parsed from the bitstream, and state is a variable maintained by the encoder and decoder in DQ. The symbol prediction region is extended to a maximum of 32×32. Predict the symbols of the top-left M×N block. The values of M and N are calculated as follows: οM = min(w, maxW) οN = min(h, maxH) where w and h are the width and height of the transform block. The maximum region for symbol prediction is not always set to 32×32. The encoder sets the maximum region (maxW, maxH) based on the configuration, sequence class, and QP, and signals the region in the SPS. The maximum number of predicted symbols remains unchanged. Symbol prediction is also applied to the LFNST block. And for the LFNST block, symbol prediction is allowed for the 4 largest coefficients in the top-left 4×4 region. 3. Problems / Issues There are several problems in existing video coding and decoding technologies, which can be further improved for higher coding and decoding gains. 1. In ECM-6.0, affine candidates can be derived from adjacent affine-based candidates, history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. A similarity check is performed for affine candidate derivation. However, a differential similarity check rule is used for affine candidate derivation. This may not be optimal. 2. In ECM-6.0, hybrid modes such as CIIP and OBMC are applied to both videos captured by cameras and screen content videos, which may not be efficient. 3. In ECM-6.0, KLT is allowed for the explicit inter-frame MTS mode. Specifically, if the TU for inter-frame coding and decoding is less than or equal to 16×16, two KLT options (i.e., KLT0 and KLT1) of the inter-frame MTS kernel are used to replace DST7 and DCT8. This design can be changed for higher coding and decoding efficiency. 4. In ECM-6.0, KLT is allowed for the inter-frame MTS mode, but the use of KLT does not depend on which inter-frame prediction technique is used for the video unit, which can be further improved. 5. In ECM-6.0, the following intra-mode derivation / mapping for intra-frame MTS and LFNST indices can be improved. 1) In the case of the dual-tree, LFNST is applied to the chrominance component. For the CCLM mode, the co-located luma mode is used for the LFNST transform set and the transpose flag index. 2) For the MIP mode, the planar mode is used for the LFNST transform set and the transpose flag index. 3) MIP is regarded as a special mode for intra MTS transform classes and intra MTS transform pair indices. 4) Allow IntraTMP to use implicit MTS (e.g., DST7) and LFNST (regarded as planar mode). 5) In hybrid modes such as TIMD hybrid mode and DIMD hybrid mode, only the first intra mode is considered for MTS / LFNST indices. 4. Embodiments of the present disclosure The following detailed embodiments should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow way. In addition, these embodiments can be combined in any way. The term "video unit" or "coding unit" or "block" may represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. The term "KLT" may refer to a type of transform. For example, it may refer to the Karhunen - Loève transform. For example, it may refer to any transform type that is not DCT or DST or Hadmard. The coefficient matrix associated with a specific KLT can be trained online or predefined (e.g., offline training) according to some prior knowledge (e.g., residuals / coefficients from already decoded neighboring blocks). In the present disclosure, regarding "a block coded using mode N", here "mode N" can be a specific prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a specific prediction technique (e.g., AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, GPM intra, MHP, OBMC, LIC, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, sub - block coding, hypothesis coding, etc.), or a specific transform process (IDTX, DCT - X, DST - Y, KLT - Z, where X / Y / Z are constants) or a specific filter process (de - blocking, SAO, bilateral filter, adaptive loop filter, CCSAO, CC - ALF, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards. Any change in coding tools is also applicable. 4.1. Regarding the first problem of the derivation of affine candidates, the following method is proposed: a. The same logic / rules / process for similarity / consistency / deduplication checking can be used for the derivation of all affine candidates. a. For example, it can refer to the derivation of affine Merge candidates. b. For example, it may refer to the derivation of affine AMVP candidates. c. For example, it may refer to both the derivation of affine Merge candidates and the derivation of affine AMVP candidates. d. For example, it may refer to the derivation of history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. b. The logic / rules / procedures for similarity / consistency / deduplication checking may refer to comparing one or more of the following elements associated with a first affine candidate with these elements associated with a second affine candidate: a. Inter-frame direction (prediction direction); b. Affine type (e.g., 6-parameter affine or 4-parameter affine); c. Sub-block Merge type (e.g., sbTMVP or affine); d. Bcw index; e. LIC flag; f. Reference index; g. Motion vector (e.g., horizontal component and / or vertical component); h. Control point motion vector (CPMV); i. First CPMV (e.g., upper left CPMV) and / or second CPMV (e.g., upper right CPMV) and / or third CPMV (e.g., lower left CPMV); .j. Horizontal displacement and / or vertical displacement between the first CPMV and the second CPMV (e.g., absolute difference between the horizontal components and / or vertical components of the first CPMV and the second CPMV); k. Horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV (e.g., absolute difference between the horizontal components and / or vertical components of the first CPMV and the third CPMV). c. A consistency check can be performed on the comparison of the elements listed in item b. a. For example, if the elements associated with the second candidate are the same as the elements associated with the first candidate, the second candidate is not added to the affine candidate list. d. A similarity check can be performed on the comparison listed in item b. a. For example, if the elements associated with the second candidate are similar to the elements associated with the first candidate, the second candidate is not added to the affine candidate list. b. For example, "similar" may refer to a threshold-based comparison. i. For example, the absolute difference is less than the threshold. ii. For example, the absolute difference is not greater than the threshold. c. For example, the threshold may depend on the block size, such as the width and / or height. i. For example, the threshold is adaptively determined according to the block width / height. ii. For example, a smaller threshold may be set for a smaller block size, while a larger threshold may be set for a larger block size. d. For example, the threshold may be a predefined fixed value (such as 0 or 1). e. For example, similarity checks may be performed separately for the horizontal and vertical components of the motion vector. f. For example, similarity checks may be performed separately for the horizontal and vertical components of the control point motion vector. e. Whether to apply similarity / consistency / deduplication checks to the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV may depend on the affine type (e.g., 6-parameter affine or 4-parameter affine). a. For example, similarity / consistency / deduplication checks for the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV may be performed only when the affine types of the first affine candidate and the second affine candidate are 6-parameter. 4.2. Regarding the second issue of the codec tools for screen content tools, the following method is proposed: a. At least one of the following codec tools may be disallowed or constrained or prohibited for a video unit. a) CIIP and / or its variants (e.g., CIIP PDPC, CIIP TM, CIIP TIMD, etc.). b) OBMC and / or its variants (e.g., OBMC TM, etc.). c) TIMD and / or its variants. d) DIMD and / or its variants. e) MHP and / or its variants. f) DMVR and / or its variants. g) interTM and / or its variants. h) CCALF and / or its variants. i) CCSAO and / or its variants. b. Whether to disallow or constrain or prohibit the codec tool for a video unit may depend on the profile / level / layer. c. Whether to disallow or constrain or prohibit the codec tool for a video unit may depend on whether the video unit belongs to a specific video type. a) The specific video type may refer to a screen content video. d. In addition, the codec tool may be disallowed or constrained or prohibited for a video sequence or a group of pictures or a picture or a slice. a) Such a constraint or prohibition can be reflected by a bitstream constraint. b) Such a constraint, prohibition or allowance can be reflected by a syntax element (e.g., a flag) signaled in the bitstream. e. A syntax element (e.g., a flag) can be signaled in the bitstream to impose such a constraint, prohibition or allowance on a specific codec tool listed in item a. a) A syntax element can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. f. A single syntax element (e.g., a flag) can be signaled to impose such a constraint, prohibition or allowance on more than one codec tool listed in item a. a) A syntax element can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. g. Different weight values / factors / tables / sets can be allowed to mix multiple prediction hypotheses of video units encoded with codec tool X. a) For example, which weight value / factor / table / set is allowed to mix multiple prediction hypotheses of a video unit can depend on the content type. b) For example, a first weight value / factor / table / set can be allowed to mix multiple prediction hypotheses of a first video unit, while a second weight value / factor / table / set can be allowed to mix multiple prediction hypotheses of a second video unit. c) For example, the first video unit can belong to a video sequence captured by a camera. d) For example, the second video unit can belong to a screen content video sequence. e) For example, assume that two prediction hypotheses are mixed into the final prediction of a video unit: i. For a first type of video unit, the final prediction block of the video unit encoded with codec tool X can be "exactly equal to the first prediction hypothesis" or "exactly equal to the second prediction hypothesis". 1. For example, the allowed weight value / factor / table / set of a video unit can be equal to 0 or 1. ii. For a second type of video unit, the final mixed prediction block of the video unit encoded with codec tool X can be "a fusion of both the first prediction hypothesis and the second prediction hypothesis". 1. For example, the allowed weight values / factors / tables / sets for a video unit can be a fraction between 0 and 1 (e.g., the fractional value of the weight of a specific prediction hypothesis can be quantized to an integer in the codec). f) For example, the final prediction (e.g., the mixed) of a video unit with codec tool X can be further mixed / weighted / fused with another video unit coded / decoded with another codec tool. g) For example, a specific codec tool X can be CIIP and / or its variants. h) For example, a specific codec tool X can be GPM and / or its variants. i) For example, a specific codec tool X can be MHP and / or its variants. j) For example, a specific codec tool X can be OBMC and / or its variants. k) For example, a specific codec tool X can be TIMD hybrid mode and / or its variants. l) For example, a specific codec tool X can be DIMD hybrid mode and / or its variants. 4.3. Regarding the third issue of the general use of KLT for the transform process, the following method is proposed: a. More than two KLT kernels can be allowed in the codec. b. KLT kernels can be allowed for the primary transform and / or the secondary transform. c. KLT kernels can be allowed for the chrominance components. a. For example, different KLT kernels can be used for the luminance component and the chrominance components of a video unit. i. Alternatively, all color components of a video unit can share the same KLT kernel. b. For example, different KLT kernels can be used for the chrominance Cb component and the chrominance Cr component. i. Alternatively, the chrominance Cb component and the chrominance Cr component of a video unit can share the same KLT kernel. d. A pair {KLT, flipped-KLT} can be allowed / used for the {horizontal, vertical} transform or the {vertical, horizontal} transform of a video unit. a. For example, {KLT, flipped-KLT} represents a pair of transform kernels that includes a first KLT and a second KLT, where the transform coefficient matrix of the second KLT (i.e., the flipped-KLT) can be the transpose matrix of the transform coefficient matrix of the first KLT. e. The horizontal or vertical transform type of a video unit can be selected from {KLT-X, flipped-KLT-X, DCT-Y, DST-Z}, where X / Y / Z is a constant (e.g., Y = 8, Z = 7). a. For example, if more than one KLT kernel is defined in the codec, the more than one KLT kernel can be represented as KLT-X, where X is an integer value such as 1, 2, 3, ..., n, i.e., KLT-1, KLT-2, KLT-3, ..., KLT-n. b. For example, for a transform block, KLT can be used for a horizontal (or vertical) transform, while inverse-KLT can be used for a vertical (or horizontal) transform. c. For example, for a transform block, KLT can be used for a horizontal (or vertical) transform, while non-KLT can be used for a vertical (or horizontal) transform. d. For example, DCT2-KLT, KLT-DCT2 can be allowed for a horizontal-vertical transform or a vertical-horizontal transform of a block. e. For example, DST7-DCT2, DCT2-DST7, DCT8-DCT2, DCT2-DCT8 can be allowed for a horizontal-vertical transform or a vertical-horizontal transform of a block. f. For example, a video unit can be coded / decoded using a specific prediction / transformation / filter mode / technique. f. For a video unit coded using a specific mode, KLT can be the only transform type. a. In one example, the specific mode can be SBT. i. In one example, for different SBT modes (SBT_horizonta]_split, SBT_vertical_split, SBT_half_split, SBT_quad_split), KLT can be used for both the horizontal dimension and the vertical dimension. b. In one example, the specific mode can be MIP. c. In one example, the same KLT can be used for both the horizontal dimension and the vertical dimension of a video unit coded using a specific mode. d. In one example, different KLTs can be used for the horizontal dimension and the vertical dimension of a video unit coded using a specific mode. g. A KLT-based transform type can additionally be allowed for a video unit. a. For example, in addition to existing MTS options (e.g., MTS indices from 0 to 5), a KLT-based transform type can be explicitly signaled. b. For example, a first syntax element (e.g., a flag) can be signaled to indicate whether KLT is used for a video unit. i. Additionally, alternatively, if KLT is used for a video unit, a second syntax element (e.g., a flag or an index) can be signaled to indicate whether and / or which KLT is used for the horizontal transform and / or the vertical transform. c. For example, a signaling syntax element (e.g., an index) can be used to indicate which KLT is used for a video unit. i. For example, a signaling syntax element can be used to indicate which KLT pair is used for the horizontal transform and the vertical transform for a video unit. ii. For example, a signaling syntax element can be used to indicate which KLT is used for a specific size (e.g., width or height) of a video unit. iii. For example, a signaling syntax element can be used for both non-KLT transforms and KLT transforms. 1. For example, indices 0 to 1 indicate a DCT2-DCT2 pair and transform skip; while indices 2 to N indicate non-DCT2-DCT2 pairs and non-transform skip pairs including combinations of non-KLT and KLT. iv. For example, a signaling syntax element can be used when a KLT is used for a video unit. 1. For example, a signaling index (possible values starting from 0) can be used to indicate which KLT pair is used (i.e., at least one of the horizontal or vertical directions is using a KLT). 2. For example, in this case, the intra (and / or inter) MTS index may not be signaled (e.g., disabled) to the video unit. d. For example, which KLT-based transform type is used for a video unit can be determined implicitly. i. For example, the implicit determination can be based on the dimension / shape / size of the video unit. ii. For example, for a specific length of a block size (e.g., width or height), a specific KLT type is used in this direction without signaling. h. A KLT-based transform type can be applied to replace a specific existing transform type of a video unit. a. For example, the existing transform type to be replaced can be DST7, or DCT8, or DCT2. b. For example, a separable transform can be replaced by a KLT. c. For example, a primary transform and / or a secondary transform can be replaced by a KLT. i. For example, the KLT can be an inseparable KLT. j. For example, the KLT can be a separable KLT. k. For example, the video unit to which the KLT is applied can be intra-coded. l. For example, the video unit to which the KLT is applied can be inter-coded. m. For example, the KLT can be used as a primary transform. n. For example, the KLT can be used as a secondary transform. 4.4. Regarding the fourth issue of the mode - dependent KLT for the transformation process, the following method is proposed: a. Which KLT kernel to be used for a video unit can depend on a combination of at least one of the following types of codec information: a. The codec mode of the video unit. b. The size of the video unit. c. The motion vector of the video unit. d. The quantization parameter of the video unit. e. The temporal layer of the video unit. b. Which KLT kernel to be used for a video unit can depend on the codec mode of the video unit. a. For example, it can be based on whether SBT is used for the video unit. b. For example, it can be based on whether implicit MTS is used for the video unit. c. For example, it can be based on whether explicit MTS is used for the video unit. d. For example, it can be based on whether intra MTS is used for the video unit. e. For example, it can be based on whether inter MTS is used for the video unit. f. For example, it can be based on whether LFNST is used for the video unit. g. For example, it can be based on whether IBC and / or its variant modes are used for the video unit. h. For example, it can be based on whether PLT and / or its variant modes are used for the video unit. i. For example, it can be based on whether the intra prediction mode is used for the video unit. .j. For example, it can be based on whether ISP and / or its variant modes are used for the video unit. k. For example, it can be based on whether MIP and / or its variant modes are used for the video unit. l. For example, it can be based on whether DIMD and / or its variant modes are used for the video unit. m. For example, it can be based on whether TIMD and / or its variant modes are used for the video unit. n. For example, it can be based on whether LM / CCLM / CCCM / GLM and / or its variant modes are used for the video unit. o. For example, it can be based on whether the inter prediction mode is used for the video unit. p. For example, it can be based on whether the AMVP mode is used for the video unit. q. For example, it can be based on whether the Merge mode is used for the video unit. r. For example, it can be based on whether inter-frame / intra-frame / IBC template matching and / or its variant modes are used for the video unit. s. For example, it can be based on whether DMVR and / or its variant modes are used for the video unit. t. For example, it can be based on whether the sub-block prediction mode is used for the video unit. u. For example, it can be based on whether affine and / or its variant modes are used for the video unit. v. For example, it can be based on whether sbTMVP and / or its variant modes are used for the video unit. w. For example, it can be based on whether the hybrid / fusion / multi-hypothesis mode and / or its variant modes are used for the video unit. i. In one example, it can be based on whether the hybrid / fusion / multi-hypothesis mode includes an intra-coding part, such as GPM inter-intra, GPM intra, CIIP, MHP using intra, partitioned GPM, etc. x. For example, it can be based on whether GPM and / or its variant modes are used for the video unit. y. For example, it can be based on whether CIIP and / or its variant modes are used for the video unit. z. For example, it can be based on whether MHP and / or its variant modes are used for the video unit. aa. For example, it can be based on whether OBMC and / or its variant modes are used for the video unit. bb. For example, it can be based on whether LIC and / or its variant modes are used for the video unit. c. Whether KLT is used and / or which KLT kernel is used for the video unit can depend on the size of the video unit. a. Different KLTs can be applied to blocks with different sizes. b. For example, it can be based on whether the width (W) and / or height (H) of the video unit satisfy predefined conditions, such as one or more combinations of the following: i. W < T1 or W <= T1, where T1 can be 8 or 16 or 32 or 64. ii. W > T2 or W >= T2, where T2 can be 2 or 4 or 8. iii. H < T3 or H <= T3, where T3 can be 8 or 16 or 32 or 64. iv. H > T4 or H >= T4, where T4 can be 2 or 4 or 8. v. W / H < T5 or W / H <= T5, where T5 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vi. W / H > T6 or W / H >= T6, where T6 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vii. H / W < T7 or H / W <= T7, where T7 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. viii. H / W > T8 or H / W >= T8, where T8 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. ix. W == T9, where T9 can be 8 or 16 or 32 or 64. x. H == T10, where T10 can be 8 or 16 or 32 or 64. d. Which KLT kernel is used for a video unit can depend on the motion vector of the video unit. a. In one example, it depends on the magnitude of the motion vector. e. Which KLT kernel is used for a video unit can depend on the quantization parameter of the video unit. a. In one example, it depends on the base QP derived / signaled at a syntax level higher than the slice level (e.g., PPS or SPS). b. In one example, it depends on the slice QP. c. In one example, it depends on the QP of the coding / decoding unit. f. Which KLT kernel is used for a video unit can depend on the temporal layer of the video unit. a. In one example, it depends on whether it is in temporal layer 0 or in temporal layer 1, or in temporal layer 2, or.... g. Depending on the coding / decoding mode and / or size of the video unit, different KLT kernels can be allowed for different video units. a. In one example, a first KLT set can be used for a first mode set, while a second KLT set can be used for a second mode set. i. For example, the first KLT set contains at least one type of KLT kernel. ii. For example, the second KLT set contains at least one type of KLT kernel. iii. For example, the first mode set contains at least one type of prediction / transformation / filter mode. iv. For example, the second mode set contains at least one prediction / transformation / filter mode. b. In one example, a video unit coded / decoded using more than one different type of prediction / transformation / filter mode can use the same KLT kernel. c. Alternatively, video units coded with different types of prediction / transformation / filter modes may use different KLT kernels. d. For example, one type of prediction / transformation / filter mode may be a sub-block based prediction mode (e.g., affine, sbTMVP, etc.). e. For example, one type of prediction / transformation / filter mode may be an affine based prediction mode (e.g., affine AMVP, affine Merge, etc.). f. For example, one type of prediction / transformation / filter mode may be a hybrid / fusion / multi-hypothesis prediction mode (e.g., GPM inter-intra, GPM intra, CIIP, etc.). g. For example, one type of prediction / transformation / filter mode may be SBT and its variants. h. For example, one type of prediction / transformation / filter mode may be ISP and its variants. i. For example, one type of prediction / transformation / filter mode may be IBC and its variants. 4.5. Regarding the fifth issue of intra mode derivation / mapping for intra MTS and LFNST indices, the following methods are proposed: a. The intra mode derived from the information of neighboring samples can be used for the chroma transformation process. a. For example, in the case of single-tree and / or double-tree, the chroma transformation process may refer to the primary transformation of the chroma component (e.g., intra MTS, inter MTS, KLT, DCT2,......). b. For example, in the case of single-tree and / or double-tree, the chroma transformation process may refer to the secondary transformation of the chroma component (e.g., LFNST). c. For example, a template constructed from the upper neighboring samples and / or the left neighboring samples can be used to derive the intra mode. i. The gradient of the template samples can be used. ii. The maximum histogram amplitude value is constructed from the gradient (e.g., gradient histogram). iii. A DIMD-based method can be used. iv. A TIMD-based method can be used. v. The intra mode can be determined from the template samples (e.g., according to the measurement based on SAD / SATD cost) based on the preset intra mode candidates. vi. How many rows / columns and / or which neighboring samples to use can be based on the multi-reference row index. vii. How many rows / columns and / or which neighboring samples to use can be predefined (e.g., One row / column, or four rows / columns, etc.) viii. Adjacent samples in different rows / columns may have different weights for cost calculation. d. A derived intra mode may be generated for a video unit encoded / decoded using the following modes: i. CCLM mode and / or its variants. ii. CCCM mode and / or its variants. iii. GLM mode and / or its variants. iv. TIMD chroma mode and / or its variants. v. DIMD chroma mode and / or its variants. vi. Planar mode and / or its variants. vii. A chroma intra mode with a mode index greater than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). e. The derived intra mode may be used for MTS transform class index derivation. f. The derived intra mode may be used for MTS transform pair index derivation. g. The derived intra mode may be used for MTS transform index derivation. h. The derived intra mode may be used for LFNST transform set index derivation. i. The derived intra mode may be used for LFNST transpose flag derivation. b. An intra mode derived from information of adjacent samples may be used for a luminance transform process. a. For example, the luminance transform process may refer to a primary transform of the luminance component (e.g., intra MTS, inter MTS, KLT, DCT2,......). b. For example, the luminance transform process may refer to a secondary transform of the luminance component (e.g., LFNST). c. For example, a template constructed from upper adjacent samples and / or left adjacent samples may be used to derive an intra mode. i. The gradient of the template samples may be used. ii. The maximum histogram amplitude value is constructed from the gradient (e.g., gradient histogram). iii. A DIMD-based method may be used. iv. A TIMD-based method may be used. v. An intra mode may be determined from template samples (e.g., according to measurements based on SAD / SATD cost) based on a preset intra mode candidate. vi. How many rows / columns and / or which adjacent samples to use may be based on a multi-reference row index. vii. How many rows / columns and / or which neighboring samples can be predefined (e.g., one row / column, or four rows / columns, etc.). viii. Neighboring samples of different rows / columns can have different weights for cost calculation. d. Derived intra modes can be generated for video units encoded / decoded using the following modes: i. Intra template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Predictions obtained by mixing at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content encoding / decoding tools with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. e. Derived intra modes can be used for MTS transform class index derivation. f. Derived intra modes can be used for MTS transform pair index derivation. g. Derived intra modes can be used for MTS transform index derivation. h. Derived intra modes can be used for LFNST transform set index derivation. i. Derived intra modes can be used for LFNST transpose flag derivation. c. Intra modes derived from predicted samples of the current block can be used for the transform process. a. Predicted samples can refer to the prediction of the current video unit. b. Predicted samples can refer to the prediction of the template samples of the current video unit. c. The transform process can refer to the primary transform (e.g., intra MTS, inter MTS, KLT, DCT2,...). d. The transform process can refer to the secondary transform (e.g., LFNST). e. Derived intra modes can be generated for video units encoded / decoded using the following modes: i. Intra-template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Predictions obtained by mixing at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content coding / decoding tools with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. xv. CCLM mode and / or its variants. xvi. CCCM mode and / or its variants. xvii. GLM mode and / or its variants. xviii. Chrominance intra mode with a mode index larger than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). f. The derived intra mode can be used for MTS transform class index derivation. g. The derived intra mode can be used for MTS transform pair index derivation. h. The derived intra mode can be used for MTS transform index derivation. i. The derived intra mode can be used for LFNST transform set index derivation. .j. The derived intra mode can be used for LFNST transpose flag derivation. d. For example, the derived intra mode can be used for indexing the intra MTS transform class and / or transform pair and / or transform set for the block coded / decoded by MIP. a. The derived intra mode can be based on the predicted samples of the MIP block before MIP prediction upsampling. b. The derived intra mode can be based on the predicted samples of the MIP block after MIP matrix vector multiplication. c. The derived intra mode can be based on the horizontal / vertical gradient of the predicted samples of the MIP block. d. The derived intra mode may be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the MIP block. e. The derived intra mode may be based on the neighboring sample values of the MIP block. f. Alternatively, the derived intra mode may be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. e. For example, the derived intra mode may be used to index the LFNST transform set and / or the LFNST transpose flag for the blocks encoded / decoded for intra template matching (e.g., intraTMP, intraTM, etc.). a. The derived intra mode may be based on the prediction samples of the block encoded / decoded by intraTMP. b. The derived intra mode may be based on the horizontal / vertical gradients of the prediction samples of the block encoded / decoded by intraTMP. c. The derived intra mode may be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block encoded / decoded by intraTMP. d. The derived intra mode may be based on the neighboring sample values of the block encoded / decoded by intraTMP. e. The derived intra mode may be based on the intra mode of the reference block of the block encoded / decoded by intraTMP. f. The derived intra mode may be based on the gradients of the reference block of the block encoded / decoded by intraTMP. g. The derived intra mode may be based on the angle of the block vector (or motion vector) of the block encoded / decoded by intraTMP. h. The derived intra mode may be based on the intra mode information stored in the history-based intra mode cache for the block encoded / decoded by intraTMP. i. Alternatively, the derived intra mode may be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. f. For example, the MTS index may be signaled for the blocks encoded / decoded for intra template matching (e.g., intraTMP, intraTM, etc.). a. For example, the intra MTS index may be signaled for the blocks encoded / decoded by intraTMP. b. For example, the inter MTS index may be signaled for the blocks encoded / decoded by intraTMP. c. The derived intra mode may be used to index the MTS transform set, MTS transform class, and MTS transform pair for the blocks encoded / decoded by intraTMP. i. The derived intra mode may be based on the predicted samples of the block coded / decoded by intraTMP. ii. The derived intra mode may be based on the horizontal / vertical gradient of the predicted samples of the block coded / decoded by intraTMP. iii. The derived intra mode may be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block coded / decoded by intraTMP. iv. The derived intra mode may be based on the neighboring sample values of the block coded / decoded by intraTMP. v. The derived intra mode may be based on the intra mode of the reference block of the block coded / decoded by intraTMP. vi. The derived intra mode may be based on the gradient of the reference block of the block coded / decoded by intraTMP. vii. The derived intra mode may be based on the angle of the block vector (or motion vector) of the block coded / decoded by intraTMP. viii. The derived intra mode may be based on the intra mode information stored in the history-based intra mode cache for the block coded / decoded by intraTMP. ix. Alternatively, the derived intra mode may be a fixed / predefined mode (e.g., in addition to the planar mode), regardless of the current predicted samples and neighboring samples of the current block. d. Alternatively, the implicit MTS kernel may be applied to the block coded / decoded by intraTMP. i. For example, it may be determined based on the derived intra mode. ii. For example, it may be determined based on the block shape / size / dimension (e.g., width / height). iii. For example, it may be determined based on the values of the transform coefficients. iv. For example, for such an MTS kernel, no signaling syntax element (e.g., index) is passed. e. In addition, KLT may be used for the block coded / decoded by intraTMP. f. Alternatively, the main transform kernels other than DCT2 may be restricted for the block coded / decoded by intraTMP. g. For example, the intra MTS transform class and / or intra MTS transform pair and / or LFNST transform set and / or LFNST transpose flag of the TIMD hybrid mode may be derived based on the final predicted samples of the TIMD block. a. Alternatively, it may be derived based on neighboring sample information such as gradient, TIMD information, DIMD information, etc. b. Alternatively, it can be based on a fixed / predefined pattern (e.g., other than the planar pattern), regardless of the current predicted samples and neighboring samples of the current block. h. For example, the intra MTS transform class and / or intra MTS transform pair and / or LFNST transform set and / or LFNST transpose flag of the DIMD hybrid mode can be derived based on the final predicted samples of the DIMD block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, etc.). b. Alternatively, it can be based on a fixed / predefined pattern (e.g., other than the planar pattern), regardless of the current predicted samples and neighboring samples of the current block. i. For example, the intra MTS transform class and / or intra MTS transform pair and / or LFNST transform set and / or LFNST transpose flag of the planar (and / or planar horizontal and / or planar vertical) mode can be derived based on the final predicted samples of the block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, TIMD information, DIMD information, etc.). b. Alternatively, it can be based on a fixed / predefined pattern (e.g., other than the planar pattern), regardless of the current predicted samples and neighboring samples of the current block. 4.6. In one example, how to encode and decode the residual block can depend on the selected transform. 4.7. In one example, the KLT can be separable or non - separable. 4.8. In one example, whether or how to apply the KLT can depend on the sum of the absolute values of the transform coefficients. a. For example, whether to allow the KLT for intra (and / or inter) encoded / decoded blocks can be based on the sum of the absolute values of the transform coefficients of the block. b. For example, whether to allow the KLT for a specific block size / dimension / shape / orientation can be based on the sum of the absolute values of the transform coefficients of the block. c. For example, whether to apply a specific KLT - based transform kernel / set / pair to intra (and / or inter) encoded / decoded blocks can be based on the sum of the absolute values of the transform coefficients of the block. d. For example, the number of allowed KLT transform kernels / sets / pairs can be determined based on the sum of the absolute values of the transform coefficients of the block. e. For example, the sum of the absolute values of the transform coefficients of the block can be compared with at least one threshold. a. The threshold can be predefined. b. The threshold can be equal to a fixed value. c. The threshold can be based on predefined rules (e.g., block size / dimension, sequence resolution, Adaptively determined based on prediction mode, whether it is screen content, etc. General Aspects 4.9. Whether and / or how to apply the methods disclosed above can be signaled at the sequence level / picture group level / picture level / strip level / slice group level, for example, in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.10. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel. 4.11. The bitstream of the syntax elements signaled in the disclosed methods can be context-coded or bypass-coded. 4.12. Whether and / or how to apply the methods disclosed above can depend on the decoded information, such as block size, color format, single / double-tree segmentation, color component, strip / picture type.

[0105] More details of embodiments of the present disclosure related to transformation, screen content coding (SCC), and affine in image / video coding will be described below. The embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be construed in a narrow manner. In addition, these embodiments can be applied alone or in combination in any way.

[0106] As used herein, the term "block" can represent a color component, sub-picture, picture, strip, slice, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), sub-block of a video block, sub-region within a video block, a video processing unit including multiple samples / pixels, etc. The block can be rectangular or non-rectangular.

[0107] Figure 42 A flowchart of a method 4200 for video processing according to some embodiments of the present disclosure is shown. Method 4200 can be implemented during the conversion between the current video block of a video and the bitstream of the video. As Figure 42 shown, method 4200 starts at 4202, where the intra mode for the current video block is obtained. The intra mode is determined based on information associated with neighboring samples of the current video block and / or a prediction associated with the current video block.

[0108] In some embodiments, information associated with neighboring samples of a current video block may include the reconstructed values of the neighboring samples, the prediction mode of the neighboring samples, the transform type of the neighboring samples, and / or any other suitable coded or reconstructed information of the neighboring samples.

[0109] In some additional or alternative embodiments, the prediction associated with a current video block may include the prediction of the current video block or the prediction of the template samples of the current video block. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0110] At 4204, an intra mode is applied to a current video block based on an intra mode. As an example, a transform set or transform kernel used in the transform process may be selected from a predefined look-up table based on the index of the intra mode. It should be understood that the above description is for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0111] At 4206, a transformation is performed based on the application. In some embodiments, the transformation may include encoding the current video block into a bitstream. Alternatively or additionally, the transformation may include decoding the current video block from the bitstream.

[0112] In view of the above, an intra mode for a transform process is determined by considering information associated with neighboring samples of a video block and / or the prediction associated with the video block. Compared with conventional solutions, the proposed method can advantageously improve the coding and decoding efficiency.

[0113] In some embodiments, the intra mode may be determined based on a prediction associated with a current video block. In this case, the transformation process includes a primary transformation (e.g., intra MTS, inter MTS, KLT, DCT2, etc.) and / or a secondary transformation (e.g., LFNST, etc.). Additionally, the current video block may be encoded / decoded using at least one of the following modes: intra template matching mode, matrix weighted intra prediction (MIP) mode, intra sub - division (ISP) mode, a mode in which a prediction can be obtained by mixing from at least one intra mode and another mode, a template - based intra mode derivation (TIMD) hybrid mode, a decoder - side intra mode derivation (DIMD) hybrid mode, a geometric partition mode (GPM) intra mode, a partitioned GPM mode, a combined inter and intra prediction (CIIP) mode, a multi - hypothesis prediction (MHP) with intra mode mixing, a screen content encoding / decoding tool with intra mode mixing, a planar mode, a planar horizontal mode, a planar vertical mode, a cross - component linear model (CCLM) mode, a convolutional cross - component model (CCCM) mode, a gradient - based linear model (GLM) mode, or a chrominance intra mode having a mode index larger than a predetermined number, where the predetermined number represents an angular intra mode. By way of example and not limitation, the predetermined number may be 80. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0114] Additionally or alternatively, the intra mode may be determined based on information associated with neighboring samples of the current video block. In this case, the transformation process may include a luminance transformation process and / or a chrominance transformation process. For example, the luminance transformation process may include a primary transformation of the luminance component and / or a secondary transformation of the luminance component. The chrominance transformation process may include: a primary transformation of the chrominance component in a single - tree case, a primary transformation of the chrominance component in a dual - tree case, a secondary transformation of the chrominance component in a single - tree case, and / or a secondary transformation of the chrominance component in a dual - tree case.

[0115] In some embodiments, the neighboring samples may include a template that includes upper neighboring samples and / or left - hand neighboring samples of the current video block. For example, the information associated with the neighboring samples of the current video block may include: the gradient of the samples in the template, and / or the maximum histogram amplitude value determined based on the gradient of the samples in the template. Additionally or alternatively, the intra mode may be determined based on the template from a set of intra mode candidates.

[0116] In some embodiments, a DIMD - based scheme or a TIMD - based scheme may be used to determine the intra mode. In some additional embodiments, information regarding at least one of the following may depend on a multi - reference row index or be predefined: the number of rows of neighboring samples used to determine the intra mode, the number of columns of neighboring samples used to determine the intra mode, or the neighboring samples used to determine the intra mode.

[0117] In some embodiments, different rows of neighboring samples may be associated with different weights for cost calculation, and / or different columns of neighboring samples may be associated with different weights for cost calculation.

[0118] In some embodiments, the intra mode may be determined based on information associated with neighboring samples of the current video block. In such a case, the current video block may be encoded and decoded using at least one of the following: intra-template matching mode, MIP mode, ISP mode, a mode in which prediction can be obtained by mixing from at least one intra mode and another mode, TIMD hybrid mode, DIMD hybrid mode, GPM intra mode, partitioned GPM mode, CIIP mode, MHP with intra mode mixing, screen content coding and decoding tools with intra mode mixing, planar mode, planar horizontal mode, or planar vertical mode. It should be understood that the above examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0119] In some embodiments, the intra mode may be used to determine: multi-transform selection (MTS) transform class index, MTS transform pair index, MTS transform index, low-frequency non-separable transform (LFNST) transform set index, and / or LFNST transpose flag.

[0120] In some embodiments, the current video block may be encoded and decoded using the MIP mode, and at least one of the following may be selected for the current video block based on the intra mode: intra MTS transform class, transform pair, or transform set.

[0121] In some embodiments, the prediction associated with the current video block for determining the intra mode may include: prediction samples of the current video block before MIP prediction upsampling, and / or prediction samples of the current video block after MIP matrix vector multiplication.

[0122] In some embodiments, the current video block may be encoded and decoded using the MIP mode. In this case, the intra mode may be determined based on at least one of the following: horizontal gradient or vertical gradient of samples in the prediction for the current video block, maximum histogram amplitude value determined based on the gradient of the current video block, or neighboring sample values of the current video block.

[0123] In some embodiments, the current video block may be encoded and decoded using the intra-template matching mode. In one example, the intra-template matching mode may be a template matching-based prediction mode (intraTMP) for intra-coded blocks. In another example, the intra-template matching mode may be intraTM. It should be understood that the above examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0124] In this case, at least one of the following may be selected for the current video block based on an intra mode: an LFNST transform set, an LFNST transpose flag, an MTS transform set, an MTS transform class, or an MTS transform pair. Additionally, the intra mode may be determined based on at least one of the following: the predicted samples of the current video block, the horizontal gradient of the predicted samples of the current video block, the vertical gradient of the predicted samples of the current video block, the maximum histogram magnitude value determined based on the gradient of the current video block, the neighboring sample values of the current video block, the intra mode of a reference video block for the current video block, the gradient of the reference video block, the angle of a block vector for the current video block, the angle of a motion vector for the current video block, or the intra mode information stored in a history-based intra mode cache for the current video block.

[0125] In some embodiments, the intra mode of another video block of the video different from the current video block may be predefined and / or fixed. For example, the intra mode may be independent of the predicted samples and neighboring samples of the other video block.

[0126] In some embodiments, the current video block may be coded and decoded using intra template matching, and an MTS index for the current video block may be indicated in the bitstream. As an example, the MTS index may include an intra MTS index. Additionally or alternatively, the MTS index may include an inter MTS index. Further, an MTS kernel may be applied to the current video block. For example, the MTS kernel may be determined based on: the intra mode, the shape of the current video block, the size of the current video block, the dimension of the current video block, the width of the current video block, the height of the current video block, the transform coefficient values of the current video block. In some embodiments, the MTS kernel may not be signaled by syntax elements (such as indices, flags, etc.).

[0127] In some embodiments, the current video block may be coded and decoded using an intra template matching mode. In this case, the Karhunen - Loève transform (KLT) may be applied to the current video block. As used herein, the term "KLT" may refer to a type of transform. For example, it may refer to the Karhunen - Loève transform. For example, it may include any suitable transform type that is not DCT or DST or Hadamard. Additionally, the primary transform kernel other than the discrete cosine transform type 2 (DCT2) may be disabled or restricted for the current video block.

[0128] In some embodiments, the current video block may be coded and decoded using the TIMD hybrid mode. In this case, at least one of the following may be determined based on the final predicted samples of the current video block: the intra MTS transform class of the TIMD hybrid mode, the intra MTS transform pair of the TIMD hybrid mode, the LFNST transform set of the TIMD hybrid mode, or the LFNST transpose flag of the TIMD hybrid mode.

[0129] In some alternative embodiments, the current video block may be coded and decoded using the DIMD hybrid mode. In this case, at least one of the following may be determined based on the final predicted samples of the current video block: the intra MTS transform class of the DIMD hybrid mode, the intra MTS transform pair of the DIMD hybrid mode, the LFNST transform set of the DIMD hybrid mode, or the LFNST transpose flag of the DIMD hybrid mode.

[0130] In some additional embodiments, at least one of the following may be determined based on the final predicted samples of the current video block: the intra MTS transform class of the planar mode, the intra MTS transform pair of the planar mode, the LFNST transform set of the planar mode, or the LFNST transpose flag of the planar mode.

[0131] In some embodiments, at least one of the following may be determined based on additional information associated with neighboring samples of the current video block: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag. By way of example and not limitation, the additional information may include: the gradient of the neighboring samples, the TIMD information of the neighboring samples, the DIMD information of the neighboring samples.

[0132] In some additional embodiments, at least one of the following may be determined based on a predefined mode: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag.

[0133] In some embodiments, the coded binary bits of the syntax elements related to the proposed method may be context-coded. Alternatively, the coded binary bits of the syntax elements related to the proposed method may be bypass-coded. It should be understood that the above description and / or examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0134] According to a further embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. In the method, an intra mode for a current video block of the video is obtained. The intra mode is determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block. A transform process is applied to the current video block based on the intra mode. Further, a bitstream is generated based on the application.

[0135] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, an intra mode for a current video block of the video is obtained. The intra mode is determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block. A transform process is applied to the current video block based on the intra mode. Further, a bitstream is generated based on the application, and the bitstream is stored in a non-transitory computer-readable recording medium.

[0136] Figure 43 A flowchart of a method 4300 for video processing according to some embodiments of the present disclosure is shown. The method 4300 may be implemented during the conversion between a current video block of a video and a bitstream of the video. As Figure 43 shown, the method 4300 starts at 4302, where information about applying the Karhunen-Loeve transform (KLT) to the current video block is obtained. The information depends on the transform coefficients of the current video block.

[0137] In some embodiments, the information may include: whether to apply KLT to the current video block, how to apply KLT to the current video block. This will be described in detail below. By way of example and not limitation, the information may depend on the sum of the absolute values of the transform coefficients, the number of non-zero coefficients among the transform coefficients, the largest non-zero coefficient among the transform coefficients, the position of the last non-zero coefficient among the transform coefficients, etc.

[0138] At 4304, a conversion is performed based on the information. In some embodiments, the conversion may include encoding the current video block into the bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.

[0139] In view of the above, the information about applying KLT to a video block depends on the transform coefficients of the video block. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency.

[0140] In some embodiments, a current video block may be intra-coded or inter-coded. In this case, the information may include at least one of the following: whether KLT is allowed to be applied to the current video block, whether a specific KLT-based transform kernel is applied to the current video block, whether a specific KLT-based transform set is applied to the current video block, or whether a specific KLT-based transform pair is applied to the current video block.

[0141] Additionally or alternatively, the information may include: whether KLT is allowed for a specific block size, whether KLT is allowed for a specific block dimension, whether KLT is allowed for a specific block shape, whether KLT is allowed for a specific direction, the number of allowed KLT transform kernels, the number of allowed KLT transform sets, or the number of allowed KLT transform pairs.

[0142] In some embodiments, the sum of the absolute values of the transform coefficients may be compared with at least one threshold. In one example, the at least one threshold may be predefined. In another example, the at least one threshold may be equal to a fixed value. In another example, the at least one threshold may be determined based on predefined rules. By way of example and not limitation, the at least one threshold may be determined based on the following: the size of the current video block, the dimension of the current video block, the sequence resolution of the current video block, the prediction mode of the current video block, information on whether the current video is screen content.

[0143] In some embodiments, the KLT-based transform type for the current video block may be implicitly determined. For example, the KLT-based transform type may be determined based on the following: the size of the current video block, the dimension of the current video block, the shape of the current video block, etc. Alternatively, for a specific length of a block size (e.g., width or height), a specific KLT-based transform type may be used in this direction without signaling. In one example, if the width of the current video block is equal to a first value, the KLT-based transform type may be used in the direction along the width of the current video block without being indicated in the bitstream. In another example, if the height of the current video block is equal to a second value, the KLT-based transform type may be used in the direction along the height of the current video block without being indicated in the bitstream.

[0144] In some embodiments, a separable transform, a primary transform, and / or a secondary transform may be replaced by KLT. For example, KLT may be an inseparable KLT or a separable KLT. Additionally, KLT may be applied to the current video block, and the current video block may be intra-coded or inter-coded. In some embodiments, KLT may be applied to the current video block as a primary transform or a secondary transform.

[0145] In some embodiments, a first syntax element indicating whether the KLT is applied to the current video block may be included in the bitstream. If the KLT is applied to the current video block, the bitstream may include a second syntax element indicating at least one of the following: whether the KLT is used for horizontal transform, whether the KLT is used for vertical transform, the KLT for horizontal transform, or the KLT for vertical transform.

[0146] In some alternative or additional embodiments, a third syntax element indicating the KLT for the current video block may be included in the bitstream. Alternatively, the third syntax element indicating the KLT may be included in the bitstream only if the KLT is applied to the current video block.

[0147] In some embodiments, the third syntax element may indicate one of the following: a KLT pair for horizontal and vertical transforms for the current video block, the KLT for horizontal transform for the current video block, or the KLT for vertical transform for the current video block.

[0148] In some embodiments, the third syntax element may be signaled for both non-KLT and KLT transforms. By way of example and not limitation, the third syntax element may be an index. In this case, indices 0 to 1 may indicate a DCT2-DCT2 pair and transform skip. Additionally, indices 2 to N may indicate non-DCT2-DCT2 pairs and non-transform skip pairs including combinations of non-KLT and KLT, and N may be an integer.

[0149] In some embodiments, there is no intra-MTS index for the current video block in the bitstream, or there is no inter-MTS index for the current video block in the bitstream. For example, the intra-MTS index and / or the inter-MTS index may not be signaled, and the intra-MTS index and / or the inter-MTS index may be disabled for the current video block.

[0150] In some embodiments, the entropy-coded binary bits of the syntax elements related to the proposed method may be context-coded. Alternatively, the entropy-coded binary bits of the syntax elements related to the proposed method may be bypass-coded. It should be understood that the above description and / or examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.

[0151] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing. In the method, information about applying the Karhunen-Loeve transform (KLT) to a current video block of the video is obtained. The information depends on the transform coefficients of the current video block. Additionally, a bitstream is generated based on the information.

[0152] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, information regarding applying the Karhunen - Loève transform (KLT) to a current video block of the video is obtained. This information depends on the transform coefficients of the current video block. Further, a bitstream is generated based on the information; and the bitstream is stored in a non - transitory computer - readable recording medium.

[0153] The embodiments of the present disclosure can be described according to the following items, and the features can be combined in any reasonable manner.

[0154] Item 1. A method for video processing, comprising: obtaining an intra - mode for a current video block of a video for conversion between the current video block of the video and a bitstream of the video, the intra - mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra - mode; and performing the conversion based on the application.

[0155] Item 2. The method according to Item 1, wherein the prediction associated with the current video block includes a prediction of the current video block or a prediction of template samples of the current video block.

[0156] Item 3. The method according to any one of Items 1 to 2, wherein the intra - mode is determined based on the prediction associated with the current video block, and the transform process includes at least one of a primary transform or a secondary transform.

[0157] Item 4. The method according to any one of Items 1 to 3, wherein the intra - mode is determined based on the prediction associated with the current video block, and the current video block is encoded and decoded using at least one of the following: an intra - template matching mode, a matrix - weighted intra - prediction (MIP) mode, an intra - sub - division (ISP) mode, a mode in which a prediction is obtained by mixing from at least one intra - mode and another mode, a template - based intra - mode derivation (TIMD) hybrid mode, a decoder - side intra - mode derivation (DIMD) hybrid mode, a geometric - partition mode (GPM) intra - mode, a partitioned GPM mode, a combined inter - and intra - prediction (CIIP) mode, a multi - hypothesis prediction (MHP) with intra - mode mixing, a screen - content encoding and decoding tool with intra - mode mixing, a planar mode, a planar horizontal mode, a planar vertical mode, a cross - component linear model (CCLM) mode, a convolutional cross - component model (CCCM) mode, a gradient - based linear model (GLM) mode, or a chrominance intra - mode having a mode index larger than a predetermined number, the predetermined number representing an angular intra - mode.

[0158] Item 5. The method according to Item 4, wherein the predetermined number is 80.

[0159] Item 6. The method according to any one of Items 1 to 2, wherein the intra mode is determined based on the information associated with neighboring samples of the current video block, and the transformation process includes at least one of a luminance transformation process or a chrominance transformation process.

[0160] Item 7. The method according to Item 6, wherein the luminance transformation process includes at least one of a primary transformation of the luminance component or a secondary transformation of the luminance component, or wherein the chrominance transformation process includes at least one of the following: a primary transformation of the chrominance component in the single-tree case, a primary transformation of the chrominance component in the double-tree case, a secondary transformation of the chrominance component in the single-tree case, or a secondary transformation of the chrominance component in the double-tree case.

[0161] Item 8. The method according to any one of Items 1 to 2 and 6 to 7, wherein the neighboring samples include a template, and the template includes upper neighboring samples and / or left neighboring samples of the current video block.

[0162] Item 9. The method according to Item 8, wherein the information associated with the neighboring samples of the current video block includes at least one of the following: the gradient of the samples in the template, or the maximum histogram amplitude value determined based on the gradient of the samples in the template.

[0163] Item 10. The method according to Item 8, wherein the intra mode is determined from a set of intra mode candidates based on the template.

[0164] Item 11. The method according to any one of Items 1 to 2 and 6 to 10, wherein a DIMD-based scheme or a TIMD-based scheme is used to determine the intra mode.

[0165] Item 12. The method according to any one of Items 1 to 2 and 6 to 11, wherein the information about at least one of the following depends on a multi-reference line index or is predefined: the number of rows of neighboring samples for determining the intra mode, the number of columns of neighboring samples for determining the intra mode, or the neighboring samples for determining the intra mode.

[0166] Item 13. The method according to any one of Items 1 to 2 and 6 to 12, wherein different rows of neighboring samples are associated with different weights for cost calculation, or different columns of neighboring samples are associated with different weights for cost calculation.

[0167] Item 14. The method according to any one of Items 1 to 2 and 6 to 13, wherein the intra mode is determined based on the information associated with neighboring samples of the current video block, and the current video block is encoded and decoded using at least one of the following: intra template matching mode, MIP mode, ISP mode, a mode in which prediction is obtained by mixing from at least one intra mode and another mode, TIMD mixing mode, DIMD mixing mode, GPM intra mode, partitioned GPM mode, CIIP mode, MHP with intra mode mixing, screen content coding tool with intra mode mixing, planar mode, planar horizontal mode, or planar vertical mode.

[0168] Item 15. The method according to any one of Items 1 to 14, wherein the intra mode is used to determine at least one of the following: multi-transform selection (MTS) transform class index, MTS transform pair index, MTS transform index, low-frequency non-separable transform (LFNST) transform set index, or LFNST transpose flag.

[0169] Item 16. The method according to any one of Items 1 to 15, wherein the current video block is encoded and decoded using the MIP mode, and at least one of the following is selected for the current video block based on the intra mode: intra MTS transform class, transform pair, or transform set.

[0170] Item 17. The method according to Item 16, wherein the prediction associated with the current video block includes at least one of the following: prediction samples of the current video block before MIP prediction upsampling, or prediction samples of the current video block after MIP matrix vector multiplication.

[0171] Item 18. The method according to any one of Items 1 to 17, wherein the current video block is encoded and decoded using the MIP mode, and the intra mode is determined based on at least one of the following: horizontal gradient or vertical gradient of samples in the prediction for the current video block, maximum histogram amplitude value determined based on the gradient of the current video block, or neighboring sample values of the current video block.

[0172] Item 19. The method according to any one of Items 1 to 15, wherein the current video block is encoded and decoded using the intra template matching mode, and at least one of the following is selected for the current video block based on the intra mode: LFNST transform set, LFNST transpose flag, MTS transform set, MTS transform class, or MTS transform pair.

[0173] Item 20. The method according to Item 19, wherein the intra mode is determined based on at least one of the following: the predicted samples of the current video block, the horizontal gradient of the predicted samples of the current video block, the vertical gradient of the predicted samples of the current video block, the maximum histogram amplitude value determined based on the gradient of the current video block, the neighboring sample values of the current video block, the intra mode of the reference video block for the current video block, the gradient of the reference video block, the angle of the block vector for the current video block, the angle of the motion vector for the current video block, or the intra mode information stored in the history-based intra mode cache for the current video block.

[0174] Item 21. The method according to any one of Items 1 to 20, wherein the intra mode of another video block of the video different from the current video block is predefined.

[0175] Item 22. The method according to any one of Items 1 to 21, wherein the current video block is encoded and decoded using intra template matching, and the MTS index for the current video block is indicated in the bitstream.

[0176] Item 23. The method according to Item 22, wherein the MTS index includes an intra MTS index or an inter MTS index.

[0177] Item 24. The method according to any one of Items 1 to 21, wherein the current video block is encoded and decoded using an intra template matching mode, and an MTS kernel is applied to the current video block.

[0178] Item 25. The method according to Item 24, wherein the MTS kernel is determined based on at least one of the following: the intra mode, the shape of the current video block, the size of the current video block, the dimension of the current video block, the width of the current video block, the height of the current video block, or the transform coefficient value of the current video block.

[0179] Item 26. The method according to any one of Items 24 to 25, wherein no signaling syntax element is used for the MTS kernel.

[0180] Item 27. The method according to any one of Items 1 to 26, wherein the current video block is encoded and decoded using an intra template matching mode, and the Karhunen-Loeve transform (KLT) is applied to the current video block.

[0181] Item 28. The method according to any one of Items 1 to 26, wherein the current video block is encoded and decoded using an intra template matching mode, and the main transform kernel other than the discrete cosine transform type 2 (DCT2) is disabled for the current video block.

[0182] Item 29. The method according to any one of Items 19 to 28, wherein the intra-template matching mode includes a template matching-based prediction mode (intraTMP) for intra-coded / decoded blocks.

[0183] Item 30. The method according to any one of Items 1 to 29, wherein the current video block is coded / decoded using the TIMD hybrid mode, and at least one of the following is determined based on the final prediction samples of the current video block: the intra MTS transform class of the TIMD hybrid mode, the intra MTS transform pair of the TIMD hybrid mode, the LFNST transform set of the TIMD hybrid mode, or the LFNST transpose flag of the TIMD hybrid mode.

[0184] Item 31. The method according to any one of Items 1 to 29, wherein the current video block is coded / decoded using the DIMD hybrid mode, and at least one of the following is determined based on the final prediction samples of the current video block: the intra MTS transform class of the DIMD hybrid mode, the intra MTS transform pair of the DIMD hybrid mode, the LFNST transform set of the DIMD hybrid mode, or the LFNST transpose flag of the DIMD hybrid mode.

[0185] Item 32. The method according to any one of Items 1 to 29, wherein at least one of the following is determined based on the final prediction samples of the current video block: the intra MTS transform class of the planar mode, the intra MTS transform pair of the planar mode, the LFNST transform set of the planar mode, or the LFNST transpose flag of the planar mode.

[0186] Item 33. The method according to any one of Items 1 to 29, wherein at least one of the following is determined based on additional information associated with neighboring samples of the current video block: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag.

[0187] Item 34. The method according to Item 31, wherein the additional information includes at least one of the following: the gradient of the neighboring samples, the TIMD information of the neighboring samples, or the DIMD information of the neighboring samples.

[0188] Item 35. The method according to any one of Items 1 to 29, wherein at least one of the following is determined based on a predefined mode: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag.

[0189] Item 36. A method for video processing, comprising: obtaining information about applying the Karhunen-Loeve transform (KLT) to a current video block of a video for conversion between the current video block of the video and a bitstream of the video, the information depending on transform coefficients of the current video block; and performing the conversion based on the information.

[0190] Item 37. The method according to item 26, wherein the information depends on a sum of absolute values of the transform coefficients.

[0191] Item 38. The method according to any one of items 36 to 37, wherein the information includes at least one of the following: whether to apply the KLT to the current video block, or how to apply the KLT to the current video block.

[0192] Item 39. The method according to any one of items 36 to 38, wherein the current video block is intra-coded or inter-coded, and the information includes at least one of the following: whether to allow the KLT to be applied to the current video block, whether to apply a specific KLT-based transform kernel to the current video block, whether to apply a specific KLT-based transform set to the current video block, or whether to apply a specific KLT-based transform pair to the current video block.

[0193] Item 40. The method according to any one of items 36 to 39, wherein the information includes at least one of the following: whether to allow the KLT for a specific block size, whether to allow the KLT for a specific block dimension, whether to allow the KLT for a specific block shape, whether to allow the KLT for a specific direction, a number of allowed KLT transform kernels, a number of allowed KLT transform sets, or a number of allowed KLT transform pairs.

[0194] Item 41. The method according to any one of items 37 to 40, wherein the sum of the absolute values of the transform coefficients is compared with at least one threshold.

[0195] Item 42. The method according to item 41, wherein the at least one threshold is predefined or equal to a fixed value.

[0196] Item 43. The method according to item 41, wherein the at least one threshold is determined based on a predefined rule.

[0197] Item 44. The method according to item 41, wherein the at least one threshold is determined based on at least one of the following: a size of the current video block, a dimension of the current video block, a sequence resolution of the current video block, a prediction mode of the current video block, or information about whether the current video is screen content.

[0198] Item 45. The method according to any one of Items 36 to 44, wherein the KLT-based transform type for the current video block is implicitly determined.

[0199] Item 46. The method according to Item 45, wherein the KLT-based transform type is determined based on at least one of the following: the size of the current video block, the dimension of the current video block, or the shape of the current video block.

[0200] Item 47. The method according to Item 45, wherein if the width of the current video block is equal to a first value, the KLT-based transform type is used in the direction along the width of the current video block without being indicated in the bitstream, or if the height of the current video block is equal to a second value, the KLT-based transform type is used in the direction along the height of the current video block without being indicated in the bitstream.

[0201] Item 48. The method according to any one of Items 36 to 47, wherein at least one of the following is replaced by the KLT: separable transform, primary transform, or secondary transform.

[0202] Item 49. The method according to any one of Items 36 to 48, wherein the KLT is an inseparable KLT or a separable KLT.

[0203] Item 50. The method according to any one of Items 36 to 49, wherein the KLT is applied to the current video block, and the current video block is intra-coded or inter-coded.

[0204] Item 51. The method according to any one of Items 36 to 49, wherein the KLT is applied to the current video block as a primary transform or a secondary transform.

[0205] Item 52. The method according to any one of Items 36 to 51, wherein a first syntax element indicating whether the KLT is applied to the current video block is included in the bitstream.

[0206] Item 53. The method according to any one of Items 36 to 52, wherein if the KLT is applied to the current video block, the bitstream includes a second syntax element indicating at least one of the following: whether the KLT is used for horizontal transform, whether the KLT is used for vertical transform, the KLT for the horizontal transform, or the KLT for the vertical transform.

[0207] Item 54. The method according to any one of Items 36 to 53, wherein a third syntax element indicating the KLT for the current video block is included in the bitstream.

[0208] Item 55. The method according to any one of Items 36 to 53, wherein if the KLT is applied to the current video block, a third syntax element indicating the KLT is included in the bitstream.

[0209] Item 56. The method according to any one of Items 54 to 55, wherein the third syntax element indicates one of the following: a KLT pair for horizontal and vertical transforms for the current video block, a KLT for the horizontal transform for the current video block, or a KLT for the vertical transform for the current video block.

[0210] Item 57. The method according to any one of Items 54 to 56, wherein the third syntax element is signaled for both non-KLT and KLT transforms.

[0211] Item 58. The method according to Item 57, wherein the third syntax element includes an index, where index 0 to 1 indicates a DCT2-DCT2 pair and transform skip, and index 2 to N indicates a non-DCT2-DCT2 pair and a non-transform skip pair including a combination of non-KLT and KLT, and N is an integer.

[0212] Item 59. The method according to any one of Items 54 to 58, wherein at least one of the following is not present in the bitstream: an intra MTS index for the current video block, or an inter MTS index for the current video block.

[0213] Item 60. The method according to any one of Items 1 to 59, wherein the entropy coded binary bits of the syntax elements associated with the method are context coded or bypass coded.

[0214] Item 61. The method according to any one of Items 1 to 60, wherein the transform includes encoding the current video block into the bitstream.

[0215] Item 62. The method according to any one of Items 1 to 60, wherein the transform includes decoding the current video block from the bitstream.

[0216] Item 63. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 62.

[0217] Item 64. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 62.

[0218] Article 65. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing, wherein the method includes: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and generating the bitstream based on the application.

[0219] Article 66. A method for storing a bitstream of a video includes: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.

[0220] Article 67. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing, wherein the method includes: obtaining information about applying a Karhunen-Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; and generating the bitstream based on the information.

[0221] Article 68. A method for storing a bitstream of a video includes: obtaining information about applying a Karhunen-Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; generating the bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device

[0222] Figure 44 A block diagram of a computing device 4400 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4400 may be implemented as the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).

[0223] It should be understood that Figure 44 the computing device 4400 shown is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.

[0224] AsFigure 44 As shown, computing device 4400 includes general computing device 4400. Computing device 4400 may at least include one or more processors or processing units 4410, a memory 4420, a storage unit 4430, one or more communication units 4440, one or more input devices 4450, and one or more output devices 4440.

[0225] In some embodiments, computing device 4400 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that computing device 4400 may support any type of interface to the user (such as a "wearable" circuitry, etc.).

[0226] Processing unit 4410 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in memory 4420. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of computing device 4400. Processing unit 4410 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0227] Computing device 4400 generally includes various computer storage media. Such media may be any media accessible by computing device 4400, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 4420 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 4430 may be any removable or non-removable media, and may include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in computing device 4400.

[0228] The computing device 4400 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 44 , a disk drive for reading from and / or writing to a removable non-volatile disk and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0229] The communication unit 4440 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4400 may be implemented by a single computing cluster or multiple computer machines, which may communicate via a communication connection. Thus, the computing device 4400 may operate in a networked environment using a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0230] The input device 4450 may be one or more of a variety of input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 4460 may be one or more of a variety of output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 4440, the computing device 4400 may also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 4400 may also communicate with one or more devices that enable a user to interact with the computing device 4400, or if needed, the computing device 4400 may also communicate with any device that enables the computing device 4400 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be carried out via an input / output (I / O) interface (not shown).

[0231] In some embodiments, some or all components of computing device 4400 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across the locations of remote data centers. The cloud computing infrastructure may provide services through shared data centers, although to a user they appear as a single access point. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0232] In embodiments of the present disclosure, computing device 4400 may be used to implement video encoding / decoding. Memory 4420 may include one or more video codec modules 4425 having one or more program instructions. These modules are accessible and executable by processing unit 4410 to perform the functions of the various embodiments described herein.

[0233] In an example embodiment of performing video encoding, input device 4450 may receive video data as input 4470 to be encoded. The video data may be processed, for example, by video codec module 4425 to generate an encoded bitstream. The encoded bitstream may be provided as output 4480 via output device 4460.

[0234] In an example embodiment of performing video decoding, input device 4450 may receive the encoded bitstream as input 4470. The encoded bitstream may be processed, for example, by video codec module 4425 to generate decoded video data. The decoded video data may be provided as output 4480 via output device 4460.

[0235] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Accordingly, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: obtaining an intra mode for a current video block of a video with respect to a conversion between the current video block of the video and a bitstream of the video, the intra mode being determined based on at least one of the following: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and performing the conversion based on the application.

2. The method according to claim 1, wherein the prediction associated with the current video block includes a prediction of the current video block or a prediction of template samples of the current video block.

3. The method according to any one of claims 1 to 2, wherein the intra mode is determined based on the prediction associated with the current video block, and the transform process includes at least one of a primary transform or a secondary transform.

4. The method according to any one of claims 1 to 3, wherein the intra mode is determined based on the prediction associated with the current video block, and the current video block is encoded and decoded using at least one of the following: intra template matching mode, matrix weighted intra prediction (MIP) mode, intra sub - partition (ISP) mode, a mode in which a prediction is obtained by mixing from at least one intra mode and another mode, template - based intra mode derivation (TIMD) hybrid mode, decoder - side intra mode derivation (DIMD) hybrid mode, geometric partition mode (GPM) intra mode, partitioned GPM mode, combined inter and intra prediction (CIIP) mode, multi - hypothesis prediction (MHP) with intra mode mixing, screen content encoding and decoding tools with intra mode mixing, plane mode, plane horizontal mode, plane vertical mode, cross - component linear model (CCLM) mode, convolutional cross - component model (CCCM) mode, gradient - based linear model (GLM) mode, or a chrominance intra mode having a mode index larger than a predetermined number, the predetermined number representing an angular intra mode.

5. The method according to claim 4, wherein the predetermined number is 80.

6. The method according to any one of claims 1 to 2, wherein the intra mode is determined based on the information associated with neighboring samples of the current video block, and the transform process includes at least one of a luminance transform process or a chrominance transform process.

7. The method according to claim 6, wherein the luminance transform process includes at least one of a primary transform of a luminance component or a secondary transform of a luminance component, or wherein the chrominance transform process includes at least one of the following: a primary transform of a chrominance component in a single - tree case, a primary transform of a chrominance component in a dual - tree case, a secondary transform of a chrominance component in a single - tree case, or a secondary transform of a chrominance component in a dual - tree case.

8. The method according to any one of claims 1 to 2 and 6 to 7, wherein the neighboring samples include a template, the template including upper neighboring samples and / or left - hand neighboring samples of the current video block.

9. The method according to claim 8, wherein the information associated with the neighboring samples of the current video block includes at least one of the following: the gradient of the samples in the template, or the maximum histogram amplitude value determined based on the gradient of the samples in the template.

10. The method according to claim 8, wherein the intra mode is determined from a set of intra mode candidates based on the template.

11. The method according to any one of claims 1 to 2 and 6 to 10, wherein a DIMD-based scheme or a TIMD-based scheme is used to determine the intra mode.

12. The method according to any one of claims 1 to 2 and 6 to 11, wherein the information about at least one of the following depends on a multi-reference line index or is predefined: the number of rows of neighboring samples used to determine the intra mode, the number of columns of neighboring samples used to determine the intra mode, or the neighboring samples used to determine the intra mode.

13. The method according to any one of claims 1 to 2 and 6 to 12, wherein different rows of neighboring samples are associated with different weights for cost calculation, or different columns of neighboring samples are associated with different weights for cost calculation.

14. The method according to any one of claims 1 to 2 and 6 to 13, wherein the intra mode is determined based on the information associated with the neighboring samples of the current video block, and the current video block is encoded and decoded using at least one of the following: intra template matching mode, MIP mode, ISP mode, where a mode in which prediction is obtained by mixing at least one intra mode and another mode, TIMD mixing mode, DIMD mixing mode, GPM intra mode, partitioned GPM mode, CIIP mode, MHP with intra mode mixing, screen content encoding and decoding tool with intra mode mixing, plane mode, plane horizontal mode, or plane vertical mode.

15. The method according to any one of claims 1 to 14, wherein the intra mode is used to determine at least one of the following: multi-transform selection (MTS) transform class index, MTS transform pair index, MTS transform index, low-frequency non-separable transform (LFNST) transform set index, or LFNST transpose flag.

16. The method according to any one of claims 1 to 15, wherein the current video block is encoded and decoded using the MIP mode, and at least one of the following is selected for the current video block based on the intra mode: intra MTS transform class, transform pair, or transform set.

17. The method according to claim 16, wherein the prediction associated with the current video block includes at least one of the following: the predicted samples of the current video block before MIP prediction upsampling, or the predicted samples of the current video block after MIP matrix vector multiplication.

18. The method according to any one of claims 1 to 17, wherein the current video block is encoded and decoded using the MIP mode, and the intra mode is determined based on at least one of the following: the horizontal gradient or vertical gradient of the samples in the prediction for the current video block, The maximum histogram amplitude value determined based on the gradient of the current video block, or The neighboring sample values of the current video block.

19. The method according to any one of claims 1 to 15, wherein the current video block is encoded and decoded using an intra-template matching mode, and at least one of the following is selected for the current video block based on the intra mode: LFNST transform set, LFNST transpose flag, MTS transform set, MTS transform class, or MTS transform pair.

20. The method according to claim 19, wherein the intra mode is determined based on at least one of the following: The predicted samples of the current video block, The horizontal gradient of the predicted samples of the current video block, The vertical gradient of the predicted samples of the current video block, The maximum histogram amplitude value determined based on the gradient of the current video block, The neighboring sample values of the current video block, The intra mode of the reference video block for the current video block, The gradient of the reference video block, The angle of the block vector for the current video block, The angle of the motion vector for the current video block, or The intra mode information stored in the history-based intra mode cache for the current video block.

21. The method according to any one of claims 1 to 20, wherein the intra mode of another video block of the video different from the current video block is predefined.

22. The method according to any one of claims 1 to 21, wherein the current video block is encoded and decoded using intra-template matching, and the MTS index for the current video block is indicated in the bitstream.

23. The method according to claim 22, wherein the MTS index includes an intra MTS index or an inter MTS index.

24. The method according to any one of claims 1 to 21, wherein the current video block is encoded and decoded using an intra-template matching mode, and an MTS kernel is applied to the current video block.

25. The method according to claim 24, wherein the MTS kernel is determined based on at least one of the following: The intra mode, The shape of the current video block, The size of the current video block, The dimension of the current video block, The width of the current video block, The height of the current video block, or The transform coefficient value of the current video block.

26. The method according to any one of claims 24 to 25, wherein no signaling syntax element is used for the MTS kernel.

27. The method according to any one of claims 1 to 26, wherein the current video block is encoded and decoded using an intra-template matching mode, and the Karhunen-Loeve transform (KLT) is applied to the current video block.

28. The method according to any one of claims 1 to 26, wherein the current video block is encoded and decoded using an intra-template matching mode, and the main transform kernels other than the discrete cosine transform type 2 (DCT2) are disabled for the current video block.

29. The method according to any one of claims 19 to 28, wherein the intra-template matching mode includes an intra-template matching based prediction mode (intraTMP) for intra-coded / decoded blocks.

30. The method according to any one of claims 1 to 29, wherein the current video block is coded / decoded using the TIMD hybrid mode, and at least one of the following is determined based on the final predicted samples of the current video block: the intra MTS transform class of the TIMD hybrid mode, the intra MTS transform pair of the TIMD hybrid mode, the LFNST transform set of the TIMD hybrid mode, or the LFNST transpose flag of the TIMD hybrid mode.

31. The method according to any one of claims 1 to 29, wherein the current video block is coded / decoded using the DIMD hybrid mode, and at least one of the following is determined based on the final predicted samples of the current video block: the intra MTS transform class of the DIMD hybrid mode, the intra MTS transform pair of the DIMD hybrid mode, the LFNST transform set of the DIMD hybrid mode, or the LFNST transpose flag of the DIMD hybrid mode.

32. The method according to any one of claims 1 to 29, wherein at least one of the following is determined based on the final predicted samples of the current video block: the intra MTS transform class of the planar mode, the intra MTS transform pair of the planar mode, the LFNST transform set of the planar mode, or the LFNST transpose flag of the planar mode.

33. The method according to any one of claims 1 to 29, wherein at least one of the following is determined based on additional information associated with neighboring samples of the current video block: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag.

34. The method according to claim 31, wherein the additional information includes at least one of the following: the gradient of the neighboring samples, the TIMD information of the neighboring samples, or the DIMD information of the neighboring samples.

35. The method according to any one of claims 1 to 29, wherein at least one of the following is determined based on a predefined mode: the intra MTS transform class, the intra MTS transform pair, the LFNST transform set, or the LFNST transpose flag.

36. A method for video processing, comprising: obtaining information regarding the application of the Karhunen-Loeve transform (KLT) to a current video block of a video for conversion between the current video block of the video and a bitstream of the video, the information being dependent on the transform coefficients of the current video block; and performing the conversion based on the information.

37. The method according to claim 26, wherein the information is dependent on the sum of the absolute values of the transform coefficients.

38. The method according to any one of claims 36 to 37, wherein the information includes at least one of the following: whether to apply the KLT to the current video block, or how to apply the KLT to the current video block.

39. The method according to any one of claims 36 to 38, wherein the current video block is intra-coded or inter-coded, and the information includes at least one of the following: whether the KLT is allowed to be applied to the current video block, whether a specific KLT-based transform kernel is applied to the current video block, whether a specific KLT-based transform set is applied to the current video block, or whether a specific KLT-based transform pair is applied to the current video block.

40. The method according to any one of claims 36 to 39, wherein the information includes at least one of the following: whether KLT is allowed for a specific block size, whether KLT is allowed for a specific block dimension, whether KLT is allowed for a specific block shape, whether KLT is allowed for a specific direction, the number of allowed KLT transform kernels, the number of allowed KLT transform sets, or the number of allowed KLT transform pairs.

41. The method according to any one of claims 37 to 40, wherein the sum of the absolute values of the transform coefficients is compared with at least one threshold.

42. The method according to claim 41, wherein the at least one threshold is predefined or equal to a fixed value.

43. The method according to claim 41, wherein the at least one threshold is determined based on predefined rules.

44. The method according to claim 41, wherein the at least one threshold is determined based on at least one of the following: the size of the current video block, the dimension of the current video block, the sequence resolution of the current video block, the prediction mode of the current video block, or information on whether the current video is screen content.

45. The method according to any one of claims 36 to 44, wherein the KLT-based transform type for the current video block is implicitly determined.

46. The method according to claim 45, wherein the KLT-based transform type is determined based on at least one of the following: the size of the current video block, the dimension of the current video block, or the shape of the current video block.

47. The method according to claim 45, wherein if the width of the current video block is equal to a first value, the KLT-based transform type is used in the direction along the width of the current video block without being indicated in the bitstream, or if the height of the current video block is equal to a second value, the KLT-based transform type is used in the direction along the height of the current video block without being indicated in the bitstream.

48. The method according to any one of claims 36 to 47, wherein at least one of the following is replaced by the KLT: separable transform, primary transform, or secondary transform.

49. The method according to any one of claims 36 to 48, wherein the KLT is an inseparable KLT or a separable KLT.

50. The method according to any one of claims 36 to 49, wherein the KLT is applied to the current video block, and the current video block is intra-coded or inter-coded.

51. The method according to any one of claims 36 to 49, wherein the KLT is applied to the current video block as a primary transform or a secondary transform.

52. The method according to any one of claims 36 to 51, wherein a first syntax element indicating whether the KLT is applied to the current video block is included in the bitstream.

53. The method according to any one of claims 36 to 52, wherein if the KLT is applied to the current video block, the bitstream includes a second syntax element indicating at least one of the following: whether the KLT is used for horizontal transform, whether the KLT is used for vertical transform, the KLT for the horizontal transform, or the KLT for the vertical transform.

54. The method according to any one of claims 36 to 53, wherein a third syntax element indicating the KLT for the current video block is included in the bitstream.

55. The method according to any one of claims 36 to 53, wherein if the KLT is applied to the current video block, the third syntax element indicating the KLT is included in the bitstream.

56. The method according to any one of claims 54 to 55, wherein the third syntax element indicates one of the following: a KLT pair for horizontal and vertical transforms for the current video block, the KLT for the horizontal transform for the current video block, or the KLT for the vertical transform for the current video block.

57. The method according to any one of claims 54 to 56, wherein the third syntax element is signaled for both non-KLT and KLT transforms.

58. The method according to claim 57, wherein the third syntax element includes an index, where indices 0 to 1 indicate a DCT2-DCT2 pair and transform skip, and indices 2 to N indicate non-DCT2-DCT2 pairs and non-transform skip pairs including a combination of non-KLT and KLT, and N is an integer.

59. The method according to any one of claims 54 to 58, wherein at least one of the following is not present in the bitstream: an intra MTS index for the current video block, or an inter MTS index for the current video block.

60. The method according to any one of claims 1 to 59, wherein the entropy-coded binary bits of the syntax elements related to the method are context-coded or bypass-coded.

61. The method according to any one of claims 1 to 60, wherein the transformation includes encoding the current video block into the bitstream.

62. The method according to any one of claims 1 to 60, wherein the transformation includes decoding the current video block from the bitstream.

63. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 62.

64. A non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to any one of claims 1 to 62.

65. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing, where the method comprises: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; and generating the bitstream based on the application.

66. A method for storing a bitstream of a video, comprises: obtaining an intra mode for a current video block of the video, the intra mode being determined based on at least one of: information associated with neighboring samples of the current video block, or a prediction associated with the current video block; applying a transform process to the current video block based on the intra mode; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.

67. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing, where the method comprises: obtaining information about applying a Karhunen - Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; and generating the bitstream based on the information.

68. A method for storing a bitstream of a video, comprises: obtaining information about applying a Karhunen - Loeve transform (KLT) to a current video block of the video, the information depending on transform coefficients of the current video block; generating the bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium.