Method and device for video processing and medium

By acquiring multiple affine candidates between the video block and the bitstream and performing deduplication checks, the problem of insufficient video encoding and decoding efficiency in the prior art is solved, and efficient encoding and decoding of different types of video blocks is realized.

CN120035988APending Publication Date: 2025-05-23DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072902.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-14
Filing Date
2023-10-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has shortcomings in improving the encoding and decoding efficiency, especially when dealing with different types of video blocks, it is difficult to effectively support efficient encoding and decoding.

Method used

By converting the video block and the bitstream, multiple affine candidates are acquired and deduplication checks are performed, and efficient encoding and decoding of the video block is realized based on this.

Benefits of technology

The encoding and codec efficiency of video encoding and codec technology is improved, especially when processing different types of video blocks, it can support efficient encoding and codec more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035988A_ABST
    Figure CN120035988A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises the following steps: for conversion between a current video block of a video and a bit stream of the video, acquiring a plurality of affine candidates for the current video block; applying a deduplication check to the plurality of affine candidates according to a check process that is common to a plurality of different types of affine candidates; and performing the transition based on the application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to video encoding and decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Versatile Video Codec (VVC) standard. However, it is generally expected to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided. The method includes: obtaining a plurality of affine candidates for a current video block of a video for conversion between the current video block and a bitstream of the video; applying a deduplication check to the plurality of affine candidates according to a check process, the check process being common to a plurality of different types of affine candidates; and performing conversion based on the application.

[0005] According to the method of the first aspect of the present disclosure, the deduplication checking process is unified for multiple different types of affine candidates. Compared with conventional solutions, the proposed method can advantageously improve encoding and decoding efficiency.

[0006] In a second aspect, another method for video processing is provided, the method comprising: for conversion between a current video block of a video and a bitstream of the video, obtaining information about applying a fusion-based codec tool to the current video block, the information depending on a video type of the current video block; and performing conversion based on the information.

[0007] According to the method of the second aspect of the present disclosure, information about applying the fusion-based codec tool to the video block depends on the video type of the video block. Compared with conventional solutions, the proposed method can advantageously better support coding and decoding of different types of videos, thereby improving coding efficiency.

[0008] In a third aspect, another method for video processing is provided, the method comprising: for conversion between a current video block of a video and a bitstream of the video, selecting at least one Karhuning-Love transform (KLT) core from a plurality of KLT cores based on codec information of the current video block; and performing conversion based on the at least one KLT core.

[0009] According to the method of the third aspect of the present disclosure, the KLT kernel is selected by considering the coding information of the video block. Compared with conventional solutions, the proposed method can advantageously improve coding efficiency.

[0010] In a fourth aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0011] In a fifth aspect, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores instructions, which enable a processor to execute the method according to the first aspect of the present disclosure.

[0012] In a sixth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: obtaining multiple affine candidates for a current video block of the video; applying a deduplication check to the multiple affine candidates according to a check process, the check process being common to multiple different types of affine candidates; and generating a bitstream based on the application.

[0013] In a seventh aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining a plurality of affine candidates for a current video block of the video; applying a deduplication check to the plurality of affine candidates according to a check process, the check process being common to a plurality of different types of affine candidates; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] In an eighth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on a video type of the current video block; and generating a bitstream based on the information.

[0015] In a ninth aspect, a method for storing a bitstream of a video is provided, the method comprising: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on a video type of the current video block; generating a bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium.

[0016] In a tenth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: selecting at least one Karhunen-Loeve transform (KLT) core from a plurality of KLT cores based on codec information of a current video block of the video; and generating a bitstream based on the at least one KLT core.

[0017] In an eleventh aspect, a method for storing a bitstream of a video is provided, the method comprising: selecting at least one Karhuning-Love transform (KLT) core from a plurality of KLT cores based on codec information of a current video block of the video; generating a bitstream based on the at least one KLT core; and storing the bitstream in a non-transitory computer-readable recording medium.

[0018] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0020] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0021] Figure 2 A block diagram illustrating a first example video encoder is shown according to some embodiments of the present disclosure;

[0022] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0023] Figure 4 The location of the spatial Merge candidate is shown;

[0024] Figure 5 shows candidate pairs considered for redundancy check of spatial merge candidates;

[0025] Figure 6 The motion vector scaling for the temporal Merge candidate is shown;

[0026] Figure 7 The candidate position C for the time domain Merge candidate is shown 0 and C1 ;

[0027] Figure 8 The MMVD search point is shown;

[0028] Fig. 9 shows the extended CU area used in BDOF;

[0029] Fig.10 A symmetric MVD pattern is shown;

[0030] Fig.11 An affine motion model based on control points is shown;

[0031] Fig.12 The affine MVF of each sub-block is shown;

[0032] Fig.13 The positions of the inherited affine motion prediction values ​​are shown;

[0033] Fig.14 Control point motion vector inheritance is shown;

[0034] Fig.15 The positions of candidate positions for constructing the affine Merge pattern are shown;

[0035] Fig.16 is a diagram of the use of motion vectors for the proposed combination method;

[0036] Fig.17 Sub-block MV VSB and pixel Δv(i, j) are shown;

[0037] Fig.18A shows the spatial neighboring blocks used by ATVMP;

[0038] Fig.18B The sub-CU motion field is derived by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs;

[0039] Fig.19 shows the extended CU area used in BDOF;

[0040] Fig. 20 Decoding side motion vector refinement is shown;

[0041] Fig.21 The top and left neighbor blocks used in the derivation of CIIP weights are shown;

[0042] Fig. 22 An example of GPM partitions grouped at the same angle is shown;

[0043] Fig.23Unidirectional prediction MV selection for geometric partitioning mode is shown;

[0044] Fig.24 shows the warp weights w using the geometric segmentation mode 0 An exemplary generation of

[0045] Fig.25 The spatial neighboring blocks used to derive spatial Merge candidates are shown;

[0046] Fig.26 It shows that template matching is performed on the search area around the initial MV;

[0047] Fig. 27 A diamond-shaped area in the search area is shown;

[0048] Fig.28 shows the frequency response of the interpolation filter and the VVC interpolation filter at half-pixel phase;

[0049] Fig.29 The template and the reference sample points of the template in the reference picture are shown;

[0050] Fig.30 A template for a block with sub-block motion and reference samples of the template using motion information of a sub-block of a current block are shown;

[0051] Fig.31 Fill candidates for replacing zero vectors in the IBC list are shown.

[0052] Fig.32 shows the IBC reference area depending on the current CU position;

[0053] Fig.33 The reference area for IBC when CTU (m, n) is encoded and decoded is shown. The blue block represents the current CTU; the green block represents the reference area; and the white block represents the invalid reference area;

[0054] Fig.34 A first HPT and a second HPT are shown;

[0055] Fig.35 The spatial neighbors used to derive the affine Merge candidate / AMVP candidate are shown;

[0056] Fig.36 shows the constructed affine Merge candidate / AMVP candidate from non-adjacent neighbors to the first type;

[0057] Fig.37 The low frequency non-separable transform (LFNST) process is shown;

[0058] Fig.38The SBT position, type and transformation type are shown;

[0059] Fig.39 The ROI for LFNST 16 is shown;

[0060] Fig.40 The ROI for LFNST 8 is shown;

[0061] Fig.41 Discontinuity measurements are shown;

[0062] Fig.42 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;

[0063] Fig.43 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;

[0064] Fig.44 A flowchart showing a method for video processing according to an embodiment of the present disclosure; and

[0065] Fig.45 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.

[0066] Same or similar reference numbers generally refer to same or similar elements throughout the drawings. DETAILED DESCRIPTION

[0067] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0068] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0069] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.

[0070] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0071] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment

[0072] Figure 1 1 is a block diagram illustrating an example video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0073] The video source 112 may include a source such as a video acquisition device. Examples of a video acquisition device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0074] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a bit sequence that forms a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0075] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or may be outside the destination device 120, which is configured to be connected to an external display device interface.

[0076] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0077] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video encoder 114 in the system 100 is shown.

[0078] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0079] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0080] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.

[0081] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail below. Figure 2 are shown separately in the example.

[0082] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0083] The mode selection unit 203 may select one of a plurality of coding modes (intra-frame coding or inter-frame coding), for example, based on the error result, and provide the resulting intra-frame coded block or inter-frame coded deblock to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra-frame and inter-frame prediction (CIIP) mode in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select a resolution for a motion vector (e.g., sub-pixel precision or integer pixel precision) for the block.

[0084] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0085] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, a "P slice" and a "B slice" may refer to a portion of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0086] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures of list 0 or list 1 to find the reference video block for the current video block. The motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0087] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indexes indicating a plurality of reference pictures in list 0 and list 1 containing a plurality of reference video blocks and a plurality of motion vectors indicating a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block of the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0088] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process by a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal motion information of a current video block with reference to motion information of another video block. For example, motion estimation unit 204 may determine that motion information of a current video block is sufficiently similar to motion information of a neighboring video block.

[0089] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0090] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0091] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0092] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.

[0093] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of samples in the current video block.

[0094] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0095] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0096] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0097] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0098] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0099] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0100] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video decoder 124 in the system 100 is shown.

[0101] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of , video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0102] exist Figure 3 In the example of , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0103] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-encoded video data, the motion information including motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data and reference pictures from adjacent PBs. The motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction areas in B strips, also includes an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatial neighboring blocks or temporal neighboring blocks.

[0104] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter.Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0105] The motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 may use the interpolation filters to generate a prediction block.

[0106] The motion compensation unit 302 may use at least part of the syntax information to determine the size of blocks used to encode (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of a picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice may be an entire picture, or it may be a region of a picture.

[0107] The intra prediction unit 303 may use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0108] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.

[0109] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the section titles used in this document are for ease of understanding, and the embodiments disclosed in the section are not limited to that section. In addition, although some embodiments are described with reference to multifunctional video codecs or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video coding and decoding steps in detail, it should be understood that the corresponding decoding steps of de-coding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview Embodiments of the present disclosure relate to video coding technology. Specifically, it is about the coding technology of transformation, screen content coding and local illumination compensation in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, ECM, etc. It can also be applied to future video coding and decoding standards or video codecs. 2. Introduction Video codec standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards are based on a hybrid video codec structure, which utilizes temporal prediction plus transform codec. In order to explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held simultaneously every quarter, and in the April 2018 JVET meeting, the new video codec standard was officially named Versatile Video Codec (VVC), and the first version of the VVC Test Model (VTM) was released at this time. The VVC working draft and the test model VTM are then updated after each meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. In January 2021, JVET established Exploratory Experiments (EEs) to target enhanced compression efficiency beyond VVC capabilities using novel general algorithms. Shortly thereafter, ECM was built as a public software library for long-term exploratory work toward the next generation of video codec standards. 2.1. Existing inter-frame prediction codec tools For each inter-frame prediction CU, the motion parameters include motion vectors, reference picture indices and reference picture list usage indices, as well as additional information required by the new coding features of VVC that will be used for inter-frame prediction sample generation. Motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded using skip mode, the CU is associated with one PU and has no significant residual coefficients, no encoded and decoded motion vector increments or reference picture indices. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs (including spatial and temporal candidates, and additional lists introduced in VVC). Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other required information are explicitly transmitted through signals per CU. In addition to the inter-frame coding features in HEVC, VVC includes several new and refined inter-frame prediction coding tools listed below: – Extended Merge prediction; – Merge mode with MVD (MMVD); – Symmetric MVD (SMVD) signaling; – Affine motion compensated prediction; – Sub-block based temporal motion vector prediction (SbTMVP); – Adaptive Motion Vector Resolution (AMVR); – Motion field storage: 1 / 16 brightness sample MV storage and 8×8 motion field compression; – Bidirectional prediction with CU-level weights (BCW); – Bidirectional Optical Flow (BDOF); – Decoder-side motion vector refinement (DMVR); – Geometric Partitioning Mode (GPM); – Combined Inter and Intra Prediction (CIIP). The following text provides details about those inter prediction methods specified in VVC. 2.1.1. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Airspace MVP from airspace neighboring CUs; 2) Temporal MVP from the co-located CU; 3) History-based MVP from FIFO table; 4) Paired average MVP; 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU codec in Merge mode, the index of the best Merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the Merge index is coded using context, and bypass coding is used for other binary bits. The derivation process of each category of Merge candidates is provided in this session. As done in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a certain size. 2.1.1.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Figure 4 Select up to four Merge candidates from the candidates in the positions depicted in . The order of derivation is B 0、 A 0、 B 1、 A 1 and B 2 Only when position B 0 , A0 , B 1 , A 1 When one or more CUs of are not available (for example, because they belong to another slice or slice) or are intra-coded, position B 2 Considered. In position A 1 After the candidates at are added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, so that the encoding and decoding efficiency is improved. In order to reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs with Figure 5 The arrows in are linked to pairs, and a candidate is added to the list only if the corresponding candidate for redundancy check does not have the same motion information. 2.1.1.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, when deriving this time domain Merge candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for deriving the co-located CU is explicitly signaled in the slice header. Figure 6 As shown by the dotted line in , a scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set equal to zero. like Figure 7 As depicted in FIG. 1 , a position C for a temporal candidate is selected among the candidates. 0 With C 1 If position C 0 The CU at position C is unavailable, intra-coded, or outside the current row of the CTU, then 1 is used. Otherwise, position C 0 Used to derive time domain Merge candidates. Merge candidate derivation based on history After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previously coded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that at most 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if the same HMVP exists in the table. If found, the same HMVP is removed from the table, and all HMVP candidates are then moved forward. HMVP candidates can be used in the Merge candidate list construction process. Check the latest HMVP candidates in the table in order and insert them into the candidate list after the TMVP candidate. Apply redundancy check to the HMVP candidate to the spatial or temporal Merge candidate. To reduce the number of redundant checking operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N<=4)? M: (8- N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process from HMVP is terminated. 2.1.1.3. Pairwise Average Merge Candidate Derivation The pairwise average candidates are generated by averaging the predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the number represents the Merge index to the Merge candidate list. The average motion vector is calculated separately for each reference list. If both motion vectors are available in one list, the two motion vectors are averaged even when they point to different reference images; if only one motion vector is available, the one vector is used directly; if no motion vector is available, the list is kept invalid. When the merge list is not full after adding pairwise average merge candidates, zero MVPs are inserted at the end until the maximum number of merge candidates is encountered. 2.1.1.4.Merge estimated region The Merge Estimation Region (MER) allows the Merge candidate list for a CU to be independently derived in the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for the generation of the Merge candidate list for the current CU. In addition, the update process for the history-based motion vector prediction candidate list is updated only when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel) and where (xCb, yCb) is the upper left luminance sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and is signaled as log2_parallel_merge_level_minus2 in the sequence parameter set. 2.1.2. Merge Mode with MVD (MMVD) In addition to the Merge mode, a Merge mode with motion vector difference (MMVD) is introduced in VVC when implicitly derived motion information is directly used for prediction sample generation of the current CU. The MMVD flag is signaled immediately after the Skip flag and the Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after the Merge candidate is selected, it is further refined by the MVD information transmitted by the signal. Further information includes the Merge candidate flag, an index for specifying the motion size, and an index for indicating the motion direction. In MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The Merge candidate flag is transmitted by the signal to specify which one to use. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 8 As shown, the offset is added to the horizontal component or the vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1. Table 1 - Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 2. It should be noted that the meaning of the MVD symbol may vary depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV in which both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), Error! The symbol in the reference source was not found. Specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV with two MVs pointing to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the symbol in Table 2 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 2 - Sign of MV offsets specified by direction index Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.1. Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, a bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow a weighted average of the two prediction signals. P bi-pred =((8-w)*P 0 +w*P 1 +4)>>3 (2-1) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CU, the weight index is transmitted by signal after the motion vector difference; 2) For Merge CU, the weight index is inferred from the neighboring block based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights are only conditionally checked for 1-pixel and 4-pixel motion vector precision if the current picture is a low-latency picture. – When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using one context codec bit, followed by a bypass codec bit. The first context codec bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled using bypass codec to indicate the use of unequal weights. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficient encoding and decoding of video content in fading conditions. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in list L0 and list L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. In order to avoid interaction between WP and BCW (which will complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index of a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded and decoded using CIIP mode, the BCW index of the current CU is set to 2, for example, equal weight. 2.1.2.2. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, is included in JEM. BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers, compared to the JEM version. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is encoded and decoded using “true” bi-prediction mode, i.e., one of the two reference pictures precedes the current picture in display order, and the other reference picture follows the current picture in display order. – The distances from the two reference pictures to the current picture (ie, the POC difference) are the same. – Both reference images are short-term reference images. –CU is not encoded or decoded using Affine mode or ATMVP Merge mode. – The CU has more than 64 luma samples. – Both CU height and CU width are greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. –Do not enable WP for the current CU. –CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. x ,v y ). Motion refinement is then used to adjust the bidirectional prediction sample values ​​in the 4×4 sub-block. The following steps are applied in the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly calculating the difference between two adjacent samples. and k=0,1, that is, Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k (k=0, 1), and shift1 is calculated based on the luma bit depth (bitDepth) as shift1=max(6, bitDepth-6). Then, the gradient S 1 , S 2 , S 3 , S 5 and S 6 The autocorrelation and cross-correlation of are calculated as: in where Ω is a 6×6 window around a 4×4 sub-block, and n a and n b The values ​​of are set equal to min(1, bitDepth - 11) and min(4, bitDepth - 8). The cross-correlation and autocorrelation terms are then used to derive the motion refinement (v x ,v y): in th′ BIO =2 max(5,BD-7) , is the floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4×4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )>>shift (2-7) These values ​​are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits. In order to derive the gradient value, some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j) needs to be generated. Fig. 9 As depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. In order to control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended area (white positions) are generated by directly obtaining reference samples at nearby integer positions (using floor() operations on coordinates) without interpolation, and a normal 8-tap motion compensated interpolation filter is used to generate prediction samples within the CU (gray positions). These extended sample values ​​are only used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.1.2.3. Symmetric MVD Codec (SMVD) In VVC, in addition to normal uni-prediction and bi-prediction mode MVD signaling, a symmetric MVD mode for bi-prediction MVD signaling is applied. In symmetric MVD mode, motion information including reference picture indices of both list 0 and list 1 and MVD of list 1 is not signaled but derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. – Otherwise, if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a forward and backward reference picture pair or a backward and forward reference picture pair, BiDirPredFlag is set to 1, and both the list-0 reference picture and the list-1 reference picture are short-term reference pictures. Otherwise, BiDirPred- Flag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference index for list0 and list1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown below. Fig.10 Symmetrical MVD mode is shown. In the encoder, symmetric MVD motion estimation starts with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for symmetric MVD motion search. 2.1.3. Affine Motion Compensated Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Fig.11 As shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: Where (mv 0x ,mv 0y ) is the motion vector of the upper left control point, (mv 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point. In order to simplify motion compensation prediction, block-based affine transformation prediction is applied. In order to derive the motion vector of each 4×4 luminance sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation (such as Fig.12 The MV of the 4×4 chroma subblock is calculated as the average of the MVs of the upper left and lower right luminance subblocks in the same 8×8 luminance region. Like translational motion inter prediction, there are two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.1.3.1. Affine Merge Prediction AF_MERGE mode can be applied to CUs with width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatially adjacent CUs. There can be up to five CPMVP candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPVM candidates are used to form the affine Merge candidate list: – Inherited affine merge candidates inferred from the CPMV of neighboring CUs; – Affine Merge candidate CPMVP constructed using the translation MV of the neighboring CU; – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Fig.13 As shown. For the prediction value on the left, the scanning order is A0->A1, and for the prediction value on the top, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown Fig.30 As shown, if the adjacent lower left block A is encoded and decoded using the affine mode, the motion vectors v of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. 2 ,v 3 and v 4 When block A is encoded and decoded using a 4-parameter affine model, according to v 2 and v 3 Calculate the two CPMVs of the current CU. In the case where block A is encoded and decoded using a 6-parameter affine model, according to v 2 ,v 3 and v 4 Calculate the three CPMVs of the current CU. Fig.14 Control point motion vector inheritance is shown. The constructed affine candidate is constructed by combining the neighboring translation motion information of each control point. The motion information of the control point is obtained from Fig.15 The specified spatial and temporal neighbors shown in are derived, CPMV k (k=1,2,3,4) represents the kth control point. 1, the B2->B3->A2 blocks are checked and the MV of the first available block is used. For CPMV 2 , the B1->B0 block is checked, and for CPMV 3 , the A1->A0 block is checked. TMVP is used as CPMV 4 (if available). After getting the MVs of the four control points, construct the affine merge candidates based on those motion information. The following combinations of control point MVs are used to construct in order: {CPMV 1 ,CPMV 2 ,CPMV 3},{CPMV 1 ,CPMV 2 ,CPMV 4},{CPMV 1 ,CPMV 3 ,CPMV 4},{CPMV 2 ,CPMV 3 ,CPMV 4},{CPMV 1 ,CPMV 2},{CPMV 1 ,CPMV 3}. The combination of 3 CPMVs constructs a 6-parameter affine Merge candidate, and the combination of 2 CPMVs constructs a 4-parameter affine Merge candidate. To avoid the motion scaling process, the relevant combination of control point MVs is discarded if the reference indexes of the control points are different. After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not full, a zero MV is inserted at the end of the list. 2.1.3.2. Affine AMVP prediction Affine AMVP mode can be applied to CUs with width and height greater than or equal to 16. A CU-level affine flag is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by using the following four types of CPVM candidates in order: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; – Affine AMVP candidate CPMVP constructed using the translation MV of the neighboring CU; – Translated MV from neighboring CU. – Zero MV. The order in which inherited affine AMVP candidates are checked is the same as that of inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When an inherited affine motion prediction value is inserted into the candidate list, the deduplication process is not applied. The AMVP candidates constructed from Fig.15 The specified spatial neighbors are derived as shown. The same check order is used as in the affine merge candidate construction. In addition, the reference picture index of the neighboring blocks is checked. The first block in the check order is used, which is inter-coded and has the same reference picture as the current CU. There is only one. When the current CU is coded using the 4-parameter affine mode and mv 0 and mv 1 When all three CPMVs are available, they are added as a candidate in the affine AMVP list. When the current CU is encoded and decoded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list of candidates is still less than 2 after inserting the valid inherited affine AMVP candidates and the constructed AMVP candidates, then mv 0 ,mv 1 and mv 2 Will be added as translation MVs in order to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, the affine AMVP list is filled with zero MVs. 2.1.3.3. Affine motion information storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is only used to generate the inherited CPMV in affine Merge mode and the inherited CPMV in affine AMVP mode for the most recently encoded CU. The sub-block MV derived from the CPMV is used for motion compensation, MV derivation of the Merge / AMVP list of translation MVs, and deblocking. To avoid picture row cache for additional CPMV, the affine motion data inheritance from the CU above the CTU is handled differently from the inheritance from the normal neighboring CU. If the candidate CU for affine motion data inheritance is in the row above the CTU, the lower left sub-block MV and the lower right sub-block MV in the row cache are used for affine MVP derivation instead of CPMV. In this way, CPMV is only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to a 4-parameter model. Fig.16As shown, along the top boundary of the CTU, the bottom left sub-block motion vector and the bottom right sub-block motion vector of the CU are used for affine inheritance of the CU in the bottom of the CTU. 2.1.3.4. Prediction refinement using optical flow for affine mode (PROF) Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the expense of prediction accuracy. In order to achieve finer motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine sub-block-based affine motion compensation predictions without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the brightness prediction samples are refined by adding the differences derived from the optical flow equation. PROF is described as the following four steps: Step 1) Sub-block based affine motion compensation is performed to generate a sub-block prediction I(i,j). Step 2) Use a 3-tap filter [-1, 0, 1], the spatial gradient g of the sub-block prediction x (i,j) and g y (i,j) is calculated at each sample point. The gradient calculation is exactly the same as the gradient calculation in BDOF. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j))shift1) (2-11) g y (i,j)=(I(i,j+1))shift1)-(I(i,j-1))shift1) (2-12) Shift1 is used to control the accuracy of the gradient. The sub-block (i.e. 4×4) prediction is extended by one sample on each side of the gradient calculation. To avoid extra memory bandwidth and extra interpolation calculations, those extended samples on the extended boundary are copied from the nearest integer pixel position in the reference picture. Step 3) The brightness prediction refinement is calculated by following the optical flow equation. ΔI(i,j)= g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (2-13) Among them Fig.17 As shown, Δv(i,j) is the difference between the sample MV calculated for the sample position (i,j), denoted as v(i,j), and the sub-block MV of the sub-block to which the sample (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 luma sample accuracy. Since the affine model parameters and the sample position relative to the subblock center do not change from subblock to subblock, Δv(i,j) can be calculated for the first subblock and reused for other subblocks in the same CU. Let dx(i,j) and dy(i,j) be the distance from the sample position (i,j) to the subblock center (x SB ,y SB )’s horizontal and vertical offsets, Δv(x,y) can be derived by the following equations. To maintain accuracy, the sub-block (x SB ,y SB ) is calculated as ((W SB –1) / 2,(H SB –1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For the 4-parameter affine model, For the 6-parameter affine model, Where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the control point motion vectors of the upper left, upper right and lower left, and w and h are the width and height of the CU. Step 4) Finally, the luma prediction refinement ΔI(i,j) is added to the sub-block prediction I(i,j). The final prediction I' is generated as follows. I′(i,j)=I(i,j)+ΔI(i,j) PROF is not applicable to affine codec CUs in two cases: 1) all control point MVs are the same, which indicates that the CU has only translational motion; 2) the affine motion parameters are larger than the specified limit, because the sub-block based affine MC is downgraded to CU based MC to avoid large memory access bandwidth requirements. A fast codec method is applied to reduce the codec complexity of affine motion estimation using PROF. PROF is not applied to the affine motion estimation stage in the following two cases: a) If the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, PROF is not applied because the possibility that the current CU selects the affine mode as the best mode is low; b) If the sizes of the four affine parameters (C, D, E, F) are all less than a predefined threshold and the current picture is not a low-latency picture, PROF is not applied because the improvement introduced by PROF for this case is small. In this way, affine motion estimation using PROF can be accelerated. 2.1.4. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the CU in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: –TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; – While TMVP pre-fetches the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the center block relative to the current CU), SbTMVP applies motion displacement before pre-fetching the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in Figure 18. SbTMVP predicts the motion vector of a sub-CU within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 18(a) is checked. If A1 has a motion vector that uses the co-located picture as its reference picture, then that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the co-located picture as shown in Figure 18(b). The example in Figure 18(b) assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the minimum motion grid covering the center sample point) in the co-located picture is used to derive the motion information of the sub-CU. After the motion information of the co-located sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block-based Merge list containing both SbTMVP candidates and affine Merge candidates is used to signal the sub-block-based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block-based Merge candidates, followed by the affine Merge candidates. The size of the sub-block-based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8×8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8. The encoding and decoding logic of the additional SbTMVP Merge candidate is the same as that of other Merge candidates, that is, for each CU in a P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.1.5. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is transmitted by signal in units of quarter-luminance samples. In VVC, the CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU mode (normal AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows: – Normal AMVP mode: quarter brightness samples, half brightness samples, full brightness samples or four brightness samples. – Affine AMVP mode: quarter brightness samples, full brightness samples, or 1 / 16 brightness samples. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e. both horizontal and vertical MVD for reference list L0 and reference list L1) are zero, then quarter luma sample MVD resolution is inferred. For a CU with at least one non-zero MVD component, the first flag is signaled to indicate whether quarter luma sample MVD precision is used for the CU. If the first flag is 0, no further signal is required and quarter luma sample MVD precision is used for the current CU. Otherwise, the second flag is signaled to indicate that half luma samples or other MVD precision (integer or four luma samples) are used for normal AMVP CUs. In the case of half luma samples, a 6-tap interpolation filter is used for the half luma sample position instead of the default 8-tap interpolation filter. Otherwise, the third flag is signaled to indicate whether whole luma samples or four luma sample MVD precision is used for normal AMVP CUs. In the case of an affine AMVP CU, the second flag is used to indicate whether whole luma sample MVD precision or 1 / 16 luma sample MVD precision is used. In order to ensure that the reconstructed MV has the expected precision (quarter luma sample, half luma sample, whole luma sample or four luma samples), the motion vector prediction value of the CU will be rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction values ​​are rounded towards zero (that is, negative motion vector prediction values ​​are rounded towards positive infinity, and positive motion vector prediction values ​​are rounded towards negative infinity). The encoder uses RD check to determine the motion vector resolution of the current CU. In order to avoid always performing four CU-level RD checks for each MVD resolution, in VTM13, only the RD check of MVD precision other than quarter-luminance samples is conditionally called. For normal AVMP mode, the RD cost of quarter-luminance sample MVD precision and the RD cost of full-luminance sample MV precision are first calculated. Then, the RD cost of full-luminance sample MVD precision is compared with the RD cost of quarter-luminance sample MVD precision to determine whether it is necessary to further check the RD cost of four-luminance sample MVD precision. When the RD cost of quarter-luminance sample MVD precision is much smaller than the RD cost of full-luminance sample MVD precision, the RD check of four-luminance sample MVD precision is skipped. Then, if the RD cost of full-luminance sample MVD precision is significantly greater than the best RD cost of the previously tested MVD precision, the check of half-luminance sample MVD precision is skipped. For affine AMVP mode, if affine inter mode is not selected after checking the rate-distortion cost of affine Merge / Skip mode, Merge / Skip mode, quarter luma sample MVD accuracy normal AMVP mode, and quarter luma sample MVD accuracy affine AMVP mode, 1 / 16 luma sample MV accuracy and 1 pixel MV accuracy affine inter mode are not checked. In addition, in 1 / 16 luma sample and quarter luma sample MV accuracy affine inter mode, the affine parameters obtained in quarter luma sample MV accuracy affine inter mode are used as the starting search point. 2.1.6. Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, a bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow a weighted average of two prediction signals. P bi-pred =((8-w)*P 0 +w*P 1 +4)>>3 (2-18) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CU, the weight index is transmitted by signal after the motion vector difference; 2) For Merge CU, the weight index is inferred from the neighboring block based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-latency pictures, all 5 weights will be used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights for 1-pixel and 4-pixel motion vector precision are only conditionally checked if the current picture is a low-delay picture. - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using one context encoded bit followed by a bypass encoded bit. The first context encoded bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficient encoding and decoding of video content in fading conditions. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in list L0 and list L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. In order to avoid interaction between WP and BCW (which will complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not signaled, and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index of a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded and decoded using CIIP mode, the BCW index of the current CU is set to 2, for example, equal weight. 2.1.7. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, is included in JEM. BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers, compared to the JEM version. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is encoded and decoded using “true” bi-prediction mode, i.e., one of the two reference pictures precedes the current picture in display order, and the other reference picture follows the current picture in display order. – The distances from the two reference pictures to the current picture (ie, the POC difference) are the same. – Both reference images are short-term reference images. –Do not use affine mode or SbTMVP Merge mode to encode or decode the CU. – The CU has more than 64 luma samples. – Both CU height and CU width are greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. –Do not enable WP for the current CU. –CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. x ,v y ). Motion refinement is then used to adjust the bidirectional prediction sample values ​​in the 4×4 sub-block. The following steps are applied in the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly calculating the difference between two adjacent samples. and k=0,1, that is, Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k (k=0, 1), and shift1 is calculated based on the luma bit depth (bitDepth) as shift1=max(6, bitDepth-6). Then, the gradient S 1 , S 2 , S 3 , S 5 and S 6 The autocorrelation and cross-correlation of are calculated as: in where Ω is a 6×6 window around a 4×4 sub-block, and n a and n b The values ​​of are set equal to min(1, bitDepth - 11) and min(4, bitDepth - 8). The cross-correlation and autocorrelation terms are then used to derive the motion refinement (v x ,v y ): in th′ BIO =2 max(5,BD-7) , is the floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4×4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )>>shift (2-24) These values ​​are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits. In order to derive the gradient value, some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j) needs to be generated. Fig. 9 As depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. In order to control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended area (white positions) are generated by directly obtaining reference samples at nearby integer positions (using floor() operations on coordinates) without interpolation, and a normal 8-tap motion compensated interpolation filter is used to generate prediction samples within the CU (gray positions). These extended sample values ​​are only used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.1.8. Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of Merge mode MV, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In the bidirectional prediction operation, the refined MV is searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. Fig. 20 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs that are coded or decoded using the following modes and functions: – CU level Merge mode with bi-prediction MV. – Relative to the current picture, one reference picture is in the past and the other reference picture is in the future. – The distances from the two reference pictures to the current picture (i.e., POC difference) are the same. – Both reference images are short-term reference images. –CU has more than 64 luma samples. – CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. – The current block is not WP enabled. –CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction for future picture encoding and decoding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.1.8.1. Search scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) obeys the following two equations: MV0′=MV0+MV_offset (2-25) MV1′=MV1-MV_offset (2-26) Wherein, MV_offset represents the refinement offset between the initial MV and the refinement MV in one of the reference pictures. The refinement search range is two integer luminance samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. In order to reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV in the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation instead of an additional search using SAD comparison. Fractional sample refinement is conditionally called based on the output of the integer sample search stage. Fractional sample refinement is further applied when the integer sample search stage ends with the center with the minimum SAD in the first or second iteration search. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y)=A(x min ) 2 +B(yy min ) 2 +C (2-27) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values ​​of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-28) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-29) x min and min The value of is automatically clamped between -8 and 8, since all cost values ​​are positive, and the minimum value is E(0,0). This corresponds to a half-pixel shift with 1 / 16 pixel MV accuracy in VVC. The calculated score (x min ,y min ) is added to the integer distance refinement MV to obtain a sub-pixel accurate refinement delta MV. 2.1.8.2. Bilinear interpolation and sample filling In VVC, the resolution of MV is 1 / 16 luma sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so these fractional position samples need to be interpolated for the DMVR search process. In order to reduce the computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples than the normal motion compensation process within the 2 sample search range. After the refined MV is obtained through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. In order not to access more reference samples for the normal MC process, samples will be filled from those available samples, which are not required for the interpolation process based on the original MV, but are required for the interpolation process based on the refined MV. 2.1.8.3. Maximum DMVR processing unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16×16. 2.1.9. Combined Inter and Intra Prediction (CIIP) In VVC, when a CU is encoded and decoded using Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is transmitted by signal to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter The intra prediction signal P is derived using the same inter prediction process applied to the conventional Merge mode; and intra It is derived after the conventional intra prediction process with planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values ​​depend on the top neighboring block and the left neighboring block (in Fig.21 The codec mode described in the figure is calculated as follows: – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) is equal to 2, set wt to 3; – Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, set wt to 2; – Otherwise, set wt to 1. The CIIP forecast is formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (2-30) 2.1.10. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning modes are supported for inter prediction. The geometric partitioning mode is signaled using a CU level flag as a type of Merge mode, where other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. Among the total 64 partitions, for each possible CU size w×h=2 with the exclusion of 8×64 and 64×8 m ×2 n , m,n∈{3…6}, segmentation is supported through geometric segmentation mode. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Fig. 22 ). The position of the dividing line is mathematically derived from the angle parameters and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion compensated predictions are required for each CU, the same as regular bidirectional prediction. If the current CU uses geometric partitioning mode, a geometric partitioning index and two Merge indexes (one Merge index for each partition) indicating the partitioning mode of the geometric partitioning (angle and offset) are further transmitted by signal. The number of maximum GPM candidate sizes is explicitly transmitted by signal in the SPS and the syntax binarization of the GPM Merge index is specified. After predicting each of the parts of the geometric partition, a hybrid process with adaptive weights is used to adjust the sample values ​​along the edges of the geometric partition. This is the prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.10.1. One-way prediction candidate list construction The unidirectional prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector of the geometric partition mode. These motion vectors are Fig.23 In the case where there is no corresponding LX motion vector for the n-th extended Merge candidate, the L(1-X) motion vector of the same candidate is used instead of the unidirectional prediction motion vector for the geometric partition mode. 2.1.10.2. Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. The blending weight for each position of the CU is derived based on the distance between each position and the partition edge. The distance from position (x, y) to the segmentation edge is derived as: where i,j are the angle and offset indices of the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j Depends on the angle index. The weight of each part of the geometric segmentation is derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) w 1 (x,y)=1-w 0 (x,y) partIdx depends on the angle index i. The weight w 0 An example of Fig.24 Shown in. 2.1.10.3. Motion field storage for geometric partitioning mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1–partIdx):partIdx) where motionIdx is equal to d(4x+2,4y+2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType is equal to 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and another from L1), then simply combine Mv1 and Mv2 to form a bi-directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.1.11. Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique to model the local illumination variation between the current block and its prediction block as a function of the local illumination variation between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, which form a linear equation, i.e., α*p[x]+β to compensate for illumination variation, where p[x] is the reference sample pointed to by MV at position x on the reference picture. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for AMVP mode to indicate the use of LIC. Local illumination compensation is used for uni-directionally predicted inter CUs with the following modifications. Neighboring samples within the frame can be used to derive LIC parameters; Disable LIC for blocks with less than 32 luma samples; For both non-subblock and affine modes, LIC parameter derivation is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16×16 unit; Generate samples of the reference block template by using MC with the block MV without rounding it to integer pixel precision. 2.1.12. Non-adjacent airspace candidates Insert non-adjacent spatial merge candidates after TMVP in the regular merge candidate list. The mode of spatial merge candidates is Fig.25 The distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. No line buffer limit is applied. 2.1.13. Template Matching I Template matching I (TM) is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between the template in the current picture (i.e., the top neighboring block and / or the left neighboring block of the current CU) and the block in the reference picture (i.e., the same size as the template). Fig.26 As shown, a better MV is searched around the initial motion of the current CU in the [–8, +8] pixel search range. The template matching method is used for the following modifications: the search step size is determined based on the AMVR mode, and the TM can be cascaded with the bilateral matching process in the Merge mode. In AMVP mode, MVP candidates are determined based on template matching error to select the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM is performed only for this specific MVP candidate for MV refinement. TM refines the MVP candidate by using an iterative diamond search starting with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) within a search range of [–8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode), followed by searching sequentially by half-pixel and quarter-pixel depending on the AMVR mode as specified in Table 3. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 3. Search mode of AMVR and Merge mode with AMVR. In Merge mode, a similar search method is applied to the Merge candidate indicated by the Merge index. As shown in Table 3, according to the merged motion information, depending on the alternative interpolation filter (the interpolation filter used when AMVR is half-pixel mode), TM can be performed all the way up to 1 / 8 pixel MVD accuracy or skip the accuracy beyond half-pixel MVD accuracy. In addition, when TM mode is enabled, template matching can be used as an independent process or an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enabling condition check. 2.1.14. Multi-pass decoder-side motion vector refinement (mpDMVR) Multiple passes of decoder side motion vector refinement are applied. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16×16 sub-block within the codec block. In the third pass, the MVs in each 8×8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction. 2.1.14.1. First pass—Block-based bilateral matching MV refinement In the first pass, refined MVs are derived by applying BM to the codec blocks. Similar to decoder side motion vector refinement (DMVR), in bi-directional prediction operation, refined MVs are searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. Refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer point precision intDeltaMV. The local search applies a 3×3 square search pattern to loop in the horizontal search range [–sHor, sHor] and the vertical search range [–sVer, sVer], where the values ​​of sHor and sVer are determined by the block size and the maximum values ​​of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. When the bilCost of the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern and continues to search for the minimum cost until it reaches the end of the search range. The existing fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as: MV0_pass1 = MV0 + deltaMV; MV1_pass1=MV1-deltaMV. 2.1.14.2. Second pass—sub-block based bilateral matching MV refinement In the second pass, refined MVs are derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample point precision intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and a search range of [–sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block size and the maximum values ​​of sHor and sVer are 8. The bilateral matching cost is calculated by applying the cost factor to the SATD cost between the two reference sub-blocks, such as: bilCost = satdCost * costFactor. The search area (2*sHor+1)*(2*sVer+1) is divided into 5 diamond search areas, such as Fig. 27 As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order starting from the upper left corner of the area to the lower right corner. When the minimum bilCost in the current search area is less than a threshold equal to sbW*sbH, the integer pixel full search is terminated, otherwise, the integer pixel full search continues to the next search area until all search points are checked. The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). Then, the MV of the second pass refinement is derived as: ·MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2); ·MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2). 2.1.14.3. Third pass - sub-block-based bidirectional optical flow MV refinement In the third pass, refined MVs are derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied starting from the refined MV of the parent sub-block of the second pass to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32. The refined MVs of the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as: ·MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv; ·MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv. 2.1.15.OBMC When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information of neighboring blocks with weighted prediction. The conditions under which OBMC should not be applied are as follows: When OBMC is disabled at SPS level. When the current block has intra mode or IBC mode. When the current block applies LIC. When the current luminance block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top sub-block boundary pixels, left sub-block boundary pixels, bottom sub-block boundary pixels and right sub-block boundary pixels using neighboring sub-blocks. It is enabled for sub-block based codec tools: Affine AMVP mode; Affine Merge mode and sub-block based temporal motion vector prediction (SbTMVP); Sub-block based bilateral matching. 2.1.16. Sample-based BDOF In sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The codec block is divided into 8×8 sub-blocks. For each sub-block, it is determined whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, for each sample in the sub-block, a sliding 5×5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the center sample of the window. 2.1.17. Interpolation The 8-tap interpolation filter used in VVC is replaced by a 12-tap filter. The interpolation filter is derived from a sine function where the frequency response is truncated at the Nyquist frequency and clipped by a cosine window function. Table 4 gives the filter coefficients for all 16 phases. Fig.28 The frequency response of the interpolation filter is compared with the VVC interpolation filter, all at half-pixel phase. Table 4. Filter coefficients for 12-tap interpolation filter 2.1.18. Multiple Hypothesis Prediction (MHP) In the multi-hypothesis inter-frame prediction mode, in addition to the conventional bi-prediction signal, one or more additional motion compensated prediction signals are transmitted by signal. The resulting overall prediction signal is obtained by weighted superposition on a sample-by-sample basis. bi and the first additional inter-frame prediction signal / hypothesis h 3 , the predicted signal p 3 Obtain as follows. p 3 =(1-α)p bi +αh 3 The weighting factor α is specified by a new syntax element add_hyp_weight_idx according to the following mapping. add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 Similarly, more than one additional prediction signal may be used. The resulting overall prediction signal is iteratively accumulated using each additional prediction signal. p n+1 =(1-α n+1 ) n +α n+1 h n+1 The resulting overall prediction signal is obtained as the final p n (ie, with the largest index). Within this EE, up to two additional prediction signals may be used (ie, n is limited to 2). The motion parameters for each additional prediction hypothesis can be signaled explicitly by specifying the reference index, motion vector predictor index and motion vector difference or implicitly by specifying the Merge index. A separate Multi-Hypothesis Merge flag distinguishes these two signaling modes. For inter-AMVP mode, MHP is only applied if unequal weights in BCW are selected in bi-prediction mode. A combination of MHP and BDOF is possible, however BDOF is only applied to the bidirectional prediction signal part of the prediction signal (ie the ordinary first two hypotheses). 2.1.19. Adaptive Reordering of Merge Candidates Using Template Matching (ARMC-TM) Merge candidates are adaptively reordered using Template Matching (TM). The reordering method is applied to the regular Merge mode, Template Matching (TM) Merge mode, and Affine Merge mode (excluding SbTMVP candidates). For TM Merge mode, Merge candidates are reordered before the refinement process. After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5 for the regular Merge mode and TMMerge mode. The subgroup size is set to 3 for the affine Merge mode. The Merge candidates in each subgroup are reordered in ascending order according to the cost value based on template matching. For simplicity, the Merge candidates in the last subgroup instead of the first subgroup are not reordered. The template matching cost of the Merge candidate is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate. When the Merge candidate uses bidirectional prediction, the reference samples of the template of the Merge candidate are also as follows: Fig.29 The bidirectional prediction shown in is generated. For a sub-block-based Merge candidate with a sub-block size equal to Wsub×Hsub, the above template includes several sub-templates with a size of Wsub×1, and the left template includes several sub-templates with a size of 1×Hsub. Fig.30 As shown, the reference sample of each sub-template is derived using the motion information of the sub-blocks in the first row and the first column of the current block. 2.1.20. Geometric Partitioning Mode (GPM) with Merged Motion Vector Difference (MMVD) GPM in VVC is extended by applying motion vector refinement on top of the existing GPM unidirectional MV. A flag for the GPM CU is first signaled to specify whether this mode is used. If the mode is used, each geometric partition of the GPM CU can further decide whether to signal MVD. If MVD is signaled for the geometric partition, the motion of the partition is further refined by the signaled MVD information after selecting the GPMMeter candidate. All other processes remain the same as GPM. The MVD is signaled as a pair of distance and direction, similar to MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) related to MMVD in GPM (GPM-MMVD). In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD is left-shifted by 2 in MMVD. 2.1.21. Geometric Partitioning Model (GPM) using Template Matching (TM) Template matching is applied to GPM. When GPM mode is enabled for a CU, a CU level flag is signaled to indicate whether TM is applied to both geometric partitions. TM is used to refine the motion information of each geometric partition. When TM is selected, the template is constructed using the left, top, or left and top neighboring samples depending on the partition angle, as shown in Table 5. The motion is then refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern of Merge mode with disabled half-pixel interpolation filter. Table 5 is a template for the first geometric partition and the second geometric partition, where A indicates using the upper sample, L indicates using the left sample, and L+A indicates using both the left sample and the upper sample. Split Angle 0 2 3 4 5 8 11 12 13 14 First division A A A A L+A L+A L+A L+A A A Second division L+A L+A L+A L L L L L+A L+A L+A Split Angle 16 18 19 20 21 24 27 28 29 30 First division A A A A L+A L+A L+A L+A A A Second division L+A L+A L+A L L L L L+A L+A L+A The GPM candidate list is constructed as follows: 1. Directly derive interleaved List 0 MV candidates and List 1 MV candidates from the regular Merge candidate list, where List 0 MV candidates have higher priority than List 1 MV candidates. Apply a deduplication method with adaptive threshold based on the current CU size to remove redundant MV candidates. 2. Interleaved List 1 MV candidates and List 0 MV candidates are further derived directly from the regular Merge candidate list, where List 1 MV candidates have higher priority than List 0 MV candidates. The same deduplication method with adaptive threshold is also applied to remove redundant MV candidates. 3. Zero MV candidates are filled until the GPM candidate list is full. GPM-MMVD and GPM-TM are enabled specifically for one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), a GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is presumed to be false. 2.1.22. GPM with inter-frame and intra-frame prediction (GPM inter-intra) With GPM inter-intra, in addition to the Merge candidates for each non-rectangular partitioned area in the CU to which GPM is applied, a predetermined intra prediction mode for the geometric partition line can also be selected. In the proposed method, whether the intra prediction mode or the inter prediction mode is determined for each GPM separable area with a flag from the encoder. When in inter prediction mode, a unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, when in intra prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra prediction mode specified by the index from the encoder. The variation of possible intra prediction modes is limited by the geometric shape. Finally, the two unidirectional prediction signals are mixed in the same way as ordinary GPM. 2.1.23. Adaptive Decoder-side Motion Vector Refinement (Adaptive DMVR) The adaptive decoder side motion vector refinement method consists of two new Merge modes, which are introduced to refine the MV in only one direction (L0 or L1) of the bidirectional prediction of the Merge candidate that meets the DMVR condition. A multi-pass DMVR process is applied to the selected Merge candidate to refine the motion vector, however, MVD0 or MVD1 is zero in the first pass (i.e., PU level) DMVR. Similar to the conventional Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially adjacent coded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list is used by both proposed Merge modes, and the Merge index is encoded and decoded in the conventional Merge mode. 2.1.24. Bilateral Matching AMVP-MERGE Mode (AMVP-MERGE) In AMVP Merge mode, the bidirectional prediction value consists of the AMVP prediction value in one direction and the Merge prediction value in the other direction. The AMVP part of the proposed mode is signaled as regular unidirectional AMVP, i.e., the reference index and MVD are signaled, and it has a derived MVP index (TM_AMVP) if template matching is used, or the MVP index is signaled when template matching is disabled. The Merge index is not signaled, and the Merge prediction value is selected from the candidate list with the minimum template or bilateral matching cost. When the selected Merge prediction value and AMVP prediction value meet the DMVR condition (which is at least one reference picture from the past and one reference picture from the future relative to the current picture) and the distance from the two reference pictures to the current picture is the same, bilateral matching MV refinement is applied to the Merge MV candidate and AMVP MVP as a starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the Merge prediction value or AMVP prediction value with a higher template matching cost. The third pass of 8×8 sub-PU BDOF refinement as multi-pass DMVR is enabled as AMVP Merge mode codec block. 2.1.25.IBC Merge / AMVP List Construction IBC Merge / AMVP list construction is modified as follows: An IBC Merge / AMVP candidate can be inserted into the IBC Merge / AMVP candidate list only if it is valid. The upper right spatial domain candidate, the lower left spatial domain candidate, the upper left spatial domain candidate and a pairwise average candidate may be added to the IBC Merge / AMVP candidate list. Adaptive Reordering Based on Template (ARMC-TM) is applied to the IBC Merge list. The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC Merge candidates with full deduplication, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list. Candidates for zero vectors used to populate the IBC Merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in IBC Merge mode, and therefore, are discarded as BVPs in the IBC candidate list. Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Fig.31 Depicted in. 2.1.26. IBC using template matching Template matching is used in IBC for both IBC Merge mode and IBC AMVP mode. The IBC-TM Merge list is modified compared to the list used by the conventional IBC Merge mode so that candidates are selected according to the deduplication method using the motion distance between candidates as in the conventional TM Merge mode. The end zero motion fulfillment is replaced by the motion vectors of the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU. In IBC-TM Merge mode, template matching methods are utilized to refine the selected candidates before the RDO or decoding process.IBC-TM Merge mode has been made competitive with the regular IBC Merge mode and the TM-Merge flag is signaled. In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM Merge list. Each of these 3 selected candidates is refined using a template matching method and ranked according to their resulting template matching cost. Only the first two candidates are then typically considered in the motion estimation process. Template matching refinement for both IBC-TM Merge mode and AMVP mode is very simple, since IBC motion vectors are constrained (i) to be integers and (ii) to be Fig.32 . Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, they are performed with integer or 4-pixel precision depending on the AMVR value. Such refinements are only accessed for samples that are not interpolated. In both cases, the refined motion vectors and the template used in each refinement step must respect the constraints of the reference region. 2.1.27.IBC reference area The reference area of ​​the IBC extends to the upper two CTU rows. Fig.33 The reference area for encoding and decoding CTU (m, n) is shown. Specifically, for the CTU (m, n) to be encoded and decoded, the reference area includes CTUs with indices (m-2, n-2)...(W, n-2), (0, n-1), (W, n-1), (0, n)...(m, n), where W represents the maximum horizontal index within the current slice, stripe or picture. This setting ensures that for a CTU size of 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is horizontally limited to [-(C<<1), C>>2] and vertically limited to [-C, C>>2] to accommodate reference area expansion, where C represents the CTU size. 2.1.28.MVD symbol prediction In this method, possible MVD symbol combinations are sorted according to the template matching cost, and the index corresponding to the real MVD symbol is derived and context-coded. At the decoder side, the MVD symbol is derived as follows: 1. Analyze the size of the MVD components; 2. Parse the MVD symbol prediction index for context encoding and decoding; 3. Construct MV candidates by creating combinations between possible signs and absolute MVD values ​​and adding them to the MV prediction values; 4. Derive the MVD symbol prediction cost for each derived MV based on the template matching cost and ranking; 5. Use the MVD symbol prediction index to pick the true MVD symbol. MVD sign prediction is applied to inter-AMVP mode, affine AMVP mode, MMVD mode, and affine MMVD mode. 2.1.29. Enhanced Bidirectional Motion Compensation In bidirectional motion compensation, out-of-boundary (OOB) prediction samples are discarded and only non-OOB prediction values ​​are used to generate the final prediction value. Specifically, assuming Pos_x i,j and Pos_y i,j Indicates the position of a predicted sample point in a current block, and (x=0,1) represents the MV of the current block; Pos LeftBdry 、Pos RightBdry 、Pos TopBdry and Pos BottomBdry are the positions of the four boundaries of the picture. A prediction sample is considered OOB when at least one of the following conditions is met: Where half_pixel is equal to 8, which represents the half-pixel sample distance in 1 / 16 pixel sample accuracy. After checking the OOB condition for each sample, the final prediction sample for a bidirectional block is generated as follows: if is OOB and Yes or No OOB Otherwise, if is non-OOB and It's OOB. otherwise When BCW is enabled, the OOB checking process also applies. 2.1.30. Block-level reference picture list reordering A block-level reference picture reordering method based on template matching is used. For uni-predictive AMVP mode, the reference pictures in list 0 and list 1 are interleaved to generate a joint list. For each hypothesis of the reference pictures in the joint list, a match is performed to calculate the cost. The joint list is reordered based on ascending order of template matching cost. The index of the selected reference picture in the reordered joint list is signaled in the bitstream. For bi-predictive AMVP mode, a list of reference picture pairs from list 0 and list 1 is generated and similarly reordered based on template matching cost. The index of the selected pair is signaled. 2.1.31. Affine model inheritance and non-adjacent affine patterns based on history parameters History Parameter-Based Affine Model Inheritance (HAMI) allows affine models to be inherited from previously affine-encoded blocks that may not be adjacent to the current block. Similar to the enhanced regular Merge mode, a non-adjacent affine mode (NA-AFF) is introduced. A first history parameter table (HPT) is established. The entries of the first HPT store affine parameter sets: a, b, c, and d, each of which is represented by a 16-bit signed integer. The entries in the HPT are classified by reference list and reference index. Five reference indexes are supported for each reference list in the HPT. In a formal way, the category of the HPT (denoted as HPTCat) is calculated as HPTCat(RefList, RefIdx)=5×RefList+min(RefIdx, 4), Where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index, respectively. For each category, up to seven entries can be stored, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to zero. cur and RefIdx cur After the affine codec is decoded, the affine parameters are used to update the category HPTCat (RefList cur ,RefIdx cur ) in the . The candidate based on historical affine parameters (HAPC) is obtained from Figure 4 The MV of the neighboring 4×4 block is used as the base MV. In a formal way, the MV of the current block at position (x, y) is calculated as: Where (mv h base ,mv v base ) represents the MV of the adjacent 4×4 block, (x base ,y base ) represents the center position of the neighboring 4×4 blocks. (x, y) can be the top left, top right, and bottom left corners of the current block to obtain the angular position MV (CPMV) of the current block, or it can be the center of the current block to obtain the normal MV of the current block. A second history parameter table (HPT) with base MV information is also appended. There are nine entries in the second HPT, where the entry includes the base MV, the reference index and four affine parameters for each reference list, and the base position. An additional Merge HAPC can be generated from the second HPT with base MV information, and the corresponding affine model is stored in the entry. The difference between the first HPT and the second HPT is Fig.34 Shown in. In addition, a pair of affine merge candidates is generated from two affine merge candidates, which are either historically derived or non-historically derived. The pair of affine merge candidates is generated by averaging the CPMVs of the existing affine merge candidates in the list. In response to the introduction of the new HAPC, the size of the sub-block based Merge candidate list was increased from 5 to 15, both of which involve the ARMC

[10] process. In NA-AFF, the mode of obtaining non-adjacent spatial neighbors is Figure 6 As shown in , the distance between the non-adjacent spatial neighbors and the current codec block in NA-AFF is also defined based on the width and height of the current CU, similar to the existing non-adjacent conventional Merge candidates [8]. use Figure 6 The motion information of the non-adjacent spatial neighbors in VVC is used to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for the inherited candidates, the same derivation process of the inherited affine Merge / AMVP candidates in VVC remains unchanged, except that CPMV is inherited from non-adjacent spatial neighbors. Non-adjacent spatial neighbors are checked based on their distance from the current block (i.e., from near to far). At a certain distance, only the first available neighbor (encoded using affine mode) from each side of the current block (e.g., left and above) is included for inherited candidate derivation. Fig.35 The spatial neighbors used to derive the affine Merge / AMVP candidates are shown. In addition, Fig.35 Sub-image (a) of shows the spatial neighbors used to derive the inherited candidate, and Fig.35Sub-image (b) of shows the spatial neighbors used to derive the first type of construction candidates. Fig.35 As indicated by the dotted arrows in sub-image (a), the order of checking the neighbors on the left and above is from bottom to top and from right to left, respectively. For the first type of build candidates, such as Fig.35 As shown in the sub-image (b) of , the position of a non-adjacent spatial neighbor on the left and above is first determined independently; then, the position of the upper left neighbor can be determined accordingly, which can close the rectangular virtual block with the non-adjacent neighbors on the left and above. Then, as Fig.36 As shown, motion information of three non-adjacent neighbors is used to form CPMVs at the upper left (A), upper right (B), and lower left (C) of the virtual block, which are finally projected to the current CU to generate corresponding construction candidates. Insert the NA-AFF candidates into the existing affine merge candidate list and the affine AMVP candidate list according to the following order: Affine Merge mode: 1.SbTMVP candidate, if available. 2. Inherit from adjacent neighbors. 3. Inherit from non-adjacent neighbors. 4. Build from adjacent neighbors. 5. Construct affine candidates of the first type from non-contiguous neighbors. 6. Zero MV. Affine AMVP mode: 1. Inherit from adjacent neighbors. 2. Build from adjacent neighbors. 3. Translational MVs from adjacent neighbors. 4. Translational MV from temporal neighbors. 5. Inherit from non-adjacent neighbors. 6. Construct affine candidates of the first type from non-contiguous neighbors. 7. Zero MV. Due to the inclusion of additional candidates generated by NA-AFF, the size of the affine merge candidate list increases from 5 to 15. The subgroup size of ARMC for affine merge mode increases from 3 to 15. In NA-AFF: 1. Regions from non-adjacent neighbors are confined to the current CTU (ie, no additional storage requirement for line buffers). 2. The storage granularity of affine motion information including CPMV and reference index is reduced from 8×8 to 16×16 (That is, only the affine motion from the top left 8x8 block is saved.) In addition, the saved CPMV is projected to each 16x16 block before storage, so that position and size information is not required. 3. Only the upper left CPMV and the upper right CPMV are stored (ie, the 4-parameter affine model for NA-AFF is always used). 2.1.32. Regression-based affine candidate inference method A regression-based affine candidate derivation method is proposed. The sub-block motion field from the previously encoded affine CU and the motion vectors from the neighboring sub-blocks of the current CU are used as input to the regression process. The predicted CPMV instead of the sub-block motion field of the current block is derived as the output. The derived CPMV can be added to the sub-block Merge candidate list or the affine AMVP list. The scanning pattern of the previously encoded affine CU is the same as the non-adjacent scanning pattern used in the conventional Merge candidate list construction. 2.2. Transform and Coefficient Encoding and Decoding 2.2.1. Large Block Size Transformation with High Frequency Zeroing In VVC, large block size transforms of up to 64×64 are enabled, which are mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are cleared so that only lower-frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M is equal to 64, only the left 32 columns of transform coefficients are retained. Similarly, when N is equal to 64, only the first 32 rows of transform coefficients are retained. When the transform skip mode is used for large blocks, the entire block is used without clearing any values. In addition, the transform displacement is removed in the transform skip mode. VTM also supports a configurable maximum transform size in SPS, giving the encoder the flexibility to select up to 32 lengths or 64 lengths of transform sizes according to the needs of a specific implementation. 2.2.2. Multiple Transformation Selection (MTS) for Kernel Transformations In addition to DCT-II already adopted in HEVC, the multi-transform selection (MTS) scheme is also used for residual coding of blocks coded and decoded inter-frame and intra-frame. It uses multiple selected transforms in DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the basis functions of the selected DST / DCT. Table 2-6 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is ​​more accurate than that in HEVC. To keep the intermediate values ​​of the transform coefficients within 16 bits, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at the SPS, a CU level flag is signaled to indicate whether MTS is applied. Here, MTS applies only to luma. MTS signaling is skipped when one of the following conditions applies: – The position of the last significant coefficient of the luma TB is less than 1 (ie DC only). – The last significant coefficient of the luminance TB is located within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform types in the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 2-7. By eliminating intra-mode and block shape dependencies, a unified transform selection for ISP and implicit MTS is used. If the current block is in ISP mode or if the current block is an intra-block and both intra- and inter-frame explicit MTS are turned on, only DST7 is used for horizontal and vertical transform kernels. In terms of transform matrix accuracy, an 8-bit main transform kernel is used. Therefore, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8) all use 8-bit main transform kernels. Table 2-7 - Transformation and signaling mapping table To reduce the complexity of large-sized DST-7 and DCT-8, high-frequency transform coefficients are zeroed for DST-7 and DCT-8 blocks with size (width or height, or both) equal to 32. Only coefficients in the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, implicit MTS can still be enabled. 2.2.3. Low-frequency non-separable transform (LFNST) In VVC, Fig.37As shown, LFNST is applied between the forward main transform and quantization (at the encoder) and between dequantization and the inverse main transform (at the decoder side). In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). The following uses the input as an example to describe the application of the non-separable transform used in LFNST. To apply 4×4 LFNST, the 4×4 input block X is first represented as a vector The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. Subsequently, the 16×1 coefficient vector is reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) for the block. Coefficients with smaller indices will be placed in the 4×4 coefficient block together with smaller scan indices. 2.2.3.1. Reduced non-separable transform LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix size to minimize the computational complexity and the memory space to store the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (where N is usually equal to 64 for 8×8 NSST) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows: Wherein the R rows of the transform are the R basis of the N-dimensional space. The inverse transform matrix for RT is the transpose of its forward transform. For 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix (which is the conventional 8×8 non-separable transform matrix size) is reduced to a 16×48 direct matrix. Therefore, a 48×16 inverse RST matrix is ​​used on the decoder side to generate the core (main) transform coefficients in the 8×8 upper left region. When a 16×48 matrix is ​​applied instead of a 16×64 with the same transform set configuration, each of which takes 48 input data from three 4×4 blocks in the upper left 8×8 block except the lower right 4×4 block. With the reduced size, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, with a reasonable performance degradation. In order to reduce complexity, it is applicable to limit LFNST only when all coefficients outside the first coefficient subgroup are not significant. Therefore, when LFNST is applied, all main transform coefficients must be zero. This allows LFNST index signaling to be adjusted on the last significant position and thus avoids extra coefficient scans in current LFNST designs, which require checking significant coefficients only at specific positions. The worst-case processing of LFNST (in terms of multiplications per pixel) limits the inseparable transforms for 4×4 blocks and 8×8 blocks to 8×16 transforms and 8×48 transforms, respectively. In these cases, the last significant scan position must be less than 8 when LFNST is applied, for other sizes less than 16. For blocks with shapes of 4×N and N×4 and N>8, the proposed restrictions mean that LFNST is now applied only once and only to the upper left 4×4 region. Since all the main coefficients are zero when LFNST is applied, the number of operations required for the main transform is reduced in this case. From the encoder's perspective, the quantization of coefficients is significantly simplified when testing the LFNST transform. Rate-distortion optimized quantization must be maximally done for the first 16 coefficients (in scan order), forcing the remaining coefficients to zero. 2.2.3.2. LFNST Transform Selection There are 4 transform sets and 2 inseparable transform matrices (kernels) used in each transform set in LFNST. As shown in Table 2-8, the mapping from intra prediction mode to transform set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by the LFNST index of explicit signaling. The index is transmitted by signal once in the bitstream per intra CU after the transform coefficients. Table 2-8 - Transformation selection table IntraPredMode Transform Set Index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0 2.2.3.3.LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-encoded, but does not depend on the intra prediction mode, and only the first binary bit is context-encoded. In addition, LFNST is applied to intra CUs on both intra and inter slices, and on both luma and chroma. If dual tree is enabled, the LFNST indexes for luma and chroma are signaled separately. For inter slices (dual tree is disabled), a single LFNST index is signaled and used for luma and chroma. Considering that large CUs larger than 64×64 are implicitly partitioned (TU slicing) due to the existing maximum transform size limit (64×64), LFNST index search can increase data buffering by four times for a certain number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST is limited to 64×64. Note that LFNST is only enabled for DCT2. LFNST index signaling is placed before MTS index signaling. Using the scaling matrix for perceptual quantization is not obvious, and the scaling matrix specified for the main matrix can be used for the LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied. 2.2.4. Sub-Block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When cu_cbf is equal to 1 for an inter-predicted CU, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded and decoded. For the former case, the inter-MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by inferred adaptive transform, while another part of the residual block is cleared. When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, such as Fig.38As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is similar to the binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is similar to the asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 (in units of luminance samples), 1:3 / 3:1 partitioning along this dimension is not allowed. A CU has a maximum of 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). Different kernel transforms are associated with the two positions, SBT-H and SBT-V. More specifically, the horizontal and vertical transforms for each SBT position are selected in Fig.38 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slice, cbf, and horizontal and vertical kernel transform types for the residual block. SBT is not applied to CUs coded with combined inter-intra mode. 2.2.5. Maximum transform size and clearing of transform coefficients Both the CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the largest intra-frame codec block can have a size of 128×128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, there is no standardized zeroing operation applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are standardizedly zeroed. 2.2.6. Enhanced MTS for intra-frame coding and decoding In the current VVC design [1], for MTS, only DST7 transform core and DCT8 transform core are utilized, which are used for intra-frame and inter-frame coding and decoding. Additional main transforms including DCT5, DST4, DST1 and identity transform (IDT) are used. The MTS set also depends on the TU size and intra mode information. 16 different TU sizes are considered, and for each TU size 5, different categories are considered according to the intra mode information. For each category, 1, 4 or 6 different transform pairs are considered. Based on the absolute value sum of the transform coefficients, multiple intra MTS candidates (between 1, 4 and 6 MTS candidates) are adaptively selected. The sum is compared with two fixed thresholds to determine the total number of allowed MTS candidates: 1 candidate: sum <= th0. 4 candidates: th0<sum<=th1. 6 candidates: sum > th1. Note that although a total of 80 different classes are considered, some of these different classes often share exactly the same set of transforms. Hence there are 58 (less than 80) unique entries in the resulting LUT. For angle modes, the joint symmetry on TU shape and intra prediction is considered. Therefore, mode i (i>34) with TU shape A×B will be mapped to the same category corresponding to mode j=(68-i) with TU shape B×A. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For wide-angle mode, the closest conventional angle mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. 2.2.7. Secondary Transformation: Extension of LFNST with Large Kernel The LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is extended to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: oFor predModeIntra<2, lfnstTrSetIdx is equal to 2; o lfnstTrSetIdx=predModeIntra for predModeIntra in [0,34]; o lfnstTrSetIdx=68-predModeIntra, for predModeIntra in [35,66]. Three different kernels LFNST 4, LFNST 8, and LFNST 16 are defined to indicate LFNST kernel sets applied to 4×N / N×4 (N≥4), 8×N / N×8 (N≥8), and M×N (M, N≥16), respectively. The kernel size is specified by: (LFSNT4, LLFNST8*, LFNST16*) = (16×16, 32×64, 32×96) Forward LFNST is applied to the upper left low frequency region called the Region of Interest (ROI).When LFNST is applied, the main transform coefficients present in the region except the ROI are zeroed, which is not changed from the VVC standard. The ROI of LFNST16 is Fig.39 It consists of six 4×4 sub-blocks, which are consecutive in scanning order. Since the number of input samples is 96, the transform matrix for forward LFNST16 can be R×96. In this contribution, R is chosen to be 32, and accordingly 32 coefficients (two 4×4 sub-blocks) are generated from forward LFNST16. block), which is placed after the coefficient scanning order. The ROI of LFNST8 is Fig.40 The forward LFNST8 matrix may be R×64, and R is chosen to be 32. The generated coefficients are positioned in the same way as for LFNST 16. The mapping from intra prediction modes to these sets is shown in Table 9, Table 9. Mapping of intra prediction modes to LFNST set indices Intra prediction mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST set index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 l 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST set index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra prediction mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNsT Set Index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.2.8. Symbol Prediction The basic idea of ​​the coefficient sign prediction method is to compute the reconstruction residuals of negative and positive sign combinations for the applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal symbol, the cost function is defined as spanning Fig.41 The discontinuity measures at the block boundaries are shown above. All hypotheses are measured and the one with the smallest cost is chosen as the predicted value for the coefficient sign. The cost function is defined as the sum of the absolute second-order derivatives in the residual domain of the upper rows and left columns as follows: Where R is the reconstructed neighbor, P is the prediction for the current block, and r is the residual hypothesis. We can compute only the term (-R- 1 +2R 0 -P1) once, and only the residual assumption is subtracted. The transform coefficient with the maximum KqIdx value of the upper left 4×4 region is selected. The qIdx value is the transform system level after compensating for the effects from multiple quantizers in DQ. Larger qIdx values ​​will produce larger dequantized transform coefficient levels. qIdx is derived as follows: qIdx=(abs(level)<<1)-(state&1); where level is the transform coefficient level parsed from the bitstream, and state is a variable maintained by the encoder and decoder in DQ. The symbol prediction area is expanded to a maximum value of 32×32. The symbol of the upper left M×N block is predicted. The values ​​of M and N are calculated as follows: οM=min(w,maxW) οN=min(h,maxH) Where w and h are the width and height of the transform block. The maximum region for symbol prediction is not always set to 32×32. The encoder sets the maximum region (maxW, maxH) based on the configuration, sequence class, and QP, and signals the region in the SPS. The maximum number of prediction symbols remains unchanged. Symbol prediction is also applied to LFNST blocks. And for LFNST blocks, symbol prediction is allowed for the 4 largest coefficients in the top left 4×4 region. 3. Questions / issues There are several problems with existing video coding techniques, which can be further improved for higher coding gain. 1. In ECM-6.0, affine candidates can be derived from neighbor-based affine candidates, history-based affine candidates, non-neighbor-based affine candidates, and regression-based affine candidates. Similarity checks are performed on affine candidate derivation. However, difference similarity check rules are used for affine candidate derivation. This may not be optimal. 2. In ECM-6.0, hybrid modes such as CIIP, OBMC are applied to both the video captured by the camera and the screen content video, which may not be efficient. 3. In ECM-6.0, KLT is allowed for explicit inter-MTS mode. Specifically, if the TU of the inter-coded codec is less than or equal to 16×16, two KLT options (i.e., KLT0 and KLT1) of the inter-MTS core are used to replace DST7 and DCT8. This design can be changed for higher codec efficiency. 4. In ECM-6.0, KLT is allowed to be used in inter-frame MTS mode, but the use of KLT does not depend on which inter-frame prediction technology is used for the video unit, which can be further improved. 4. Embodiments of the present disclosure The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), CU, PU, ​​TU, PB, TB. The term "KLT" may refer to a type of transform. For example, it may refer to the Karl-Hunin-Löf transform. For example, it may refer to any transform type that is not DCT or DST or Hadamard. The coefficient matrix associated with a particular KLT may be trained online or predefined (e.g., trained offline) based on some a priori knowledge (e.g., residuals / coefficients from already decoded neighboring blocks). In the present disclosure, regarding “blocks encoded using mode N”, “mode N” here can be a specific prediction mode (for example, MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a specific prediction technology (for example, AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, Affine, CIIP, GPM, GPM intra, MHP, OBMC, LIC, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, sub-block coding, hypothesis coding, etc.), or a specific transform process (IDTX, DCT-X, DST-Y, KLT-Z, where X / Y / Z are constants) or a specific filter process (deblocking, SAO, bilateral filter, adaptive loop filter, CCSAO, CC-ALF, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes in codec tools are also applicable. 4.1. Regarding the first problem of the derivation of affine candidates, the following method is proposed: a. The same logic / rules / process of similarity / consistency / deduplication checking can be used for the derivation of all affine candidates. a. For example, it can refer to the derivation of an affine merge candidate. b. For example, it may refer to the derivation of an affine AMVP candidate. c. For example, it may refer to both the derivation of an affine Merge candidate and an affine AMVP candidate. d. For example, it may refer to the derivation of history-based affine candidates, non-neighbor-based affine candidates, and regression-based affine candidates. b. The logic / rules / process of similarity / consistency / deduplication checking may refer to comparing one or more of the following elements associated with the first affine candidate with those associated with the second affine candidate: a. Inter-frame direction (prediction direction); b. Affine type (e.g., 6-parameter affine or 4-parameter affine); c. Sub-block Merge type (e.g., sbTMVP or affine); d.Bcw index; e. LIC logo; f. Reference index; g. motion vector (eg, horizontal component and / or vertical component); h. Control point motion vector (CPMV); i. a first CPMV (eg, upper left CPMV) and / or a second CPMV (eg, upper right CPMV) and / or a third CPMV (eg, lower left CPMV); j. a horizontal displacement and / or a vertical displacement between the first CPMV and the second CPMV (eg, an absolute difference between a horizontal component and / or a vertical component of the first CPMV and the second CPMV); k. Horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV (eg, the absolute difference between the horizontal component and / or the vertical component of the first CPMV and the third CPMV). c. A comparison of the elements listed in item b can be checked for consistency. a. For example, if the element associated with the second candidate is the same as the element associated with the first candidate, the second candidate is not added to the affine candidate list. d. A similarity check can be performed on the comparisons listed in item b. a. For example, if the element associated with the second candidate is similar to the element associated with the first candidate, the second candidate is not added to the affine candidate list. b. For example, "similar" can refer to a comparison based on a threshold. i. For example, the absolute difference is less than a threshold. ii. For example, the absolute difference is not greater than a threshold. c. For example, the threshold may depend on block dimensions, such as width and / or height. i. For example, the threshold is adaptively determined according to the block width / height. ii. For example, a smaller threshold may be set for smaller block sizes, while a larger threshold may be set for larger block sizes. d. For example, the threshold value may be a predefined fixed value (such as 0 or 1). e. For example, similarity checks may be performed on the horizontal component and the vertical component of the motion vector separately. f. For example, similarity checks may be performed on the horizontal component and the vertical component of the control point motion vector respectively. e. Whether to apply similarity / consistency / deduplication check on the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV may depend on the affine type (eg, 6-parameter affine or 4-parameter affine). a. For example, only when the affine types of the first affine candidate and the second affine candidate are 6 parameters, a similarity / consistency / deduplication check of horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV may be performed. 4.2. Regarding the second problem of the codec tool for the screen content tool, the following method is proposed: a. At least one of the following codec tools may not be allowed, restricted or prohibited for a video unit. a) CIIP and / or its variants (e.g., CIIP PDPC, CIIP TM, CIIP TIMD wait). b) OBMC and / or its variants (eg, OBMC TM, etc.). c) TIMD and / or its variants. d) DIMD and / or its variants. e) MHP and / or variants thereof. f) DMVR and / or its variants. g)interTM and / or its variants. h) CCALF and / or its variants. i) CCSAO and / or its variants. b. Whether the codec is not allowed or restricted or prohibited for a video unit may depend on the profile / level / layer. c. Whether the codec tool is not allowed, restricted or prohibited for a video unit may depend on whether the video unit belongs to a specific video type. a) The specific video type may refer to screen content video. d. In addition, codec tools may not be allowed or restricted or prohibited for video sequences or groups of pictures or pictures or slices. a) This restriction or disallowance can be reflected by a bitstream constraint. b) Such constraints or disallowances or allowances may be reflected by syntax elements (eg, flags) signaled in the bitstream. e. Syntax elements (eg, flags) may be signaled in the bitstream to impose such constraints or disallow or allow for specific codecs listed in item a. a) The syntax elements may be signaled at the sequence level / picture group level / picture level / slice level / slice group level, such as in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. f. A single syntax element (eg, a flag) may be signaled to impose such a constraint or disallow or allow for more than one codec listed in item a. a) The syntax elements may be signaled at the sequence level / picture group level / picture level / slice level / slice group level, such as in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. g. Different weight values / factors / tables / sets may be allowed to mix multiple prediction hypotheses for a video unit coded with codec X. a) For example, which weight values / factors / tables / sets are allowed to mix multiple prediction hypotheses for a video unit may depend on the content type. b) For example, a first weight value / factor / table / set may be allowed to mix multiple prediction hypotheses for a first video unit, while a second weight value / factor / table / set may be allowed to mix multiple prediction hypotheses for a second video unit. c) For example, the first video unit may belong to a video sequence captured by a camera. d) For example, the second video unit may belong to a screen content video sequence. e) For example, suppose there are two prediction hypotheses that are mixed into the final prediction for a video unit: i. For a video unit of the first type, the final prediction block of the video unit encoded by codec X may be “completely equal to the first prediction hypothesis” or “completely equal to the second prediction hypothesis”. 1. For example, the allowed weight values / factors / table / set for a video unit may be equal to 0 or 1. ii. For the second type of video unit, the final hybrid prediction block of the video unit encoded by codec X may be a “fusion of the first prediction hypothesis and the second prediction hypothesis”. 1. For example, the allowed weight values / factors / table / set for a video unit may be equal to a fraction between 0 or 1 (eg, the fractional value of the weight of a particular prediction hypothesis may be quantized to an integer in the codec). f) For example, the final prediction (eg, mixed) of a video unit with codec X may be further mixed / weighted / fused with another video unit coded with another codec. g) For example, the specific codec tool X may be CIIP and / or its variants. h) For example, the specific codec tool X may be GPM and / or its variants. i) For example, the specific codec X may be MHP and / or its variants. j) For example, the specific codec X may be OBMC and / or its variants. k) For example, the specific codec tool X may be TIMD hybrid mode and / or its variants. l) For example, the specific codec tool X may be DIMD hybrid mode and / or its variants. 4.3. Regarding the third question of the general use of KLT for transformation processes, the following method is proposed: a. More than two KLT cores may be allowed in a codec. b. The KLT core may be enabled for the primary transform and / or the secondary transform. c. The KLT kernel may be allowed to be used for chroma components. a. For example, different KLT kernels may be used for the luma component and the chroma components of a video unit. i. Alternatively, all color components of a video unit may share the same KLT kernel. b. For example, different KLT kernels can be used for the chrominance Cb component and chrominance Cr component of a video unit. Quantity. i. Alternatively, the chroma Cb component and the chroma Cr component of a video unit can share the same KLT kernel. d. A pair of {KLT, flip-KLT} may be allowed / used for {horizontal, vertical} transformation or {vertical, horizontal} transformation of a video unit. a. For example, {KLT, flip-KLT} represents a pair of transform kernels including a first KLT and a second KLT, wherein the transform coefficient matrix of the second KLT (ie, flip-KLT) may be the transposed matrix of the transform coefficient matrix of the first KLT. e. The horizontal or vertical transform type of the video unit can be selected from {KLT-X, flip-KLT-X, DCT-Y, DST-Z}, where X / Y / Z are constants (e.g., Y=8, Z=7). a. For example, if more than one KLT core is defined in the codec, the more than one KLT core can be represented as KLT-X, such as X equals an integer value such as 1, 2, 3, ..., n, i.e., KLT-1, KLT-2, KLT-3, ..., KLT-n. b. For example, for a transform block, KLT can be used for horizontal (or vertical) transform, while flip-KLT can be used for vertical (or horizontal) transform. c. For example, for transform blocks, KLT can be used for horizontal (or vertical) transform instead of KLT can be used for vertical (or horizontal) transformation. d. For example, DCT2-KLT, KLT-DCT2 may be enabled for horizontal-to-vertical transform or vertical-to-horizontal transform of a block. e. For example, DST7-DCT2, DCT2-DST7, DCT8-DCT2, DCT2-DCT8 may be allowed to be used for horizontal-vertical transform or vertical-horizontal transform of blocks. f. For example, a video unit may be encoded or decoded using a specific prediction / transform / filter mode / technique. f. For a video unit in a specific mode codec, KLT can be the only transform type. a. In one example, the specific mode may be SBT. i. In one example, for different SBT modes (SBT_horizontal_split, SBT_vertical_split, SBT_half_split, SBT_quad_split), KLT can be used for both horizontal and vertical dimensions. b. In one example, the specific mode may be MIP. c. In one example, the same KLT may be used for both the horizontal dimension and the vertical dimension of a video unit for a particular mode codec. d. In one example, different KLTs may be used for the horizontal dimension and the vertical dimension of a video unit encoded in a particular mode. g. KLT based transform types may additionally be allowed for video units. a. For example, in addition to the existing MTS options (eg, MTS index from 0 to 5), the KLT-based transform type may be explicitly signaled. h. A KLT-based transform type may be applied to replace a specific existing transform type for a video unit. a. For example, the existing transform type to be replaced may be DST7, or DCT8, or DCT2. 4.4. Regarding the fourth problem of the mode-dependent KLT for the transform process, the following method is proposed: a. Which KLT core is used for a video unit may depend on a combination of at least one of the following types of codec information: a. Codec mode of the video unit. b. The size of the video unit. c. Motion vector of the video unit. d. Quantization parameters of video units. e. Temporal layer of video units. b. Which KLT core is used for a video unit may depend on the codec mode of the video unit. a. For example, it can be based on whether SBT is used for video units. b. For example, it can be based on whether implicit MTS is used for the video unit. c. For example, it can be based on whether an explicit MTS is used for the video unit. d. For example, it can be based on whether intra-MTS is used for the video unit. e. For example, it can be based on whether inter-MTS is used for the video unit. f. For example, it can be based on whether LFNST is used for video units. g. For example, it may be based on whether IBC and / or its variant modes are used for the video unit. h. For example, it may be based on whether PLT and / or its variant modes are used for video units. i. For example, it can be based on whether intra prediction mode is used for the video unit. j. For example, it can be based on whether the ISP and / or its variant mode is used for the video unit. k. For example, it may be based on whether MIP and / or its variant modes are used for the video unit. l. For example, it can be based on whether DIMD and / or its variant mode is used for the video unit. m. For example, it can be based on whether TIMD and / or its variant modes are used for the video unit. n. For example, it can be based on whether LM / CCLM / CCCM / GLM and / or its variant modes are used for the video unit. o. For example, it can be based on whether inter prediction mode is used for the video unit. p. For example, it can be based on whether AMVP mode is used for the video unit. q. For example, it can be based on whether Merge mode is used for the video unit. r. For example, it can be based on whether inter / intra / IBC template matching and / or its variant modes are used for the video unit. s. For example, it can be based on whether DMVR and / or its variant modes are used for the video unit. t. For example, it can be based on whether a sub-block prediction mode is used for a video unit. u. For example, it can be based on whether affine and / or variant modes are used for video units. v. For example, it can be based on whether sbTMVP and / or its variant modes are used for the video unit. w. For example, it may be based on whether hybrid / fusion / multi-hypothesis mode and / or its variant modes are used for the video unit. i. In one example, it can be based on whether the hybrid / fusion / multi-hypothesis mode contains intra-frame coding and decoding parts, such as GPM inter-intra, GPM intra, CIIP, MHP using intra, split GPM, etc. x. For example, it can be based on whether GPM and / or its variant modes are used for video units. y. For example, it can depend on whether CIIP and / or its variant patterns are used for the video unit. z. For example, it can depend on whether MHP and / or its variant patterns are used for the video unit. aa. For example, it can depend on whether OBMC and / or its variant patterns are used for the video unit. bb. For example, it can depend on whether LIC and / or its variant patterns are used for the video unit. c. Whether KLT is used and / or which KLT kernel is used for the video unit can depend on the size of the video unit. a. Different KLTs can be applied to blocks of different sizes. b. For example, it can depend on whether the width (W) and / or height (H) of the video unit satisfy predefined conditions, such as one or more combinations of the following: i. W < T1 or W <= T1, where T1 can be 8 or 16 or 32 or 64. ii. W > T2 or W >= T2, where T2 can be 2 or 4 or 8. iii. H < T3 or H <= T3, where T3 can be 8 or 16 or 32 or 64. iv. H > T4 or H >= T4, where T4 can be 2 or 4 or 8. v. W / H < T5 or W / H <= T5, where T5 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vi. W / H > T6 or W / H >= T6, where T6 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vii. H / W < T7 or H / W <= T7, where T7 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. viii. H / W > T8 or H / W >= T8, where T8 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. ix. W == T9, where T9 can be 8 or 16 or 32 or 64. x. H == T10, where T10 can be 8 or 16 or 32 or 64. d. Which KLT kernel is used for the video unit can depend on the motion vector of the video unit. a. In one example, it depends on the magnitude of the motion vector. e. Which KLT kernel is used for the video unit can depend on the quantization parameter of the video unit. a. In one example, it depends on the syntax level above the slice level (e.g., The base QP derived / signaled in the PPS or SPS). b. In one example, it depends on the slice QP. c. In one example, it depends on the QP of the codec unit. f. Which KLT kernel is used for a video unit may depend on the temporal layer of the video unit. a. In one example, it depends on whether it is in time domain layer 0 or in time domain layer 1 or in time domain layer 2 or ... . g. Depending on the codec mode and / or size of the video unit, different KLT cores may be enabled for different video units. a. In one example, the first KLT set can be used for the first pattern set, and the second The KLT set may be used for the second pattern set. i. For example, the first KLT set includes at least one type of KLT core. ii. For example, the second KLT set includes at least one type of KLT core. iii. For example, the first mode set contains at least one type of prediction / transform / filter mode. iv. For example, the second mode set includes at least one prediction / transform / filter mode. b. In one example, video units utilizing more than one different type of prediction / transform / filter mode codec can use the same KLT core. c. Alternatively, video units utilizing different types of prediction / transform / filter mode codecs may use different KLT cores. d. For example, one type of prediction / transform / filter mode may be a sub-block based prediction mode (eg, affine, sbTMVP, etc.). e. For example, one type of prediction / transform / filter mode may be an affine-based prediction mode (eg, Affine AMVP, Affine Merge, etc.). f. For example, a type of prediction / transform / filter mode can be based on mixing / fusion / Multiple hypothesis prediction modes (eg, GPM inter-intra, GPM intra, CIIP, etc.). g. For example, one type of prediction / transform / filter mode may be SBT and its variants. h. For example, one type of prediction / transform / filter mode may be ISP and its variants. i. For example, one type of prediction / transform / filter mode may be IBC and its variants. 4.5. In one example, how to encode and decode the residual block may depend on the selected transform. 4.6. In one example, the KLT may be divisible or indivisible. General Aspects 4.7. Whether and / or how to apply the method disclosed above can be signaled at the sequence level / GOP level / Picture level / Slice level / Slice group level, for example in the sequence header / Picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.8. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other kind of area containing more than one sample or pixel. 4.9. Whether and / or how to apply the methods disclosed above may depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type.

[0110] More details of embodiments of the present disclosure related to transformation, screen content coding (SCC) and affine in image / video coding will be described below. The embodiments of the present disclosure should be considered as examples to explain general concepts and should not be interpreted in a narrow manner. In addition, these embodiments can be applied individually or in combination in any way.

[0111] As used herein, the term "block" may refer to a color component, a sub-picture, a picture, a slice, a codec tree unit (CTU), a CTU row, a CTU group, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), a sub-block of a video block, a sub-region within a video block, a video processing unit including a plurality of samples / pixels, etc. Blocks can be rectangular or non-rectangular.

[0112] Fig.42 4 shows a flow chart of a method 4200 for video processing according to some embodiments of the present disclosure. The method 4200 may be implemented during a conversion between a current video block of a video and a bitstream of the video. Fig.42 As shown, method 4200 begins at 4202, where multiple affine candidates are obtained for a current video block.

[0113] At 4204, a deduplication check is applied to the plurality of affine candidates according to a checking process. The checking process is common to a plurality of different types of affine candidates. As an example and not limitation, the plurality of different types of affine candidates may include affine Merge candidates, affine Advanced Motion Vector Prediction (AMVP) candidates, history-based affine candidates, non-adjacent-based affine candidates, regression-based affine candidates, and the like.

[0114] At 4206, a conversion is performed based on the application. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above description and / or examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0115] In view of the above, the deduplication checking process is unified for multiple different types of affine candidates. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency.

[0116] In some embodiments, the plurality of affine candidates may include a first affine candidate and a second affine candidate. At least one of the following elements associated with the second affine candidate may be compared with the corresponding at least one element associated with the first affine candidate during the inspection process: inter-frame direction, prediction direction, affine type, sub-block Merge type, bidirectional prediction (BCW) index with codec unit-level weights, local illumination compensation (LIC) flag, reference index, motion vector, horizontal component of motion vector, vertical component of motion vector, at least one control point motion vector (CPMV), horizontal displacement between the first CPMV and the second CPMV, vertical displacement between the first CPMV and the second CPMV, horizontal displacement between the first CPMV and the third CPMV, vertical displacement between the first CPMV and the third CPMV, etc. As an example and not limitation, the first CPMV may be the upper left CPMV, or the second CPMV may be the upper right CPMV, and / or the third CPMV may be the lower left CPMV. It should be understood that the above examples are described for descriptive purposes only. The scope of the present disclosure is not limited in this regard.

[0117] In some embodiments, it may be checked in the deduplication check whether the second affine candidate is the same as the first affine candidate. In this case, the deduplication check may be a consistency check. For example, if at least one element associated with the second affine candidate is the same as the corresponding at least one element associated with the first affine candidate, the second affine candidate is not added to the affine candidate list for the current video block.

[0118] In some alternative embodiments, the similarity between the first affine candidate and the second affine candidate may be checked in a deduplication check. In this case, the deduplication check may be a similarity check. For example, if at least one element associated with the second affine candidate is similar to the corresponding at least one element associated with the first affine candidate, the second affine candidate may not be added to the affine candidate list for the current video block.

[0119] In some embodiments, if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than a threshold, the first element associated with the second affine candidate may be similar to the first element associated with the first affine candidate. Alternatively, if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than or equal to a threshold, the first element associated with the second affine candidate may be similar to the first element associated with the first affine candidate.

[0120] In some embodiments, the threshold value may depend on the size of the current video block. For example, the threshold value may be based on the width of the current video block and / or the height of the current video block.

[0121] In some embodiments, if the size of the current video block is smaller than the size of another video block of the video, the threshold for the current video block may be smaller than the threshold for the other video block. If the size of the current video block is larger than the size of the other video block, the threshold for the current video block may be larger than the threshold for the other video block. In some alternative embodiments, the threshold may be predefined and fixed.

[0122] In some embodiments, the similarity between the horizontal component of the motion vector for the first affine candidate and the horizontal component of the motion vector for the second affine candidate can be checked in the deduplication check. In addition, the similarity between the vertical component of the motion vector for the first affine candidate and the vertical component of the motion vector for the second affine candidate can be checked in the deduplication check.

[0123] In some embodiments, the similarity between the horizontal component of the CPMV for the first affine candidate and the horizontal component of the CPMV for the second affine candidate can be checked in the deduplication check. Additionally, the similarity between the vertical component of the CPMV for the first affine candidate and the vertical component of the CPMV for the second affine candidate can be checked in the deduplication check.

[0124] In some embodiments, whether the deduplication check is applied to at least one of the following may depend on the affine type of the first affine candidate and the affine type of the second affine candidate: the horizontal displacement between the first CPMV and the third CPMV, or the vertical displacement between the first CPMV and the third CPMV. For example, if the affine type of the first affine candidate and the affine type of the second affine candidate are 6-parameter affine, the deduplication check may be applied to the horizontal displacement between the first CPMV and the third CPMV and / or the vertical displacement between the first CPMV and the third CPMV.

[0125] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a device for video processing. In the method, multiple affine candidates for a current video block of the video are obtained. A deduplication check is applied to the multiple affine candidates according to a check process. The check process is common to multiple different types of affine candidates. In addition, a bitstream is generated based on the application.

[0126] According to some further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, multiple affine candidates for a current video block of the video are obtained. Deduplication check is applied to the multiple affine candidates according to a check process. The check process is common to multiple different types of affine candidates. In addition, a bitstream is generated based on the application; and the bitstream is stored in a non-transitory computer-readable recording medium.

[0127] Fig.43 4300 is a flowchart of a method 4300 for video processing according to some embodiments of the present disclosure. The method 4300 may be implemented during a conversion between a current video block of a video and a bitstream of the video. Fig.43 As shown, method 4300 starts at 4302, where information about applying a fusion-based codec tool to a current video block is obtained. The information depends on the video type of the current video block. As an example and not a limitation, if the video type of the current video block is a screen content video, the fusion-based codec tool can be disabled for the current video block.

[0128] As an example and not limitation, fusion-based codec tools may include: overlapped block motion compensation (OBMC), OMBC template matching (TM), combined inter-frame and intra-frame prediction (CIIP), CIIP position-dependent intra-frame prediction combination (PDPC), CIIP TM, CIIP template-based intra-frame mode derivation (TIMD), TIMD, decoder-side intra-frame mode derivation (DIMD), multiple hypothesis prediction (MHP), decoder-side motion vector refinement (DMVR), inter-frame template matching (interTM), cross-component adaptive loop filtering (CCALF), or cross-component sample adaptive compensation (CCSAO), or any variants thereof. It should be understood that the above examples are described for descriptive purposes only. The scope of the present disclosure is not limited in this regard.

[0129] In some embodiments, the information may include whether fusion-based codec tools are disabled for the current video block. Additionally or alternatively, the information may include how to apply fusion-based codec tools to the current video block.

[0130] At 4304, conversion is performed based on the application. In some embodiments, conversion may include encoding the current video block into a bitstream. Alternatively or additionally, conversion may include decoding the current video block from a bitstream.

[0131] In view of the above, the information about applying the fusion-based codec tool to the video block depends on the video type of the video block. Compared with conventional solutions, the proposed method can advantageously better support coding and decoding of different types of videos, thereby improving coding efficiency.

[0132] In some embodiments, the information may also depend on a grade, level or tier.

[0133] In some embodiments, the fusion-based codec tool may be disabled for a video sequence, a group of pictures, a picture, or a slice, etc. For example, a bitstream constraint may indicate that the fusion-based codec tool is disabled for at least one of the following: a video sequence, a group of pictures, a picture, or a slice. Alternatively, a syntax element in the bitstream may indicate that the fusion-based codec tool is disabled for at least one of the following: a video sequence, a group of pictures, a picture, or a slice.

[0134] In some embodiments, a syntax element indicating information about applying a fusion-based codec tool to a current video block may be included in the bitstream. As an example and not limitation, the syntax element may be indicated at a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level. Additionally or alternatively, the syntax element may be indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a slice group header.

[0135] In some embodiments, information about applying a fusion-based codec tool to a current video block may be indicated by a single syntax element in a bitstream, and the fusion-based codec tool may include more than one codec tool.

[0136] In some embodiments, the fusion-based codec tool may include a first codec tool in which multiple predictions for a video block are fused to generate a target prediction for the video block. As an example, the first codec tool may include CIIP, GPM, MHP, OBMC, TIMD, DIMD, etc. The first codec tool is applied to the current video block, and multiple sets of weights may be allowed to fuse multiple predictions for the current video block. It should be noted that the multiple sets of weights may be implemented as weight values, weight factors, weight tables, weight sets, etc.

[0137] In some embodiments, information about which set of weights is used to fuse multiple predictions for the current video block may depend on the content type of the current video block.

[0138] In some embodiments, the plurality of groups of weights may include a first group of weights and a second group of weights. The first group of weights may be used to fuse multiple predictions for a current video block, and the second group of weights may be used to fuse multiple predictions for another video block different from the current video block. For example, the content type of the current video block may be different from the content type of the other video block. In one example, the current video block belongs to a video sequence captured by a camera, and the other video block belongs to a screen content video sequence. Alternatively, the current video block belongs to a screen content video sequence, and the other video block belongs to a video sequence captured by a camera.

[0139] In some embodiments, the multiple predictions for the current video block may include two predictions, and the target prediction generated for the current video block may be the same as one of the two predictions. By way of example and not limitation, the weights in the first set of weights may be equal to 0 or 1. Additionally, the multiple predictions for another video block may include two predictions, and the target prediction generated for the another video block may be a mixture of the two predictions. By way of example and not limitation, the weights in the second set of weights may be equal to a fraction between 0 and 1.

[0140] In some further embodiments, the target prediction generated for the current video block may be merged with another video block encoded using another codec different from the first codec.

[0141] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, information about applying a fusion-based codec tool to a current video block of the video is obtained. The information depends on the video type of the current video block. In addition, a bitstream is generated based on the information.

[0142] According to some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, information about applying a fusion-based codec tool to a current video block of the video is obtained. The information depends on the video type of the current video block. In addition, a bitstream is generated based on the information, and the bitstream is stored in a non-transitory computer-readable recording medium.

[0143] Fig.44 4400 is a flowchart of a method 4400 for video processing according to some embodiments of the present disclosure. The method 4400 may be implemented during a conversion between a current video block of a video and a bitstream of the video. Fig.44 As shown, method 4400 starts at 4402, where at least one Karl-Hunin-Lew transform (KLT) kernel is selected from a plurality of KLT kernels based on codec information of a current video block. As used herein, the term "KLT" may refer to a type of transform. For example, it may refer to a Karl-Hunin-Lew transform. For example, it may include any suitable transform type that is not DCT or DST or Hadamard. By way of example and not limitation, the codec information may include a codec mode, size, motion vector, quantization parameter (QP), temporal layer, etc.

[0144] At 4404, conversion is performed based on at least one KLT core. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream.

[0145] In view of the above, the KLT kernel is selected by considering the coding information of the video block. Compared with the conventional solution, the proposed method can advantageously improve the coding efficiency.

[0146] In some embodiments, the codec information may include at least one of the following: whether sub-block transform (SBT) is applied to the current video block, whether implicit multi-transform selection (MTS) is applied to the current video block, whether explicit MTS is applied to the current video block, whether intra-frame MTS is applied to the current video block, whether inter-frame MTS is applied to the current video block, whether low-frequency non-separable transform (LFNST) is applied to the current video block, whether intra-frame block copy (IBC) is applied to the current video block, and the palette mode (P whether intra prediction mode is applied to the current video block, whether intra sub-partitioning (ISP) is applied to the current video block, whether matrix weighted intra prediction (MIP) is applied to the current video block, whether DIMD is applied to the current video block, whether TIMD is applied to the current video block, whether linear model (LM) is applied to the current video block, whether cross-component linear model (CCLM) is applied to the current video block, whether convolutional cross-component model (CCCM) is applied used on the current video block, whether the gradient-based linear model (GLM) is applied to the current video block, whether the inter-frame prediction mode is applied to the current video block; whether AMVP is applied to the current video block, whether the Merge mode is applied to the current video block, whether the inter-frame template matching is applied to the current video block, whether the intra-frame template matching is applied to the current video block, whether the IBC template matching is applied to the current video block, whether the DMVR is applied to the current video block, whether the sub-block prediction mode is applied to the current video block, whether the affine mode is applied to the current video block, whether the sub-block-based temporal motion vector prediction (SbTMVP) is applied to the current video block, whether the multiple hypothesis mode is applied to the current video block, whether the multiple hypothesis mode includes the intra-frame coding and decoding part, whether the geometric partitioning mode (GPM) is applied to the current video block, whether the CIIP is applied to the current video block, whether the MHP is applied to the current video block, whether the OBMC is applied to the current video block, or whether the LIC is applied to the current video block. It should be understood that the above description is only for the purpose of illustration. The scope of the present disclosure is not limited in this regard.

[0147] In some embodiments, different KLT cores may be applied to video blocks of different sizes. For example, the codec information may include the size, and at least one KLT core may be selected based on whether at least one of the following satisfies a predefined condition: the width of the current video block; or the height of the current video block.

[0148] In some embodiments, the predefined condition may include at least one of the following: the width of the current video block is less than a first value, the width of the current video block is less than or equal to the first value, the width of the current video block is greater than a second value, the width of the current video block is greater than or equal to the second value, the height of the current video block is less than a third value, the height of the current video block is less than or equal to the third value, the height of the current video block is greater than a fourth value, the height of the current video block is greater than or equal to the fourth value, a first result of dividing the width of the current video block by the height of the current video block is less than a fifth value, the first result is less than or equal to the fifth value, the first result is greater than a sixth value, the first result is greater than or equal to the sixth value, a second result of dividing the height of the current video block by the width of the current video block is less than a seventh value, the second result is less than or equal to the seventh value, the second result is greater than an eighth value, the second result is greater than or equal to the eighth value, the width of the current video block is equal to a ninth value, or the height of the current video block is equal to a tenth value.

[0149] As an example, the first value may be 8 or 16 or 32 or 64, or the second value may be 2 or 4 or 8, or the third value may be 8 or 16 or 32 or 64, or the fourth value may be 2 or 4 or 8, or the fifth value may be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the sixth value may be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the seventh value may be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the eighth value may be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the ninth value may be 8 or 16 or 32 or 64, or the tenth value may be 8 or 16 or 32 or 64. It should be understood that the specific values ​​described herein are intended to be exemplary and not to limit the scope of the present disclosure.

[0150] In some embodiments, different KLT cores may be enabled for different video blocks depending on the codec mode or size of the different video blocks of the video.

[0151] In some embodiments, the first group of KLT cores can be used for video blocks encoded and decoded using the first group of codec modes, and the second group of KLT cores can be used for video blocks encoded and decoded using the second group of codec modes. For example, the first group of KLT cores can include at least one type of KLT core, or the second group of KLT cores can include at least one type of KLT core. Additionally, the first group of codec modes can include at least one type of prediction mode, at least one type of transform mode, or at least one type of filter mode. In addition, the second group of codec modes can include at least one type of prediction mode, at least one type of transform mode, or at least one type of filter mode.

[0152] In some embodiments, the same KLT core can be used for video blocks encoded and decoded using more than one different type of codec mode. Alternatively, different KLT cores can be used for video blocks encoded and decoded using different types of codec modes. For example, a codec mode can include a prediction mode, a transform mode, or a filter mode.

[0153] In some embodiments, a type of coding mode may include: a sub-block based prediction mode, an affine based prediction mode, a fusion based prediction mode, SBT, ISP, or IBC. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0154] In some embodiments, the KLT core may be divisible or indivisible. Additionally or alternatively, the plurality of KLT cores may include more than two KLT cores. In some additional or alternative embodiments, at least one KLT core may be allowed to be used for the following primary transformation and / or secondary transformation.

[0155] In some embodiments, the current video block may be encoded using the first encoding mode, and different KLT cores may be used for horizontal transform and vertical transform of the current video block. Alternatively, the same KLT core may be used for horizontal transform and vertical transform of the current video block.

[0156] In some embodiments, the current video block may be encoded using a first codec mode, and only KLT type transforms may be allowed for the current video block.

[0157] In some embodiments, the first coding mode may include SBT or MIP. Additionally or alternatively, the first coding mode may include at least one of SBT horizontal division mode, SBT vertical division mode, SBT half division mode, or SBT four division mode.

[0158] In some embodiments, the predetermined transform type may be replaced by a KLT-based transform type. As an example and not limitation, the predetermined transform type may include: discrete cosine transform type 2 (DCT2), discrete sine transform type 7 (DST7), or discrete sine transform type 8 (DST8).

[0159] In some embodiments, at least one KLT core may be allowed to be used for the chrominance component of the current video block. In one example, different KLT cores may be used for the luminance component and the chrominance component of the current video block. In another example, the same KLT core may be used for all color components of the current video block. In another example, different KLT cores may be used for the chrominance blue (Cb) component and the chrominance red (Cr) component of the current video block. In yet another example, the same KLT core may be used for the Cb component and the Cr component of the current video block.

[0160] In some embodiments, a pair of a first KLT kernel and a second KLT kernel may be used for horizontal transformation and vertical transformation of the current video block. The transformation coefficient matrix of the second KLT kernel may be a transposed matrix of the transformation coefficient matrix of the first KLT kernel.

[0161] In some embodiments, the horizontal transform type or the vertical transform type for the current video block may be selected from KLT-X, flip KLT-X, DCT-Y, and DST-Z, and each of X, Y, and Z may be a constant, such as Y=8, Z=7, etc. The transform coefficient matrix of the flip KLT kernel may be a transposed matrix of the transform coefficient matrix of the KLT kernel.

[0162] In some embodiments, more than one KLT core may be defined. Each of the more than one KLT cores may be represented as KLT-X, and X may be equal to an integer value such as 1, 2, 3, ..., etc., i.e., KLT-1, KLT-2, KLT-3, ...

[0163] In some embodiments, the current video block may be a transform block. Additionally, the KLT kernel may be used for a horizontal transform of the current video block, and the flip KLT kernel may be used for a vertical transform of the current video block. Alternatively, the KLT kernel may be used for a vertical transform of the current video block, and the flip KLT kernel may be used for a horizontal transform of the current video block.

[0164] In some embodiments, the current video block may be a transform block. Additionally, a KLT kernel may be used for a horizontal transform of the current video block, and a non-KLT kernel may be used for a vertical transform of the current video block. Alternatively, a KLT kernel may be used for a vertical transform of the current video block, and a non-KLT kernel may be used for a horizontal transform of the current video block. As an example and not limitation, a non-KLT kernel may be a DCT2 kernel, etc.

[0165] In some embodiments, wherein the DST7 core may be used for the horizontal transform of the current video block, and the DCT2 core may be used for the vertical transform of the current video block. Alternatively, the DCT2 core may be used for the horizontal transform of the current video block, and the DST7 core may be used for the vertical transform of the current video block. In some further embodiments, the DCT8 core may be used for the horizontal transform of the current video block, and the DCT2 core may be used for the vertical transform of the current video block. In some further embodiments, the DCT2 core may be used for the horizontal transform of the current video block, and the DCT8 core may be used for the vertical transform of the current video block.

[0166] In some embodiments, the current video block may be encoded using a specific prediction mode, a specific transform mode, or a specific filter mode. In some embodiments, in addition to a predetermined MTS option, a KLT-based transform type may be explicitly signaled.

[0167] In some embodiments, the codec information may include a motion vector, and at least one KLT kernel may be selected based on the size of the motion vector for the current video block. Additionally or alternatively, the codec information may include one of: a base QP determined at a syntax level higher than the slice level, a base QP signaled at a syntax level higher than the slice level, a slice QP, or a QP of the current video block.

[0168] In some embodiments, the codec information may include a temporal layer, and at least one KLT core may be selected based on whether the temporal layer of the current video block is a predetermined temporal layer (such as temporal layer 0, etc.).

[0169] In some embodiments, the current video block may be a residual block.In addition, the information on how to encode and decode the current video block may depend on at least one KLT core.

[0170] In some embodiments, whether and / or how to apply the method may be indicated at a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level. In some embodiments, whether and / or how to apply the method may be indicated in a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, a slice group header, etc.

[0171] In some embodiments, whether and / or how to apply a method may be indicated at a region containing more than one sample or pixel. By way of example and not limitation, a region may include at least one of a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

[0172] In some embodiments, whether and / or how to apply the method may depend on the encoded information. As an example and not limitation, the encoded information may include at least one of the following: block size, color format, single-double tree segmentation, double tree segmentation, color component, stripe type, or picture type. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0173] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, at least one Karhuning-Love transform (KLT) core is selected from a plurality of KLT cores based on codec information of a current video block of the video. In addition, a bitstream is generated based on at least one KLT core.

[0174] According to some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, at least one Karl Huning-Love transform (KLT) core is selected from a plurality of KLT cores based on codec information of a current video block of the video. In addition, a bitstream is generated based on the at least one KLT core, and the bitstream is stored in a non-transitory computer-readable recording medium.

[0175] Embodiments of the present disclosure may be described according to the following items, features of which may be combined in any reasonable way.

[0176] Item 1. A method for video processing, comprising: obtaining a plurality of affine candidates for a conversion between a current video block of a video and a bitstream of the video; applying a deduplication check to the plurality of affine candidates according to a checking process, the checking process being common to a plurality of different types of affine candidates; and performing the conversion based on the application.

[0177] Item 2. A method according to Item 1, wherein the multiple different types of affine candidates include at least one of the following: an affine merge candidate, an affine advanced motion vector prediction (AMVP) candidate, a history-based affine candidate, a non-neighbor-based affine candidate, or a regression-based affine candidate.

[0178] Item 3. A method according to any one of Items 1 to 2, wherein the multiple affine candidates include a first affine candidate and a second affine candidate, and at least one of the following elements associated with the second affine candidate is compared with the corresponding at least one element associated with the first affine candidate during the checking process: inter-frame direction, prediction direction, affine type, sub-block merge type, bidirectional prediction (BCW) index with codec unit level weight, local illumination compensation (LIC) flag, reference index, motion vector, horizontal component of the motion vector, vertical component of the motion vector, at least one control point motion vector (CPMV), horizontal displacement between the first CPMV and the second CPMV, vertical displacement between the first CPMV and the second CPMV, horizontal displacement between the first CPMV and the third CPMV, or vertical displacement between the first CPMV and the third CPMV.

[0179] Item 4. The method according to Item 3, wherein the first CPMV is a top-left CPMV, or the second CPMV is a top-right CPMV, or the third CPMV is a bottom-left CPMV.

[0180] Item 5. A method according to any one of Items 3 to 4, wherein in the deduplication check it is checked whether the second affine candidate is the same as the first affine candidate.

[0181] Item 6. A method according to Item 5, wherein if the at least one element associated with the second affine candidate is the same as the corresponding at least one element associated with the first affine candidate, the second affine candidate is not added to the affine candidate list for the current video block.

[0182] Item 7. A method according to any one of Items 3 to 4, wherein similarity between the first affine candidate and the second affine candidate is checked in the deduplication check.

[0183] Item 8. A method according to Item 7, wherein if the at least one element associated with the second affine candidate is similar to the corresponding at least one element associated with the first affine candidate, the second affine candidate is not added to the affine candidate list for the current video block.

[0184] Item 9. A method according to Item 8, wherein the first element associated with the second affine candidate is similar to the first element associated with the first affine candidate if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than a threshold.

[0185] Item 10. A method according to Item 8, wherein the first element associated with the second affine candidate is similar to the first element associated with the first affine candidate if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than or equal to a threshold.

[0186] Item 11. A method according to any one of Items 9 to 10, wherein the threshold value depends on the size of the current video block.

[0187] Item 12. A method according to any one of Items 9 to 11, wherein the threshold is determined based on at least one of: a width of the current video block, or a height of the current video block.

[0188] Item 13. A method according to Item 11, wherein if the size of the current video block is smaller than the size of another video block of the video, the threshold used for the current video block is smaller than the threshold used for the other video block, and if the size of the current video block is larger than the size of the other video block, the threshold used for the current video block is larger than the threshold used for the other video block.

[0189] Item 14. A method according to any one of Items 9 to 10, wherein the threshold is predefined and fixed.

[0190] Item 15. A method according to any one of Items 3 to 14, wherein the similarity between the horizontal component of the motion vector for the first affine candidate and the horizontal component of the motion vector for the second affine candidate is checked in the deduplication check, and the similarity between the vertical component of the motion vector for the first affine candidate and the vertical component of the motion vector for the second affine candidate is checked in the deduplication check.

[0191] Item 16. A method according to any one of Items 3 to 15, wherein the similarity between the horizontal component of the CPMV for the first affine candidate and the horizontal component of the CPMV for the second affine candidate is checked in the deduplication check, and the similarity between the vertical component of the CPMV for the first affine candidate and the vertical component of the CPMV for the second affine candidate is checked in the deduplication check.

[0192] Item 17. A method according to any one of Items 3 to 16, wherein whether the deduplication check is applied to at least one of the following depends on the affine type of the first affine candidate and the affine type of the second affine candidate: the horizontal displacement between the first CPMV and the third CPMV, or the vertical displacement between the first CPMV and the third CPMV.

[0193] Item 18. A method according to Item 17, wherein if the affine type of the first affine candidate and the affine type of the second affine candidate are 6-parameter affines, the deduplication check is applied to at least one of the following: the horizontal displacement between the first CPMV and the third CPMV, or the vertical displacement between the first CPMV and the third CPMV.

[0194] Item 19. A method for video processing, comprising: for conversion between a current video block of a video and a bitstream of the video, obtaining information about applying a fusion-based codec tool to the current video block, the information depending on the video type of the current video block; and performing the conversion based on the information.

[0195] Item 20. The method of Item 19, wherein the information comprises at least one of: whether the fusion-based codec tool is disabled for the current video block, or how to apply the fusion-based codec tool to the current video block.

[0196] Item 21. A method according to any one of Items 19 to 20, wherein the fusion-based codec tool includes at least one of the following: overlapped block motion compensation (OBMC), OMBC template matching (TM), combined inter-frame and intra-frame prediction (CIIP), CIIP position-dependent intra-frame prediction combination (PDPC), CIIP TM, CIIP template-based intra-frame mode derivation (TIMD), TIMD, decoder-side intra-frame mode derivation (DIMD), multiple hypothesis prediction (MHP), decoder-side motion vector refinement (DMVR), inter-frame template matching (interTM), cross-component adaptive loop filtering (CCALF), or cross-component sample adaptive compensation (CCSAO).

[0197] Item 22. The method of any one of Items 19 to 21, wherein if the video type of the current video block is screen content video, the fusion-based codec tool is disabled for the current video block.

[0198] Item 23. A method according to any one of items 19 to 22, wherein the information also depends on the grade, level or layer.

[0199] Item 24. A method according to any one of Items 19 to 23, wherein the fusion-based codec tool is disabled for at least one of: a video sequence, a group of pictures, a picture; or a slice.

[0200] Item 25. A method according to any one of Items 19 to 24, wherein the bitstream constraint indicates that the fusion-based codec tool is disabled for at least one of the following: a video sequence, a group of pictures, a picture; or a slice.

[0201] Item 26. A method according to any one of Items 19 to 25, wherein a syntax element in the bitstream indicates that the fusion-based codec tool is disabled for at least one of the following: a video sequence, a group of pictures, a picture; or a slice.

[0202] Clause 27. A method according to any one of clauses 19 to 26, wherein a syntax element indicating the information is included in the bitstream.

[0203] Item 28. The method of Item 27, wherein the syntax element is indicated at one of: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.

[0204] Item 29. A method according to any one of Items 27 to 28, wherein the syntax element is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0205] Item 30. A method according to any one of Items 19 to 29, wherein the information is indicated by a single syntax element in the bitstream and the fusion-based codec tool includes more than one codec tool.

[0206] Item 31. A method according to any one of Items 19 to 30, wherein the fusion-based codec tool includes a first codec tool in which multiple predictions for a video block are fused to generate a target prediction for the video block, the first codec tool is applied to the current video block, and multiple sets of weights are allowed to fuse multiple predictions for the current video block.

[0207] Item 32. The method of Item 31, wherein information about which set of weights is used to fuse the multiple predictions for the current video block depends on a content type of the current video block.

[0208] Item 33. A method according to any one of Items 31 to 32, wherein the multiple sets of weights include a first set of weights and a second set of weights, the first set of weights is used to fuse the multiple predictions for the current video block, and the second set of weights is used to fuse the multiple predictions for another video block different from the current video block.

[0209] Clause 34. The method of clause 33, wherein the content type of the current video block is different from the content type of the another video block.

[0210] Item 35. A method according to any one of Items 33 to 34, wherein the current video block belongs to a video sequence captured by a camera and the other video block belongs to a screen content video sequence, or wherein the current video block belongs to a screen content video sequence and the other video block belongs to a video sequence captured by a camera.

[0211] Item 36. A method according to any one of Items 33 to 35, wherein the multiple predictions for the current video block include two predictions, and the target prediction generated for the current video block is the same as one of the two predictions.

[0212] Item 37. A method according to Item 36, wherein the weights in the first set of weights are equal to 0 or 1.

[0213] Item 38. The method of any one of Items 33 to 37, wherein the multiple predictions for the other video block include two predictions, and the target prediction generated for the other video block is a mix of the two predictions.

[0214] Item 39. A method according to Item 38, wherein the weights in the second set of weights are equal to a fraction between 0 and 1.

[0215] Item 40. The method of any one of Items 31 to 39, wherein the target prediction generated for the current video block is fused with another video block encoded using another codec different from the first codec.

[0216] Item 41. A method according to any one of Items 31 to 40, wherein the first codec tool includes one of the following: CIIP, GPM, MHP, OBMC, TIMD, or DIMD.

[0217] Item 42. A method for video processing, comprising: for conversion between a current video block of a video and a bitstream of the video, selecting at least one Karhuning-Lew transform (KLT) core from a plurality of KLT cores based on codec information of the current video block; and performing the conversion based on the at least one KLT core.

[0218] Item 43. The method according to Item 42, wherein the codec information includes at least one of the following: codec mode, size, motion vector, quantization parameter (QP), or time domain layer.

[0219] Item 44. A method according to any one of Items 42 to 43, wherein the codec information includes at least one of the following: whether a sub-block transform (SBT) is applied to the current video block, whether an implicit multi-transform selection (MTS) is applied to the current video block, whether an explicit MTS is applied to the current video block, whether an intra-frame MTS is applied to the current video block, whether an inter-frame MTS is applied to the current video block, whether a low-frequency non-separable transform (LFNST) is applied to the current video block, whether an intra-frame block copy (IBC) is applied to the current video block, and whether the sub-block transform (SBT) is applied to the current video block. frequency block, whether the palette mode (PLT) is applied to the current video block, whether the intra prediction mode is applied to the current video block, whether the intra sub-partitioning (ISP) is applied to the current video block, whether the matrix weighted intra prediction (MIP) is applied to the current video block, whether DIMD is applied to the current video block, whether TIMD is applied to the current video block, whether the linear model (LM) is applied to the current video block, whether the cross-component linear model (CCLM) is applied to the current video block, and whether the convolutional cross-component model (CCCM) is applied to the current video block. ) is applied to the current video block, whether the gradient-based linear model (GLM) is applied to the current video block, whether the inter prediction mode is applied to the current video block, whether AMVP is applied to the current video block, whether the Merge mode is applied to the current video block, whether the inter template matching is applied to the current video block, whether the intra template matching is applied to the current video block, whether the IBC template matching is applied to the current video block, whether the DMVR is applied to the current video block, whether the sub-block prediction mode is applied to the the current video block, whether the affine mode is applied to the current video block, whether the sub-block based temporal motion vector prediction (SbTMVP) is applied to the current video block, whether the multi-hypothesis mode is applied to the current video block, whether the multi-hypothesis mode includes an intra-frame coding and decoding part, whether the geometric partitioning mode (GPM) is applied to the current video block, whether the CIIP is applied to the current video block, whether the MHP is applied to the current video block, whether the OBMC is applied to the current video block, or whether the LIC is applied to the current video block.

[0220] Item 45. A method according to any of Items 42 to 44, wherein different KLT kernels are applied on video blocks with different sizes.

[0221] Item 46. A method according to any one of Items 42 to 45, wherein the codec information includes a size, and the at least one KLT core is selected based on whether at least one of the following satisfies a predefined condition: the width of the current video block; or the height of the current video block.

[0222] Item 47. A method according to Item 46, wherein the predefined condition includes at least one of the following: the width of the current video block is less than a first value, the width of the current video block is less than or equal to the first value, the width of the current video block is greater than a second value, the width of the current video block is greater than or equal to the second value, the height of the current video block is less than a third value, the height of the current video block is less than or equal to the third value, the height of the current video block is greater than a fourth value, the height of the current video block is greater than or equal to the fourth value, a first result of dividing the width of the current video block by the height of the current video block is less than a fifth value, the first result is less than or equal to the fifth value, the first result is greater than a sixth value, and the first result is greater than or equal to the sixth value, a second result of dividing the height of the current video block by the width of the current video block is less than a seventh value, the second result is less than or equal to the seventh value, the second result is greater than an eighth value, the second result is greater than or equal to the eighth value, the width of the current video block is equal to a ninth value, or the height of the current video block is equal to a tenth value.

[0223] Item 48. A method according to Item 47, wherein the first value is 8 or 16 or 32 or 64, or the second value is 2 or 4 or 8, or the third value is 8 or 16 or 32 or 64, or the fourth value is 2 or 4 or 8, or the fifth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the sixth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the seventh value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the eighth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or the ninth value is 8 or 16 or 32 or 64, or the tenth value is 8 or 16 or 32 or 64.

[0224] Item 49. A method according to any of Items 42 to 48, wherein different KLT cores are enabled for different video blocks of the video depending on a coding mode or size of the different video blocks.

[0225] Item 50. A method according to any of Items 42 to 49, wherein a first set of KLT cores are used for video blocks encoded and decoded using a first set of codec modes, and a second set of KLT cores are used for video blocks encoded and decoded using a second set of codec modes.

[0226] Item 51. A method according to Item 50, wherein the first group of KLT cores includes at least one type of KLT core, or the second group of KLT cores includes at least one type of KLT core.

[0227] Item 52. A method according to any one of Items 50 to 51, wherein the first set of coding modes includes at least one type of prediction mode, at least one type of transform mode or at least one type of filter mode, or wherein the second set of coding modes includes at least one type of prediction mode, at least one type of transform mode or at least one type of filter mode.

[0228] Item 53. A method according to any of Items 42 to 48, wherein the same KLT core is used for video blocks encoded using more than one different type of codec mode.

[0229] Item 54. The method of any one of Items 42 to 48, wherein different KLT cores are used for video blocks encoded using different types of codec modes.

[0230] Item 55. A method according to any one of Items 53 to 54, wherein the encoding and decoding mode comprises a prediction mode, a transform mode or a filter mode.

[0231] Item 56. A method according to any one of Items 53 to 55, wherein a type of coding mode includes: sub-block based prediction mode, affine based prediction mode, fusion based prediction mode, SBT, ISP, or IBC.

[0232] Item 57. The method according to any one of items 42 to 46, wherein the KLT core is divisible or inseparable.

[0233] Item 58. The method according to any one of Items 42 to 57, wherein the plurality of KLT cores comprises more than two KLT cores.

[0234] Item 59. A method according to any one of Items 42 to 58, wherein the at least one KLT core is enabled to be used for at least one of: a primary transform, or a secondary transform.

[0235] Item 60. The method of any one of Items 42 to 59, wherein the current video block is encoded using a first codec mode, and different KLT kernels are used for horizontal transform and vertical transform of the current video block.

[0236] Item 61. A method according to any one of Items 42 to 59, wherein the current video block is encoded using a first codec mode and the same KLT kernel is used for a horizontal transform and a vertical transform of the current video block.

[0237] Item 62. A method according to any one of Items 42 to 61, wherein the current video block is encoded using a first codec mode and only KLT type transforms are allowed for the current video block.

[0238] Item 63. A method according to any one of items 60 to 62, wherein the first coding mode comprises SBT or MIP.

[0239] Item 64. The method according to Item 60, wherein the first encoding and decoding mode includes at least one of the following: SBT horizontal partition mode, SBT vertical partition mode, SBT half partition mode or SBT four partition mode.

[0240] Item 65. A method according to any one of Items 42 to 64, wherein the predetermined transform type is replaced by a KLT-based transform type.

[0241] Item 66. A method according to item 65, wherein the predetermined transform type includes one of the following: discrete cosine transform type 2 (DCT2), discrete sine transform type 7 (DST7), or discrete sine transform type 8 (DST8).

[0242] Item 67. The method of any one of Items 42 to 66, wherein the at least one KLT kernel is enabled for chroma components of the current video block.

[0243] Item 68. The method of any one of Items 42 to 67, wherein different KLT kernels are used for luma and chroma components of the current video block.

[0244] Item 69. The method of any one of Items 42 to 67, wherein the same KLT kernel is used for all color components of the current video block.

[0245] Item 70. The method of any one of Items 42 to 67, wherein different KLT kernels are used for a chrominance blue (Cb) component and a chrominance red (Cr) component of the current video block.

[0246] Item 71. A method according to any one of Items 42 to 67, wherein the same KLT kernel is used for the Cb component and the Cr component of the current video block.

[0247] Item 72. A method according to any one of Items 42 to 71, wherein a pair of a first KLT kernel and a second KLT kernel are used for horizontal transformation and vertical transformation of the current video block, and the transformation coefficient matrix of the second KLT kernel is the transposed matrix of the transformation coefficient matrix of the first KLT kernel.

[0248] Item 73. A method according to any one of items 42 to 72, wherein the horizontal transform type or the vertical transform type for the current video block is selected from KLT-X, flip KLT-X, DCT-Y and DST-Z, and each of X, Y and Z is a constant.

[0249] Item 74. A method according to Item 73, wherein the transform coefficient matrix of the flipped KLT kernel is the transposed matrix of the transform coefficient matrix of the KLT kernel.

[0250] Item 75. A method according to any one of Items 73 to 74, wherein more than one KLT core is defined, each of the more than one KLT cores is denoted as KLT-X, and X is equal to an integer value.

[0251] Item 76. A method according to any one of Items 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a horizontal transform of the current video block, and a flipped KLT kernel is used for a vertical transform of the current video block.

[0252] Item 77. A method according to any one of Items 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a vertical transform of the current video block, and a flipped KLT kernel is used for a horizontal transform of the current video block.

[0253] Item 78. A method according to any one of Items 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a horizontal transform of the current video block, and a non-KLT kernel is used for a vertical transform of the current video block.

[0254] Item 79. A method according to any one of Items 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a vertical transform of the current video block, and a non-KLT kernel is used for a horizontal transform of the current video block.

[0255] Item 80. The method according to any one of Items 78 to 79, wherein the non-KLT core is a DCT2 core.

[0256] Item 81. A method according to any one of Items 73 to 74, wherein a DST7 core is used for a horizontal transform of the current video block and a DCT2 core is used for a vertical transform of the current video block, or wherein a DCT2 core is used for the horizontal transform of the current video block and a DST7 core is used for the vertical transform of the current video block, or wherein a DCT8 core is used for the horizontal transform of the current video block and a DCT2 core is used for the vertical transform of the current video block, or wherein a DCT2 core is used for the horizontal transform of the current video block and a DCT8 core is used for the vertical transform of the current video block.

[0257] Item 82. A method according to any one of items 42 to 81, wherein the current video block is encoded using a specific prediction mode, a specific transform mode, or a specific filter mode.

[0258] Item 83. A method according to any one of Items 42 to 82, wherein in addition to a predetermined MTS option, a KLT-based transform type is explicitly signaled.

[0259] Item 84. The method of any one of Items 42 to 83, wherein the codec information comprises a motion vector, and the at least one KLT kernel is selected based on a magnitude of the motion vector for the current video block.

[0260] Item 85. A method according to any one of Items 42 to 84, wherein the codec information includes one of the following: a base QP determined at a syntax level higher than the slice level, a base QP transmitted by a signal at a syntax level higher than the slice level, a slice QP, or a QP of the current video block.

[0261] Item 86. A method according to any one of Items 42 to 85, wherein the codec information includes a time domain layer, and the at least one KLT core is selected based on whether the time domain layer of the current video block is a predetermined time domain layer.

[0262] Item 87. The method of any one of Items 42 to 86, wherein the current video block is a residual block, and the information about how to encode and decode the current video block depends on the at least one KLT kernel.

[0263] Item 88. A method according to any one of Items 1 to 87, wherein whether and / or how the method is applied is indicated at one of: sequence level, group of picture level, picture level, slice level, or slice group level.

[0264] Item 89. A method according to any one of Items 1 to 88, wherein whether and / or how the method is applied is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0265] Item 90. A method according to any one of Items 1 to 87, wherein whether and / or how to apply the method is indicated at an area containing more than one sample or pixel.

[0266] Item 91. A method according to item 90, wherein the region includes at least one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, or a sub-picture.

[0267] Item 92. A method according to any one of items 1 to 87, wherein whether and / or how the method is applied depends on the encoded and decoded information.

[0268] Item 93. The method of Item 92, wherein the encoded information comprises at least one of: block size, color format, single or double tree partitioning, double tree partitioning, color component, slice type, or picture type.

[0269] Item 94. A method according to any one of Items 1 to 93, wherein the converting includes encoding the current video block into the bitstream.

[0270] Item 95. A method according to any one of Items 1 to 93, wherein the converting comprises decoding the current video block from the bitstream.

[0271] Item 96. A device for video processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 95.

[0272] Item 97. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of Items 1 to 95.

[0273] Item 98. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: obtaining multiple affine candidates for a current video block of the video; applying a deduplication check to the multiple affine candidates according to a check process, the check process being common to multiple different types of affine candidates; and generating the bitstream based on the application.

[0274] Item 99. A method for storing a bitstream of a video, comprising: obtaining multiple affine candidates for a current video block of the video; applying a deduplication check to the multiple affine candidates according to a check process, the check process being common to multiple different types of affine candidates; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.

[0275] Item 100. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on a video type of the current video block; and generating the bitstream based on the information.

[0276] Item 101. A method for storing a bitstream of a video, comprising: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on the video type of the current video block; generating the bitstream based on the information; and storing the bitstream in a non-transitory computer-readable recording medium.

[0277] Item 102. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a device for video processing, wherein the method comprises: selecting at least one Karl-Hunin-Löw transform (KLT) kernel from a plurality of KLT kernels based on codec information of a current video block of the video; and generating the bitstream based on the at least one KLT kernel.

[0278] Item 103. A method for storing a bitstream of a video, comprising: selecting at least one Karl-Hunin-Löw transform (KLT) core from a plurality of KLT cores based on codec information of a current video block of the video; generating the bitstream based on the at least one KLT core; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device

[0279] Fig.45A block diagram of a computing device 4500 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4500 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0280] It should be understood that Fig.45 The computing device 4500 shown in the figure is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0281] like Fig.45 As shown, computing device 4500 includes a general computing device 4500. Computing device 4500 may include at least one or more processors or processing units 4510, memory 4520, storage unit 4530, one or more communication units 4540, one or more input devices 4550, and one or more output devices 4560.

[0282] In some embodiments, computing device 4500 can be implemented as any user terminal or server terminal with computing power. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including the accessories and peripherals of these devices, or any combination thereof. It is conceivable that computing device 4500 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0283] The processing unit 4510 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 4520. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to increase the parallel processing capability of the computing device 4500. The processing unit 4510 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0284] The computing device 4500 typically includes various computer storage media. Such media can be any media accessible by the computing device 4500, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4520 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 4530 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 4500.

[0285] The computing device 4500 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Fig.45 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.

[0286] The communication unit 4540 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 4500 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 4500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0287] Input device 4550 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4560 may be one or more of various output devices, such as a display, speaker, printer, etc. With the aid of communication unit 4540, computing device 4500 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and may also communicate with one or more devices that enable a user to interact with computing device 4500, or, if necessary, may also communicate with any device (e.g., a network card, a modem, etc.) that enables computing device 4500 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).

[0288] In some embodiments, some or all components of the computing device 4500 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, the cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.

[0289] In an embodiment of the present disclosure, the computing device 4500 may be used to implement video encoding / decoding. The memory 4520 may include one or more video codec modules 4525 having one or more program instructions. These modules are accessible and executable by the processing unit 4510 to perform the functions of the various embodiments described herein.

[0290] In an example embodiment performing video encoding, input device 4550 may receive video data as input 4570 to be encoded. The video data may be processed, for example, by video codec module 4525 to generate an encoded bitstream. The encoded bitstream may be provided as output 4580 via output device 4560.

[0291] In an example embodiment performing video decoding, input device 4550 may receive an encoded bitstream as input 4570. The encoded bitstream may be processed, for example, by video codec module 4525 to generate decoded video data. The decoded video data may be provided as output 4580 via output device 4560.

[0292] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, include: For conversion between a current video block of a video and a bitstream of the video, obtaining a plurality of affine candidates for the current video block; applying a deduplication check to the plurality of affine candidates according to a checking process, the checking process being common to a plurality of different types of affine candidates; as well as The converting is performed based on the application.

2. The method according to claim 1, wherein the plurality of different types of affine candidates include at least one of the following: Affine Merge candidate, Affine Advanced Motion Vector Prediction (AMVP) candidates, History-based affine candidates, Based on non-adjacent affine candidates, or Regression-based affine candidates.

3. The method according to any one of claims 1 to 2, wherein the plurality of affine candidates comprises a first affine candidate and a second affine candidate, and at least one of the following elements associated with the second affine candidate is compared with a corresponding at least one element associated with the first affine candidate during the checking process: Inter-frame direction, Prediction direction, Affine types, Sub-block Merge type, Bidirectional prediction (BCW) index with codec-level weights, Local Illumination Compensation (LIC) logo, Reference Index, Motion vector, The horizontal component of the motion vector, The vertical component of the motion vector, at least one control point motion vector (CPMV), The horizontal displacement between the first CPMV and the second CPMV, a vertical displacement between the first CPMV and the second CPMV, a horizontal displacement between the first CPMV and the third CPMV, or A vertical displacement between the first CPMV and the third CPMV. 4 . The method according to claim 3 , wherein the first CPMV is a top-left CPMV, or the second CPMV is a top-right CPMV, or the third CPMV is a bottom-left CPMV. 5 . The method according to claim 3 , wherein in the deduplication check it is checked whether the second affine candidate is identical to the first affine candidate.

6. The method of claim 5, wherein if the at least one element associated with the second affine candidate is the same as the corresponding at least one element associated with the first affine candidate, the second affine candidate is not added to the affine candidate list for the current video block. 7 . The method according to claim 3 , wherein similarity between the first affine candidate and the second affine candidate is checked in the deduplication check.

8. The method of claim 7, wherein if the at least one element associated with the second affine candidate is similar to the corresponding at least one element associated with the first affine candidate, the second affine candidate is not added to the affine candidate list for the current video block.

9. A method according to claim 8, wherein if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than a threshold, then the first element associated with the second affine candidate is similar to the first element associated with the first affine candidate.

10. The method of claim 8, wherein the first element associated with the second affine candidate is similar to the first element associated with the first affine candidate if the absolute difference between the first element associated with the second affine candidate and the first element associated with the first affine candidate is less than or equal to a threshold.

11. The method of any one of claims 9 to 10, wherein the threshold value depends on the size of the current video block.

12. The method according to any one of claims 9 to 11, wherein the threshold is determined based on at least one of: the width of the current video block, or The height of the current video block.

13. The method of claim 11, wherein if the size of the current video block is smaller than the size of another video block of the video, the threshold for the current video block is smaller than the threshold for the other video block, and If the size of the current video block is larger than the size of the other video block, then the threshold for the current video block is larger than the threshold for the other video block.

14. The method according to any one of claims 9 to 10, wherein the threshold value is predefined and fixed.

15. A method according to any one of claims 3 to 14, wherein the similarity between the horizontal component of the motion vector for the first affine candidate and the horizontal component of the motion vector for the second affine candidate is checked in the deduplication check, and the similarity between the vertical component of the motion vector for the first affine candidate and the vertical component of the motion vector for the second affine candidate is checked in the deduplication check.

16. A method according to any one of claims 3 to 15, wherein the similarity between the horizontal component of the CPMV for the first affine candidate and the horizontal component of the CPMV for the second affine candidate is checked in the deduplication check, and the similarity between the vertical component of the CPMV for the first affine candidate and the vertical component of the CPMV for the second affine candidate is checked in the deduplication check.

17. The method according to any one of claims 3 to 16, wherein whether the deduplication check is applied to at least one of the following depends on the affine type of the first affine candidate and the affine type of the second affine candidate: the horizontal displacement between the first CPMV and the third CPMV, or The vertical displacement between the first CPMV and the third CPMV.

18. The method of claim 17, wherein if the affine type of the first affine candidate and the affine type of the second affine candidate are 6-parameter affines, the deduplication check is applied to at least one of the following: the horizontal displacement between the first CPMV and the third CPMV, or The vertical displacement between the first CPMV and the third CPMV.

19. A method for video processing, include: For conversion between a current video block of a video and a bitstream of the video, obtaining information about applying a fusion-based codec tool to the current video block, the information depending on a video type of the current video block; as well as The converting is performed based on the information.

20. The method of claim 19, wherein the information includes at least one of the following: Whether the fusion-based codec tool is disabled for the current video block, or How to apply the fusion-based coding tool to the current video block.

21. The method according to any one of claims 19 to 20, wherein the fusion-based codec tool comprises at least one of the following: Overlapped Block Motion Compensation (OBMC), OMBC Template Matching(TM), Combined Inter and Intra Prediction (CIIP), CIIP Position Dependent Intra Prediction Combination (PDPC), CIIP TM, CIIP Template-based Intra Mode Derivation (TIMD), TIMD, Decoder-side intra mode derivation (DIMD), Multiple Hypothesis Prediction (MHP), Decoder-side motion vector refinement (DMVR), Inter-frame template matching (interTM), Cross-component adaptive loop filtering (CCALF), or Cross Component Sample Adaptive Offset (CCSAO).

22. The method of any one of claims 19 to 21, wherein if the video type of the current video block is screen content video, the fusion-based codec tool is disabled for the current video block.

23. A method according to any one of claims 19 to 22, wherein the information further depends on the grade, level or layer.

24. The method of any one of claims 19 to 23, wherein the fusion-based codec tool is disabled for at least one of: Video sequence, Picture Group, Pictures; or Strips.

25. The method of any one of claims 19 to 24, wherein the bitstream constraint indicates that the fusion-based codec tool is disabled for at least one of: Video sequence, Picture Group, Pictures; or Strips.

26. The method of any one of claims 19 to 25, wherein a syntax element in the bitstream indicates that the fusion-based codec tool is disabled for at least one of: Video sequence, Picture Group, Pictures; or Strips.

27. The method according to any one of claims 19 to 26, wherein a syntax element indicating the information is included in the bitstream.

28. The method of claim 27, wherein the syntax element is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.

29. The method according to any one of claims 27 to 28, wherein the syntax element is indicated in one of the following: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.

30. The method according to any one of claims 19 to 29, wherein the information is indicated by a single syntax element in the bitstream and the fusion-based codec tool comprises more than one codec tool.

31. A method according to any one of claims 19 to 30, wherein the fusion-based codec tool includes a first codec tool, multiple predictions for a video block are fused in the first codec tool to generate a target prediction for the video block, the first codec tool is applied to the current video block, and multiple sets of weights are allowed to fuse multiple predictions for the current video block.

32. The method of claim 31, wherein information about which set of weights is used to fuse the multiple predictions for the current video block depends on a content type of the current video block.

33. The method according to any one of claims 31 to 32, wherein the multiple sets of weights include a first set of weights and a second set of weights, the first set of weights being used to fuse the multiple predictions for the current video block, and the second set of weights being used to fuse multiple predictions for another video block different from the current video block.

34. The method of claim 33, wherein a content type of the current video block is different than a content type of the another video block.

35. The method according to any one of claims 33 to 34, wherein the current video block belongs to a video sequence captured by a camera, and the another video block belongs to a screen content video sequence, or The current video block belongs to a screen content video sequence, and the other video block belongs to a video sequence captured by a camera.

36. The method of any one of claims 33 to 35, wherein the multiple predictions for the current video block include two predictions, and the target prediction generated for the current video block is the same as one of the two predictions.

37. The method of claim 36, wherein the weights in the first set of weights are equal to 0 or 1.

38. The method of any one of claims 33 to 37, wherein the multiple predictions for the other video block include two predictions, and the target prediction generated for the other video block is a mixture of the two predictions.

39. The method of claim 38, wherein a weight in the second set of weights is equal to a fraction between 0 and 1.

40. The method of any one of claims 31 to 39, wherein the target prediction generated for the current video block is fused with another video block encoded using another codec different from the first codec.

41. The method of any one of claims 31 to 40, wherein the first codec tool comprises one of: CIIP, GPM, MHP, OBMC, TIMD, or DIMD.

42. A method for video processing, include: For conversion between a current video block of a video and a bitstream of the video, selecting at least one Karhuning-Love transform (KLT) kernel from a plurality of KLT kernels based on codec information of the current video block; as well as The transforming is performed based on the at least one KLT core.

43. The method according to claim 42, wherein the codec information comprises at least one of the following: Codec mode, size, Motion vector, Quantization Parameter (QP), or Time domain layer.

44. The method according to any one of claims 42 to 43, wherein the codec information comprises at least one of the following: whether a sub-block transform (SBT) is applied on the current video block, whether implicit multiple transform selection (MTS) is applied to the current video block, whether explicit MTS is applied to the current video block, Whether intra-MTS is applied to the current video block, Whether inter-frame MTS is applied to the current video block, whether a low frequency non-separable transform (LFNST) is applied to the current video block, whether intra block copy (IBC) is applied on the current video block, whether a palette mode (PLT) is applied to the current video block, whether intra prediction mode is applied to the current video block, whether intra sub-partitioning (ISP) is applied on the current video block, whether matrix-weighted intra prediction (MIP) is applied to the current video block, whether DIMD is applied to the current video block, whether TIMD is applied to the current video block, whether a linear model (LM) is applied to the current video block, whether a cross-component linear model (CCLM) is applied to the current video block, whether a convolutional cross-component model (CCCM) is applied on the current video block, whether a gradient-based linear model (GLM) is applied to the current video block, whether an inter-prediction mode is applied to the current video block, Whether AMVP is applied to the current video block, Whether the Merge mode is applied to the current video block, whether inter-frame template matching is applied to the current video block, whether intra template matching is applied to the current video block, Whether IBC template matching is applied on the current video block, whether DMVR is applied to the current video block, whether a sub-block prediction mode is applied to the current video block, whether affine mode is applied to the current video block, whether sub-block based temporal motion vector prediction (SbTMVP) is applied to the current video block, whether a multiple hypothesis mode is applied to the current video block, whether the multi-hypothesis mode includes an intra-frame coding and decoding part, whether a geometric partitioning mode (GPM) is applied to the current video block, Whether CIIP is applied to the current video block, whether MHP is applied on the current video block, Whether OBMC is applied to the current video block, or Whether LIC is applied to the current video block.

45. The method of any one of claims 42 to 44, wherein different KLT kernels are applied on video blocks of different sizes.

46. ​​The method according to any one of claims 42 to 45, wherein the codec information includes a size, and the at least one KLT core is selected based on whether at least one of the following satisfies a predefined condition: the width of the current video block; or The height of the current video block.

47. The method of claim 46, wherein the predefined condition comprises at least one of the following: The width of the current video block is smaller than a first value, the width of the current video block is less than or equal to the first value, The width of the current video block is greater than a second value, the width of the current video block is greater than or equal to the second value, The height of the current video block is less than a third value, The height of the current video block is less than or equal to the third value, The height of the current video block is greater than a fourth value, The height of the current video block is greater than or equal to the fourth value, a first result of dividing the width of the current video block by the height of the current video block is less than a fifth value, the first result is less than or equal to the fifth value, the first result is greater than a sixth value, the first result is greater than or equal to the sixth value, a second result of dividing the height of the current video block by the width of the current video block is less than a seventh value, the second result is less than or equal to the seventh value, the second result is greater than the eighth value, the second result is greater than or equal to the eighth value, The width of the current video block is equal to a ninth value, or The height of the current video block is equal to a tenth value.

48. The method according to claim 47, wherein the first value is 8 or 16 or 32 or 64, or The second value is 2 or 4 or 8, or The third value is 8 or 16 or 32 or 64, or The fourth value is 2 or 4 or 8, or The fifth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or The sixth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or The seventh value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or The eighth value is 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16, or The ninth value is 8 or 16 or 32 or 64, or The tenth value is 8 or 16 or 32 or 64.

49. The method of any one of claims 42 to 48, wherein different KLT cores are enabled for different video blocks of the video depending on their coding modes or sizes.

50. The method of any one of claims 42 to 49, wherein a first set of KLT cores are used for video blocks encoded using a first set of codec modes, and a second set of KLT cores are used for video blocks encoded using a second set of codec modes.

51. The method of claim 50, wherein the first group of KLT cores includes at least one type of KLT core, or the second group of KLT cores includes at least one type of KLT core.

52. The method of any one of claims 50 to 51, wherein the first set of coding modes comprises at least one type of prediction mode, at least one type of transform mode or at least one type of filter mode, or The second group of encoding and decoding modes includes at least one type of prediction mode, at least one type of transform mode or at least one type of filter mode.

53. The method of any one of claims 42 to 48, wherein the same KLT core is used for video blocks encoded using more than one different type of codec mode.

54. The method of any one of claims 42 to 48, wherein different KLT cores are used for video blocks encoded using different types of codec modes.

55. The method according to any one of claims 53 to 54, wherein the coding mode comprises a prediction mode, a transform mode or a filter mode.

56. A method according to any one of claims 53 to 55, wherein a type of codec mode include: Sub-block based prediction mode, Affine-based prediction model, Based on the fusion prediction model, SBT, ISP, or IBC.

57. The method of any one of claims 42 to 46, wherein the KLT core is divisible or indivisible.

58. The method of any one of claims 42 to 57, wherein the plurality of KLT cores comprises more than two KLT cores.

59. The method of any one of claims 42 to 58, wherein the at least one KLT core is enabled for at least one of: The main transformation, or Secondary transformation.

60. The method of any one of claims 42 to 59, wherein the current video block is encoded using a first codec mode, and different KLT kernels are used for horizontal transform and vertical transform of the current video block.

61. The method of any one of claims 42 to 59, wherein the current video block is encoded using a first codec mode, and the same KLT kernel is used for horizontal transform and vertical transform of the current video block.

62. The method according to any one of claims 42 to 61, wherein the current video block is encoded using a first codec mode and only KLT type transforms are allowed for the current video block.

63. The method according to any one of claims 60 to 62, wherein the first codec mode comprises SBT or MIP.

64. The method of claim 60, wherein the first encoding / decoding mode comprises at least one of the following: an SBT horizontal partition mode, an SBT vertical partition mode, an SBT half partition mode, or an SBT four partition mode.

65. The method according to any one of claims 42 to 64, wherein the predetermined transform type is replaced by a KLT based transform type.

66. The method of claim 65, wherein the predetermined transformation type comprises one of: Discrete Cosine Transform Type 2 (DCT2), Discrete Sine Transform Type 7 (DST7), or Discrete sine transform type 8 (DST8).

67. The method of any one of claims 42 to 66, wherein the at least one KLT kernel is enabled for chroma components of the current video block.

68. The method of any one of claims 42 to 67, wherein different KLT kernels are used for luma and chroma components of the current video block.

69. The method of any one of Claims 42 to 67, wherein the same KLT kernel is used for all color components of the current video block.

70. The method of any one of claims 42 to 67, wherein different KLT kernels are used for a chrominance blue (Cb) component and a chrominance red (Cr) component of the current video block.

71. The method of any one of claims 42 to 67, wherein the same KLT kernel is used for a Cb component and a Cr component of the current video block.

72. The method according to any one of claims 42 to 71, wherein a pair of a first KLT kernel and a second KLT kernel are used for horizontal transformation and vertical transformation of the current video block, and the transformation coefficient matrix of the second KLT kernel is the transposed matrix of the transformation coefficient matrix of the first KLT kernel.

73. The method of any one of claims 42 to 72, wherein a horizontal transform type or a vertical transform type for the current video block is selected from KLT-X, flip KLT-X, DCT-Y, and DST-Z, and each of X, Y, and Z is a constant.

74. The method of claim 73, wherein the transform coefficient matrix of the flipped KLT kernel is a transposed matrix of the transform coefficient matrix of the KLT kernel.

75. The method of any one of claims 73 to 74, wherein more than one KLT kernel is defined, each of the more than one KLT kernel is denoted as KLT-X, and X is equal to an integer value.

76. The method of any one of claims 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a horizontal transform of the current video block, and a flipped KLT kernel is used for a vertical transform of the current video block.

77. The method of any one of claims 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a vertical transform of the current video block, and a flipped KLT kernel is used for a horizontal transform of the current video block.

78. The method of any one of claims 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a horizontal transform of the current video block, and a non-KLT kernel is used for a vertical transform of the current video block.

79. The method of any one of claims 73 to 75, wherein the current video block is a transform block, a KLT kernel is used for a vertical transform of the current video block, and a non-KLT kernel is used for a horizontal transform of the current video block.

80. The method of any one of claims 78 to 79, wherein the non-KLT kernel is a DCT2 kernel.

81. The method of any one of claims 73 to 74, wherein a DST7 kernel is used for horizontal transform of the current video block and a DCT2 kernel is used for vertical transform of the current video block, or wherein a DCT2 core is used for the horizontal transform of the current video block, and a DCT7 core is used for the vertical transform of the current video block, or wherein a DCT8 core is used for the horizontal transform of the current video block, and a DCT2 core is used for the vertical transform of the current video block, or A DCT2 core is used for the horizontal transform of the current video block, and a DCT8 core is used for the vertical transform of the current video block.

82. The method of any one of claims 42 to 81, wherein the current video block is encoded using a specific prediction mode, a specific transform mode, or a specific filter mode.

83. A method according to any one of claims 42 to 82, wherein in addition to a predetermined MTS option, a KLT based transform type is explicitly signalled.

84. The method of any one of claims 42 to 83, wherein the codec information includes a motion vector, and the at least one KLT kernel is selected based on a magnitude of the motion vector for the current video block.

85. The method according to any one of claims 42 to 84, wherein the codec information comprises one of the following: The base QP determined at a syntax level above the slice level, a base QP signaled at a syntax level higher than the slice level, Strip QP, or The QP of the current video block.

86. The method according to any one of claims 42 to 85, wherein the codec information includes a time domain layer, and the at least one KLT core is selected based on whether the time domain layer of the current video block is a predetermined time domain layer.

87. The method of any one of claims 42 to 86, wherein the current video block is a residual block, and information about how to encode and decode the current video block depends on the at least one KLT kernel.

88. A method according to any one of claims 1 to 87, wherein whether and / or how to apply the method is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.

89. The method of any one of claims 1 to 88, wherein whether and / or how to apply the method is indicated in one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.

90. A method according to any one of claims 1 to 87, wherein whether and / or how to apply the method is indicated at an area comprising more than one sample or pixel.

91. The method of claim 90, wherein the region comprises at least one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Virtual Pipeline Data Unit (VPDU), Codec Tree Unit (CTU), CTU line, Strips, piece, or Sub-picture.

92. A method according to any one of claims 1 to 87, wherein whether and / or how the method is applied depends on coded information.

93. The method of claim 92, wherein the encoded information comprises at least one of: Block size, Color format, Single and double tree partitioning, Dual tree partitioning, Color component, Strip type, or Image type.

94. The method of any one of claims 1 to 93, wherein the converting comprises encoding the current video block into the bitstream.

95. The method of any one of claims 1 to 93, wherein the converting comprises decoding the current video block from the bitstream.

96. A device for video processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of claims 1 to 95.

97. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 95.

98. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by an apparatus for video processing, wherein the method include: Obtaining a plurality of affine candidates for a current video block of the video; applying a deduplication check to the plurality of affine candidates according to a checking process, the checking process being common to a plurality of different types of affine candidates; as well as The bitstream is generated based on the application.

99. A method for storing a bit stream of a video, include: Obtaining a plurality of affine candidates for a current video block of the video; applying a deduplication check to the plurality of affine candidates according to a checking process, the checking process being common to a plurality of different types of affine candidates; generating the bitstream based on the application; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

100. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by an apparatus for video processing, wherein the method include: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on a video type of the current video block; as well as The bitstream is generated based on the information.

101. A method for storing a bit stream of a video, include: obtaining information about applying a fusion-based codec tool to a current video block of the video, the information depending on a video type of the current video block; generating the bitstream based on the information; as well as The bit stream is stored in a non-transitory computer-readable recording medium.

102. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by an apparatus for video processing, wherein the method include: Selecting at least one Karhuning-Love transform (KLT) kernel from a plurality of KLT kernels based on codec information of a current video block of the video; as well as The bitstream is generated based on the at least one KLT core.

103. A method for storing a bit stream of a video, include: Selecting at least one Karhuning-Love transform (KLT) kernel from a plurality of KLT kernels based on codec information of a current video block of the video; generating the bitstream based on the at least one KLT core; as well as The bit stream is stored in a non-transitory computer-readable recording medium.