Method and device for video processing and medium
By integrating intra-template matching prediction mode with other codec tools, the codec performance of video codec is improved, and the problem of improving codec efficiency in the existing technology is solved, and it is suitable for HEVC and VVC standards.
Patent Information
- Application Number
- CN202480006568.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-05
- Filing Date
- 2024-01-04
- Publication Date
- 2025-08-12
AI Technical Summary
The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in the application of intra prediction mode, which makes it difficult to further improve the encoding and decoding performance.
By fusing the intra-template matching prediction mode with other codec tools, the prediction or reconstruction of video units can be derived to improve codec efficiency.
Improves the encoding and codec performance and efficiency of intra-template matching prediction, and is suitable for existing video encoding and codec standards such as HEVC and VVC.
Smart Images

Figure CN120476595A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to the fusion of intra-frame template matching predictions. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, further improvements in the encoding and decoding efficiency of video encoding and decoding technologies are generally desired. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: determining, for conversion between a video unit and a bitstream of the video unit, integrating an intra-frame template matching prediction (Intra-TMP) mode with a codec tool; deriving a prediction or reconstruction of the video unit based on the integration of the Intra-TMP mode with the codec tool; and performing conversion based on the prediction or reconstruction of the video unit. In this way, by integrating Intra-TMP with other codec tools, the codec performance and codec efficiency of Intra-TMP are improved.
[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a fusion of an intra-frame template matching prediction (intra-TMP) mode with a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode with the codec tool; and generating a bitstream based on the prediction or reconstruction of the video unit.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode and the codec tool; generating a bitstream based on the prediction or reconstruction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system according to some embodiments of the present disclosure is shown;
[0012] Figure 2 shows a block diagram of a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown;
[0014] Figure 4 An example of an encoder block diagram is shown;
[0015] Figure 5 67 intra prediction modes are shown;
[0016] Figure 6 The reference samples used for wide-angle intra prediction are shown;
[0017] Figure 7 The discontinuity problem is shown when the orientation exceeds 45°;
[0018] Figure 8 The MMVD search point is shown;
[0019] Figure 9 is a schematic diagram of the symmetric MVD pattern;
[0020] Figure 10 shows the extended CU area used in BDOF;
[0021] Figure 11 An affine motion model based on control points is shown;
[0022] Figure 12 The affine MVF of each sub-block is shown;
[0023] Figure 13 The positions of the inherited affine motion prediction values are shown;
[0024] Figure 14 Control point motion vector inheritance is shown;
[0025] Figure 15 The positions of candidate positions for the constructed affine Merge pattern are shown;
[0026] Figure 16 is a schematic diagram of the use of motion vectors for the proposed combination method;
[0027] Figure 17 The sub-block MVVSB and the pixel Δv(i, j) are shown;
[0028] Figure 18A shows the spatial neighboring blocks used by ATVMP;
[0029] Figure 18B Shows the derivation of sub-CU motion fields by applying motion shifts from spatial neighbors and scaling motion information from corresponding co-located sub-CUs;
[0030] Figure 19 Position lighting compensation is shown;
[0031] Figure 20 No downsampling for the short edges is shown;
[0032] Figure 21 Decoding side motion vector refinement is shown;
[0033] Figure 22 The diamond-shaped area in the search zone is shown;
[0034] Figure 23 Shows the location of the spatial merge candidate;
[0035] Figure 24 shows the candidate pairs considered for redundancy check of spatial Merge candidates;
[0036] Figure 25 It is a schematic diagram of motion vector scaling for time domain Merge candidates;
[0037] Figure 26 The candidate positions for the time domain Merge candidates C0 and C1 are shown;
[0038] Figure 27 Shows the VVC spatial neighboring blocks of the current block;
[0039] Figure 28 is a schematic diagram of the virtual blocks in the i-th round of search;
[0040] Figure 29 An example of GPM partitioning grouped at the same angle is shown;
[0041] Figure 30 Unidirectional prediction MV selection for geometric partitioning mode is shown;
[0042] Figure 31 shows an exemplary generation of warp weights w0 using a geometric segmentation pattern;
[0043] Figure 32 The figure shows the spatial neighboring blocks used to derive spatial Merge candidates;
[0044] Figure 33 shows the template matching performed on the search area around the initial MV;
[0045] Figure 34 It is a schematic diagram of the sub-blocks of OBMC application;
[0046] Figure 35 The SBT position, type and transformation type are shown;
[0047] Figure 36 shows the neighboring sample points used to calculate the SAD;
[0048] Figure 37 shows the neighboring samples used to calculate the SAD for sub-CU level motion information;
[0049] Figure 38 The sorting process is shown;
[0050] Figure 39 shows the recording process in the encoder;
[0051] Figure 40 shows the reordering process in the decoder;
[0052] Figure 41 is a schematic diagram of the extended reference area;
[0053] Figure 42 shows the IBC reference area depending on the current CU position;
[0054] Figure 43 An example of symmetry in a picture of screen content is shown;
[0055] Figure 44A It is a schematic diagram of BV adjustment for horizontal flip;
[0056] Figure 44B It is a schematic diagram of BV adjustment for vertical flip;
[0057] Figure 45 The intra-frame template matching search area used is shown;
[0058] Figure 46 is a schematic diagram of the template region;
[0059] Figure 47 The spatial portion of the convolution filter is shown;
[0060] Figure 48 shows the reference region (and its filling) used to derive the filter coefficients;
[0061] Figure 49 Four Sobel-based gradient modes for GLM are shown;
[0062] Figure 50 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and
[0063] Figure 51 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0064] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION
[0065] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0066] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0067] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, whether or not explicitly described, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in relation to other embodiments.
[0068] It should be understood that although the terms "first" and "second" and the like can be used to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0069] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises," "includes," and / or "having" when used herein indicate the presence of stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0070] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0071] Video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0072] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0073] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0074] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0075] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0076] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0077] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0078] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, wherein at least one reference picture is a picture in which the current video block is located.
[0079] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0080] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0081] The mode selection unit 203 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0082] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0083] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are not dependent on macroblocks in the same picture.
[0084] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0085] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0086] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0087] In one example, motion estimation unit 204 may indicate to video decoder 300 a value in a syntax structure associated with the current video block that indicates the current video block has the same motion information as another video block.
[0088] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0089] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0090] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0091] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0092] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0093] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.
[0094] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0095] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0096] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0097] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0098] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0099] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0100] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0101] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-encoded video data, which includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from adjacent blocks in the spatial or temporal domain.
[0102] The motion compensation unit 302 may generate a motion compensated block, and interpolation may be performed based on an interpolation filter. An identifier of an interpolation filter to be used with sub-pixel precision may be included in a syntax element.
[0103] The motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 may use the interpolation filters to generate a prediction block.
[0104] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the coded video sequence, partition information describing how each macroblock of the picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information for decoding the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0105] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0106] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0107] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to only that section. In addition, although specific embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video codec technologies. In addition, although some embodiments describe the video coding and decoding steps in detail, it should be understood that the corresponding decoding steps of the undo codec will be implemented by the decoder. In addition, the term "video processing" includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding techniques. Specifically, it relates to intra-frame template matching prediction and its integration with other codec tools, as well as other codec tools in image / video codecs. It can be applied to existing video codec standards such as HEVC or Versatile Video Codec (VVC). It can also be applied to future video codec standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations worked together to develop the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, targeting a 50% bitrate reduction compared to HEVC. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are investigating potential requirements for the standardization of future video codecs with compression capabilities significantly exceeding those of the current VVC standard. Such future standardization efforts could take the form of additional extensions to VVC or entirely new standards. These groups are jointly conducting this exploratory activity through a joint collaborative effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. New codec features and coding methods implemented in the Enhanced Compression Model (ECM) software are being explored in a coordinated manner by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG as potential enhancements to the video codec beyond the capabilities of VVC. 2.1. Encoding and decoding flow of typical video codecs Figure 4 An example of a VVC encoder block diagram is shown, which contains three loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original and reconstructed samples by adding offset and applying a finite impulse response (FIR) filter, respectively, and using the encoded side information to signal the offset and filter coefficients. ALF is located at the last processing stage of each picture and can be seen as a tool that attempts to capture and repair artifacts caused by previous stages. Intra-mode codec with 67 intra-prediction modes In order to capture arbitrary edge directions present in natural videos, such as Figure 5 As shown, the number of directional intra modes is extended from 33 used in HEVC to 65, while planar and DC modes remain unchanged. These more dense directional intra prediction modes are applicable to all block sizes and both luma and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value for non-square blocks. 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. Conventional angular intra prediction directions are defined as going from 45 degrees to -135 degrees clockwise. In VVC, for non-square blocks, multiple conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced mode is signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged at 67, and the intra mode encoding and decoding method also remains unchanged. To support these predictions, Figure 6 As shown, a top reference with a length of 2W+1 and a left reference with a length of 2H+1 are defined. The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra-frame prediction modes are shown in Table 2-1. Table 2-1 - Intra-frame prediction modes replaced by wide-angle mode like Figure 7As shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap Δpα. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted by these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the non-fractional mode in the conventional prediction mode with the wide-angle mode. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the entries of the mapping table to more accurately convert the prediction angles for chroma blocks. 2.3. Inter-frame prediction For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, as well as new codec features of VVC for additional information required for inter-prediction sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified, where the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and the additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, and other required information are explicitly signaled for each CU. 2.4. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is in integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra or inter prediction mode. The IBC mode is applicable to CUs with a width and height less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4×4 sub-block. For larger-sized current blocks, a hash key is determined to match the hash key of a reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At CU level, IBC mode is signaled with a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial, HMVP and pairwise candidates. - IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the upper neighbor (if encoded with IBC). When either neighbor is unavailable, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.5. IBC Movement Candidates The term "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB, or a video processing unit including multiple samples / pixels. A block may be rectangular or non-rectangular. For IBC-coded blocks, a block vector (BV) is used to indicate the displacement from the current block to a reference block that has been reconstructed within the current picture. W and H are the width and height of the current block (eg, luma block). The non-adjacent spatial candidates of the current codec block are the adjacent spatial candidates of the virtual block in the i-th round search (such as Figure 9 For the i-th search round, the width and height of the virtual block are calculated using the following formulas: newWidth = i × 2 × gridX + W, newHeight = i × 2 × gridY + H. Obviously, if search round i is 0, the virtual block is the current block. In the following, the BV prediction value is also the BV candidate. The skip mode is also the merge mode. Based on some criteria, BV candidates can be divided into several groups. Each group is called a subgroup. For example, we can use adjacent spatial and temporal BV candidates as the first subgroup and the remaining BV candidates as the second subgroup. In another example, we can use the first N (N ≥ 2) BV candidates as the first subgroup, the next M (M ≥ 2) BV candidates as the second subgroup, and the remaining BV candidates as the third subgroup. 2.6. Merge Mode with MVD (MMVD) In addition to the Merge mode (in which the implicitly derived motion information is directly used for prediction sample generation of the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. Immediately after sending the regular Merge flag, the MMVD flag is signaled to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, the merge candidate is further refined using signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The MMVD candidate flag is signaled to specify which of the first and second merge candidates to use. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 8As shown in Table 2-2, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2. Table 2-2 - Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 2-3. It should be noted that the meaning of the MVD symbol can change depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV (where both lists point to the same side of the current picture, that is, the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 2-3 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV (where both MVs point to different sides of the current picture, that is, the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbol in Table 2-3 specifies the sign of the MV offset added to the list0 MV component of the starting MV, and the sign for listlMV has an opposite value. Otherwise, if the POC difference in list 1 is greater than the POC difference in list 0, then the sign in Table 2-3 specifies the sign of the MV offset added to the list 1 MV component of the starting MV, and the sign for list 0 MV has the opposite value. The MVD is scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb, as Figure 26 If the POC difference of L1 is greater than the POC difference of L0, the MVD of list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MVs. Table 2-3 - Signs of MV offsets specified by direction index 2.7. Symmetric MVD Encoding and Decoding In VVC, in addition to the normal unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional prediction MVD signaling is also applied. In symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not transmitted through the signal but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: - If mvd_11_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, a symmetric mode flag is explicitly signaled to indicate whether the symmetric mode is used. When the symmetric mode flag is true, only mvp_10_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown in the following formula. In the encoder, symmetric MVD motion estimation starts with initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for symmetric MVD motion search. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, formerly known as BIO, is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and multiplier size. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: - the CU is coded using "true" bi-prediction mode, i.e., one of the two reference pictures precedes the current picture in display order, and the other follows the current picture in display order; - the distances from the two reference pictures to the current picture (i.e., POC differences) are the same; - Both reference images are short-term reference images; -CU is not encoded or decoded using affine mode or SbTMVP Merge mode; -CU has more than 64 luma samples; -CU height and CU width are both greater than or equal to 8 luma samples; - BCW weight index indicates equal weight; -Do not enable WP for the current CU; -CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ). Motion refinement is then used to adjust the bidirectional prediction sample values in the 4×4 sub-block. The following steps are applied in the BDOF process. First, the horizontal gradient and vertical gradient of the two prediction signals and k = 0, 1 is calculated by directly calculating the difference between two adjacent sample points, that is, Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k (k=0, 1), and shift1 is calculated as shift1=max(6, bitDepth-6) based on the luma bit depth bitDepth. Then, the autocorrelations and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as S1=∑ (i,j)∈Ω Abs(ψ x (i, j)), S3 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ x (i, j)) (2-3) S5=∑ (i,j)∈Ω Abs(ψ y (i, j)), S6 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ y (i, j)) in θ(i, j)=(I (1) (i, j)>>n b )-(I (0) (i, j)>>n b ) where Ω is a 6×6 window around the 4×4 sub-block, and n a and n b The values of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8), respectively. Then use the following formula, motion refinement (v x , v y ) is derived using the cross-correlation and autocorrelation terms: in th′ BIO =2 max(5,BD-) . is a floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4×4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o offset )>>shift (2-7). These values are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, it is necessary to generate some predicted sample points I in the list k (k = 0, 1) outside the current CU boundary (k) (i, j). Figure 10 As shown, BDOF in VVC uses an extended row / column around the CU boundary. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly obtaining the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. In order to avoid the additional complexity of the SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded in symmetric MVD mode or CIIP mode. 2.9. Joint Inter- and Intra-frame Prediction (CIIP) 2.10. Affine Motion Compensated Prediction In HEVC, only the translation motion model is applied to motion compensated prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensated prediction is applied. Figure 11 As shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: Among them (mv 0x , mv 0y ) is the motion vector of the upper left control point, (mv 1x , mv 1y ) is the motion vector of the upper right control point, and (mv 2x , mv 2y ) is the motion vector of the lower left control point. In order to simplify the motion compensation prediction, block-based affine transformation prediction is applied. In order to derive the motion vector of each 4×4 luminance sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation (such as Figure 12 The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the four corresponding 4×4 luminance subblocks. Like translational motion inter prediction, there are two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.10.1. Affine Merge Prediction AF_MERGE mode can be applied to CUs with width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatially adjacent CUs. There can be up to five CPMVP candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPVM candidates are used to form the affine Merge candidate list: - Inherited affine merge candidates inferred from the CPMV of neighboring CUs; -Affine Merge candidate CPMVP constructed using the translation MV of the neighboring CU; -Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Figure 13 As shown. For the left prediction value, the scanning order is A0->A1, and for the upper prediction value, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown in the figure, if the neighboring lower left block A is encoded and decoded in affine mode, the motion vectors v2, v3 and v4 of the upper left, upper right and lower left corners of the CU containing block A are obtained. When block A is encoded and decoded with a 4-parameter affine model, the two CPMVs of the current CU are calculated based on v2 and v3. When block A is encoded and decoded with a 6-parameter affine model, the three CPMVs of the current CU are calculated based on v2, v3 and v4. The constructed affine candidate means that the candidate is constructed by combining the neighboring translation motion information of each control point. The motion information of the control point is obtained from Figure 15 The specified spatial and temporal nearest neighbors are derived as shown. k(k=1, 2, 3, 4) represents the kth control point. For CPMV1, check B2->B3->A2 blocks and use the MV of the first available block. For CPMV2, check B1->B0 blocks, and for CPMV3, check A1->A0 blocks. If available, TMVP is used as CPMV4. After obtaining the MVs of the four control points, the affine merge candidate is constructed based on the motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}. The combination of 3 CPMVs constructs a 6-parameter affine Merge candidate, and the combination of 2 CPMVs constructs a 4-parameter affine Merge candidate. To avoid the motion scaling process, the relevant combination of control point MVs is discarded if the reference indices of the control points are different. After the inherited affine merge candidates and the constructed affine merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list. 2.10.2. Affine AMVP Prediction Affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. A CU-level affine flag is signaled in the bitstream to indicate whether affine AMVP mode is used, and another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2 and is generated by using the following four types of CPVM candidates in order: - Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs; -Affine AMVP candidate CPMVP constructed using the translation MV of the neighboring CU; -Translated MV from neighboring CU; -Zero MV. The order in which inherited affine AMVP candidates are checked is the same as the order in which inherited affine Merge candidates are checked. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as the current block are considered. When the inherited affine motion predictor is inserted into the candidate list, the deduplication process is not applied. The AMVP candidates are constructed from Figure 15The specified spatial neighbor derivation is shown. The same check order is used as in the affine Merge candidate construction. In addition, the reference picture index of the neighboring blocks is also checked. The first block in the check order that is inter-coded and has the same reference picture as the current CU is used. There is only one. When the current CU is coded with a 4-parameter affine mode and both mv0 and mv1 are available, they are added as a candidate in the affine AMVP list. When the current CU is coded with a 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list candidate is still less than 2 after the inherited affine AMVP candidate and the constructed AMVP candidate are checked, then mv0, mv1, and mv2 will be added as translation MVs in order when available to predict all control point MVs of the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.10.3. Affine motion information storage In VVC, the CPMV of an affine CU is stored in a separate buffer. The stored CPMV is only used to generate the inherited CPMVP in affine Merge mode and affine AMVP mode for the most recently coded CU. The sub-block MV derived from the CPMV is used for motion compensation, MV derivation of the Merge / AMVP list of translation MVs, and deblocking. To avoid picture row cache for additional CPMV, the affine motion data inheritance of CUs from upper CTUs is handled differently from the inheritance from regular neighboring CUs. If the candidate CU for affine motion data inheritance is in the upper CTU row, the bottom left and bottom right sub-block MVs in the row cache are used for affine MVP derivation instead of the CPMVs. In this way, the CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to a 4-parameter model. Figure 16 As shown, along the top CTU boundary, the bottom left and bottom right sub-block motion vectors of the CU are used for affine inheritance of the CU in the bottom CTU. 2.10.4. Prediction Refinement Using Optical Flow for Affine Mode Compared with pixel-based motion compensation, sub-block based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the expense of prediction accuracy loss. In order to achieve finer-grained motion compensation, prediction refinement using optical flow (PROF) is used to refine sub-block based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block based affine motion compensation is performed, the luminance prediction samples are refined by adding the difference derived from the optical flow equation. PROF is described as the following four steps: Step 1) Sub-block based affine motion compensation is performed to generate a sub-block prediction I(i, j). Step 2) Use a 3-tap filter [-1, 0, 1] to calculate the spatial gradient g of the sub-block prediction at each sample point x (i, j) and g y (i, j). The gradient calculation is exactly the same as that in BDOF. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) (2-10) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) (2-11) Shift1 is used to control the accuracy of the gradient. The sub-block (i.e. 4×4) prediction is extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extension boundary are copied from the nearest integer pixel position in the reference picture. Step 3) The brightness prediction refinement is calculated by the following optical flow equation. ΔI(i, j)=g x (i, j)*Δv x (i, j)+g y (i, j)*Δv y (i, j) (2-12) Where Δv(i, j) is the difference between the sample point MV (denoted by v(i, j)) calculated for the sample point position (i, j) and the sub-block MV of the sub-block to which the sample point (i, j) belongs, as Figure 17 As shown in FIG. Δv(i, j) is quantized in units of 1 / 32 luminance sample accuracy. Since the affine model parameters and the sample position relative to the sub-block center do not change from one sub-block to another, Δv(i, j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i, j) and dy(i, j) be the distance from the sample position (i, j) to the center of the sub-block (xSB ,y SB )’s horizontal and vertical offsets, Δv(x, y), can be derived by the following equations: To maintain accuracy, the center of the sub-block (x SB ,y SB ) is calculated as ((W SB -1) / 2, (H SB -1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For a 4-parameter affine model, For the 6-parameter affine model, Where (v 0x , v 0y )、(v 1x , v 1y )、(v 2x , v 2y ) are the control point motion vectors of the upper left, upper right, and lower left, and w and h are the width and height of the CU. Step 4) Finally, the luma prediction refinement ΔI(i, j) is added to the sub-block prediction I(i, j). The final prediction I' is generated as the following equation. I′(i,j)=I(i,j)+ΔI(i,j)(2-17) PROF is not applied to affine-coded CUs in two cases: 1) all control point MVs are the same, indicating that the CU has only translational motion; 2) the affine motion parameters are greater than the specified limit, because sub-block-based affine MC is downgraded to CU-based MC to avoid large memory access bandwidth requirements. A fast coding method is applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) If the CU is not a root block and its parent block does not select the affine mode as its best mode, PROF is not applied because the possibility of the current CU selecting the affine mode as the best mode is low; b) If the magnitudes of the four affine parameters (C, D, E, F) are all less than a predefined threshold and the current picture is not a low-latency picture, PROF is not applied because the improvement introduced by PROF in this case is small. In this way, affine motion estimation using PROF can be accelerated. 2.11. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports sub-block-based temporal motion vector prediction (SbTMVP). Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the CU in the current picture. The same co-located pictures used by TMVP are used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: -TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; -While TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right or center block relative to the current CU), SbTMVP applies motion shifting before obtaining the temporal motion information from the co-located picture, where the motion shifting is obtained from the motion vector of one of the spatial neighboring blocks from the current CU. The SbTVMP process is as follows Figure 18A and Figure 18B As shown. SbTMVP predicts the motion vector of the sub-CU in the current CU in two steps. In the first step, check Figure 18A The spatial neighbor A1 in is selected. If A1 has a motion vector that uses the co-located picture as its reference picture, then that motion vector is selected as the motion shift to be applied. If no such motion is identified, then the motion shift is set to (0, 0). In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to move from Figure 18B The co-located picture shown obtains motion information (motion vector and reference index) at the sub-CU level. Figure 18B The example in assumes that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block in the co-located picture (the minimum motion grid covering the center sample) is used to derive the motion information for the sub-CU. After the motion information of the co-located sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a sub-block based Merge list containing both SbTVMP candidates and affine Merge candidates is used for signaling of the sub-block based Merge mode. The SbTVMP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block based Merge candidates, followed by the affine Merge candidate. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8×8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8. The encoding logic of the additional SbTMVP Merge candidate is the same as that of other Merge candidates, that is, for each CU in a P slice or a B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.12. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is signaled in units of quarter-luminance samples. In VVC, the CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows: -Normal AMVP mode: quarter brightness samples, half brightness samples, integer brightness samples, or four brightness samples. -Affine AMVP mode: quarter-luminance samples, integer-luminance samples, or 1 / 16-luminance samples. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal and vertical MVD for reference list L0 and reference list L1) are zero, a quarter luma sample MVD resolution is inferred. For CUs with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required, and quarter-luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half-luma samples or other MVD precision (integer or quad luma samples) are used for normal AMVP CUs. In the case of half-luma samples, a 6-tap interpolation filter is used for half-luma sample positions instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma sample MVD precision or quad luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, a second flag is used to indicate whether integer-luma sample MVD precision or 1 / 16 luma sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma sample, half-luma sample, integer-luma sample, or quad luma sample), the CU's motion vector prediction value is rounded to the same precision as the MVD before being added to the MVD. Motion vector prediction values are rounded towards zero (ie, negative motion vector prediction values are rounded towards positive infinity, and positive motion vector prediction values are rounded towards negative infinity). The encoder uses RD check to determine the motion vector resolution for the current CU. In order to avoid always performing four CU-level RD checks for each MVD resolution, in VTMl1, the RD check of MVD precision other than quarter luma samples is only called under conditional circumstances. For normal AVMP mode, the RD cost of quarter luma sample MVD precision and integer luma sample MV precision is first calculated. Then, the RD cost of integer luma sample MVD precision is compared with the RD cost of quarter luma sample MVD precision to determine whether it is necessary to further check the RD cost of four luma sample MVD precision. When the RD cost of quarter luma sample MVD precision is much smaller than the RD cost of integer luma sample MVD precision, the RD check of four luma sample MVD precision is skipped. Then, if the RD cost of integer luma sample MVD precision is significantly greater than the best RD cost of the previously tested MVD precision, the check of half luma sample MVD precision is skipped. For the affine AMVP mode, if the affine inter mode is not selected after checking the rate-distortion cost of the affine merge / skip mode, merge / skip mode, quarter luma sample MVD accuracy normal AMVP mode, and quarter luma sample MVD accuracy affine AMVP mode, the 1 / 16 luma sample MV accuracy and 1 pixel MV accuracy affine inter modes are not checked. In addition, in the 1 / 16 luma sample and quarter luma sample MV accuracy affine inter modes, the affine parameters obtained in the quarter luma sample MV accuracy affine inter mode are used as the starting search point. 2.13. Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-18) Five weights, w∈{-2, 3, 4, 5, 10}, are allowed in weighted average bidirectional prediction. For each bidirectionally predicted CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3, 4, 5}) are used. -At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing encoder complexity. These algorithms are summarized below. For more details, please refer to the VTM software. When combined with AMVR, if the current picture is a low-latency picture, unequal weights are conditionally checked for 1-pixel and 4-pixel motion vector accuracy. - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. - Conditionally check for unequal weights when the two reference pictures in bidirectional prediction are the same. - Based on the POC distance between the current picture and its reference pictures, codec QP, and temporal level, do not search for unequal weights when certain conditions are met. The BCW weight index is encoded using a context codec bit followed by a bypass codec bit. The first context codec bit indicates whether equal weights are used; if unequal weights are used, an additional bit is signaled using bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with cross-fading. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid the interaction between WP and BCW, which would complicate the VVC decoder design, if the CU uses WP, the BCW weight index is not signaled, and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded or decoded in CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.14. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a codec tool used to address the problem of local illumination variations between a current picture and its temporal reference picture. LIC is based on a linear model where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x, y) = α·P r (x+v x , y+v y )+β Where P(x, y) is the prediction signal of the current block at the coordinate (x, y); P r (x+v x , y+v y ) is the motion vector (v x , v y ) points to the reference block; α and β are the corresponding scaling factors and offsets applied to the reference block. Figure 19 The LIC process is shown in Figure 19 In , when LIC is applied to a block, the minimum mean square error (LMSE) method is adopted to minimize the neighboring samples of the current block (i.e., Figure 19 The template T in the temporal reference picture) and its corresponding reference sample in the temporal reference picture (ie, Figure 19In addition, in order to reduce the computational complexity, both the template samples and the reference template samples are downsampled (adaptive downsampling) to derive the LIC parameters, i.e., only Figure 19 The shaded points in are used to derive α and β. In order to improve the encoding and decoding performance, such as Figure 20 As shown, no downsampling is performed on the short edges. 2.15. Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of MV in Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In bidirectional prediction operation, the refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and list L1. Figure 21 As shown, the SAD between two blocks based on each MV candidate (eg, MV0' and MV1') around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, the application of DMVR is restricted and is only applied to CUs coded with the following modes and features: -CU-level Merge mode with bi-directional prediction MV - Relative to the current picture, one reference picture is in the past and the other reference picture is in the future - The distances from the two reference pictures to the current picture (i.e., POC differences) are the same - Both reference images are short-term reference images -CU has more than 64 luma samples -CU height and CU width are both greater than or equal to 8 luminance samples -BCW weight index indicates equal weight -Do not enable WP for the current block -CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction in future picture codecs. The original MV is used in the deblocking process and is also used for spatial motion vector prediction in future CU codecs. Additional features of DMVR are mentioned in the following sub-items. 2.15.1. Search Scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) obeys the following two equations: MV0′=MV0+MV_offset (2-19) MV1′=MV1-MV_offset (2-20) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples away from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainty, a bias towards the original MV during the DMVR process is proposed. The SAD between the reference blocks referenced by the initial MV candidate is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation rather than an additional search with SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates with the center having the minimum SAD in either the first or second iteration of the search. In the sub-pixel offset estimation based on the parametric error surface, the center position cost and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form E(x, y) = A(x min ) 2 +B(yy min ) 2 +C (2-21) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-22) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-23). x min and y min The value of is automatically constrained to be between -8 and 8, since all cost values are positive and the minimum is E(0,0). This corresponds to a half-pixel shift with 1 / 16 pixel MV precision. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to get the sub-pixel accurate refinement increment MV. 2.15.2. Bilinear interpolation and sample filling In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the initial fractional pixel MV with integer sample offsets, so samples at those fractional positions need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, with a 2-sample search range, DVMR does not access more reference samples than the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV are filled in from these available samples. 2.15.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16×16. 2.16. Multi-pass decoder-side motion vector refinement In this contribution, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16×16 sub-block within the codec block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction. 2.16.1. First pass - Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MV based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive integer sample precision intDeltaMV and half-pixel sample precision halfDeltaMv. The local search applies a 3×3 square search pattern to cycle through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. The existing fractional sample refinement is further applied to derive the final deltaMV. Then the refined MV after the first pass is derived as: MV0_pass1=MV0+deltaMV MV1_pass1=MV1-deltaMV. 2.16.2. Second pass - Sub-block based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to the 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) for the reference picture lists L0 and L1 obtained by the first pass. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each subblock, BM performs a full search to derive integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated by applying the cost factor to the SATD cost between two reference sub-blocks as follows: bilCost = satdCost * costFactor. The search area (2*sHor+1)*(2*sVer+1) is divided into Figure 22 There are up to five diamond-shaped search regions shown. Each search region is assigned a costFactor, which is determined by the distance between each search point and the starting MV (intDeltaMV), and each diamond-shaped region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order, starting from the upper left corner of the region and ending at the lower right corner. When the minimum bilCost within the current search region is less than a threshold (which is equal to sbW*sbH), the integer-pixel full search is terminated. Otherwise, the integer-pixel full search continues to the next search region until all search points have been checked. BM performs a local search to derive the half-sample accuracy halfDeltaMv. The search pattern and cost function are the same as those defined in 2.9.1. The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). Then the refined MV at the second pass is derived as: ·MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2) ·MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2). 2.16.3. Third pass - sub-block-based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy without clipping from the refined MV of the parent block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample accuracy and clipped between -32 and 32. The refined MVs at the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as: ·MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv ·MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv. 2.17. Sample-based BDOF In sample-based BDOF, instead of deriving the motion refinement (Vx, Vy) on a block basis, it is performed for each sample. The codec block is divided into 8×8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks with a threshold. If BDOF is applied to a sub-block, a sliding 5×5 window is used for each sample in the sub-block, and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectionally predicted sample value for the center sample of the window. 2.18. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: (1) Spatial MVP from the adjacent CU (2) Temporal MVP from the same CU (3) History-based MVP from FIFO table (4) Paired Average MVP (5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU coded in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded using context, and bypass coding is used for the other binary bits. This section provides the derivation process of various types of Merge candidates. As done in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a certain size area. 2.18.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates located at the positions shown in the figure, a maximum of four Merge candidates are selected. The derivation order is B0, A0, B1, A1 and B2. Position B2 is only considered when one or more CUs at positions B0, A0, B1 and A1 are not available (for example, because it belongs to another slice or piece) or is intra-coded. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding and decoding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 24 The pairs are linked by arrows in , and a candidate is added to the list only if the corresponding candidates used for redundancy check do not have the same motion information. 2.18.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Figure 25 As shown by the dotted line in , the scaled motion vector of the temporal merge candidate is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to 0. like Figure 26 As shown, the position for the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. 2.18.3. History-Based Merge Candidate Derivation History-based MVP (HMVP) Merge candidates are added to the Merge list after the spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (cleared). As long as there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find whether the same HMVP exists in the table. If found, the same HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. Redundancy check is applied to HMVP candidates and spatial or temporal Merge candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: The number of HMPV candidates used for Merge list generation is set to (N<=4) × M: (8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of HMVP candidates available in the table. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process from HMVP is terminated. 2.18.4. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the number represents the merge index of the merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, the list remains invalid. When the Merge list is not full after pairwise average Merge candidates are added, a zero MVP will be inserted at the end until the maximum number of Merge candidates is reached. 2.18.5.Merge Estimation Region Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). Candidate blocks located in the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. In addition, the update process of the history-based motion vector prediction candidate list is only updated when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel), where (xCb, yCb) is the top left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.19. New Merge Candidates 2.19.1. Non-adjacent Merge Candidate Derivation In VVC, Figure 27 The five spatial neighboring blocks and one temporal neighbor are shown to be used to derive the Merge candidate. It is proposed to use the same pattern as in VVC to derive additional Merge candidates from positions non-adjacent to the current block. To achieve this, for each round of search i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated using the following formula: Offsetx=-i×gridX, Offsety=-i×gridY Where Offsetx and Offsety represent the offset of the top left corner of the virtual block relative to the top left corner of the current block, and gridX and gridY are the width and height of the search grid. Second, the width and height of the virtual block are calculated using the following formula: newWidth=i×2×gridX+currWidth newHeight=i×2×gridY+currHeight. Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block. gridX and gridY are currently set to currWidth and currHeight respectively. Figure 28 The relationship between the virtual block and the current block is shown. After generating the virtual block, block A i 、B i 、C i 、D i and E i The VVC spatial neighbors of the virtual block can be considered, and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i It is the spatial neighboring block used in VVC Merge mode. When building the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbors are utilized. Non-adjacent spatial domain Merge candidates are inserted into the Merge list after the time domain Merge candidates in the order of B1->A1->C1->D1->E1. 2.19.2.STMVP It is proposed to use three spatial domain Merge candidates and one temporal domain Merge candidate to derive the average candidate as the STMVP candidate. The STMVP is inserted before the upper left spatial merge candidate. The STMVP candidate is deduplicated along with all previous merge candidates in the merge list. For spatial candidates, the first three candidates in the current Merge candidate list are used. For temporal candidates, the same positions as VTM / HEVC co-location are used. For spatial candidates, the first candidate, the second candidate, and the third candidate inserted into the current Merge candidate list before STMVP are denoted as F, S, and T. The temporal candidate having the same position as the VTM / HEVC co-location used in TMVP is denoted as Col. The motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of the four Merge candidates are all valid in the prediction direction X (X=0 or 1) and are all equal to 0, then mvLX=(mvLX_F+mvLX_S+mvLX_T+mvLX_Col)>>2. 2) If the reference indexes of three of the four Merge candidates are valid in the prediction direction X (X=0 or 1) and equal to 0, then mvLX=(mvLX_F×3+mvLX_S×3+mvLX_Col×2)>>3 or mvLX=(mvLX_F×3+mvLX_T×3+mvLX_Col×2)>>3 or mvLX=(mvLX_S×3+mvLX_T×3+mvLX_Col×2)>>3. 3) If the reference indexes of two of the four Merge candidates are valid in the prediction direction X (X=0 or 1) and equal to 0, then mvLX=(mvLX_F+mvLX_Col)>>1 or mvLX=(mvLX_S+mvLX_Col)>>1 or mvLX=(mvLX_T+mvLX_Col)>>1. NOTE: If time domain candidates are not available, STMVP mode is turned off. 2.19.3.Merge List Size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8. 2.20. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode for inter-frame prediction is supported. The geometric partitioning mode is signaled as a Merge mode using a CU level flag. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. The geometric partitioning mode is used for each possible CU size (w×h=2 m ×2 n , where m, n ∈ {3…6} (excluding 8x64 and 64x8) supports a total of 64 splits. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Figure 29 ). The position of the partition line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-frame predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion compensated predictions are required per CU. The unidirectional prediction motion for each partition is derived using the process described in 2.20.1. If geometric partitioning mode is used for the current CU, a geometric partitioning index and two Merge indices (one for each partition) indicating the partitioning mode (angle and offset) of the geometric partitioning are further transmitted by signal. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, the sample values along the geometric partitioning edges are adjusted using a blending process with adaptive weights as in 2.20.2. This is a prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored as described in 2.20.3. 2.20.1. One-way prediction candidate list construction The unidirectional prediction candidate list is directly derived from the merge candidate list constructed according to the extended merge prediction process in 2.18. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are Figure 30 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partition mode. 2.20.2. Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. The blending weight for each position of the CU is derived based on the distance between the individual position and the partition edge. The distance from the segmentation edge for a position (x, y) is derived as: where i, j are the indices of the angle and offset for the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j The sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-28) w1(x,y)=1-w0(x,y) (2-30). partIdx depends on the angle index i. An example of weight w0 is Figure 31 is shown in . 2.20.3. Motion Field Storage for Geometric Partitioning Mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combination Mv of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2: (motionIdx≤0?(1-partIdx):partIdx) (2-31) Where motionIdx is equal to d(4x+2, 4y+2), which is recalculated from equation (2-18). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion My2 is stored. 2.21. Multi-hypothesis prediction In Multi-Hypothesis Prediction (MHP), in addition to Inter AMVP, Normal Merge, Affine Merge, and MMVD modes, up to two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to Table 2-4 below. Table 2-4-MHP weighting factors add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. Additional assumptions can be Merge or AMVP mode. In the case of Merge mode, motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric partitioning mode. In the case of AMVP mode, the reference index, MVP index, and MVD are transmitted through the signal. 2.22. Non-adjacent airspace candidates Non-adjacent spatial merge candidates are inserted after TMVP in the regular merge candidate list. The spatial merge candidate pattern is as follows: Figure 32 The distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. Template Matching (TM) Template Matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in a reference picture (i.e., the same size as the template). Figure 33 As shown in Figure 2, a better MV is searched around the initial motion of the current CU within the search range of [-8, +8] pixels. Two modified template matching methods are proposed: the search step size is determined based on the AMVR mode, and the TM can be cascaded with the bilateral matching process in the Merge mode. In AMVP mode, MVP candidates are determined based on template matching error to pick one MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM performs MV refinement only on that specific MVP candidate. TM refines the MVP candidate by using an iterative diamond search, starting with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode), and then using half-pixel and quarter-pixel in sequence according to the AMVR mode specified in Table 2-5. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 2-5 - Search Styles for AMVR and Merge Mode with AMVR In Merge mode, a similar search method is applied to the merge candidates indicated by the merge index. As shown in Table 2-5, TM can be performed all the way down to 1 / 8 pixel MVD accuracy, or skip those accuracies exceeding half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). In addition, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check. 2.24. Overlapped Block Motion Compensation (OBMC) Overlapped Block Motion Compensation (OBMC) has been previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed on all motion compensated (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it is applied to both luminance and chrominance components. In JEM, an MC block corresponds to a codec block. When a CU is encoded and decoded with sub-CU modes (including Sub-CU Merge, Affine, and FRUC modes), each sub-block of the CU is an MC block. In order to handle CU boundaries in a unified manner, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set equal to 4×4, as shown in Figure 34 shown. When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block. The prediction block based on the motion vectors of the neighboring sub-blocks is denoted as P N , where N indicates the index for the adjacent upper, lower, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is denoted as P C When P N When the motion information of the neighboring sub-block contains the same motion information as the current sub-block, OBMC is not used from P N Execute. Otherwise, P N Each sample point is added to P C The same point in the same N Four rows / columns are added to P C The weight factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and weight factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P CThe exception is small MC blocks (ie, when the height or width of the codec block is equal to 4 or the CU is coded in sub-CU mode), for which P N Only two rows / columns are added to P C In this case, weight factors {1 / 4, 1 / 8} are used for P N , and weight factors {3 / 4, 7 / 8} are used for P C For P generated based on the motion vectors of vertically (horizontally) adjacent sub-blocks N , P N The samples in the same row (column) of P are added to P with the same weighting factor. C . In JEM, for CUs with a size less than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 luma samples or not encoded in AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is considered during the motion estimation stage. The prediction signal formed by OBMC, using the motion information of the top and left neighboring blocks, is used to compensate for the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied. 2.25. Multiple Transform Selection (MTS) for Core Transformations In addition to the DCT-II used in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding for both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the selected DST / DCT basis functions. Table 2-6 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain orthogonality of the transform matrix, the transform matrix is quantized more accurately than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients must have 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is applied only to luma. MTS signaling is skipped when one of the following conditions applies: - the position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC); -The last significant coefficient of the luminance TB lies within the MTS null region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally transmitted by signal to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 2-7. By removing the intra mode and block shape dependencies, a unified transform selection for ISP and implicit MTS is used. If the current block is in ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are turned on, only DST7 is used for both the horizontal transform core and the vertical transform core. When it comes to transform matrix accuracy, an 8-bit main transform core is used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform cores (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8) use an 8-bit main transform core. Table 2-7 - Transformation and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, high-frequency transform coefficients are zeroed for DST-7 blocks and DCT-8 blocks with size (width or height, or both) equal to 32. Only coefficients in the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. When MTS is enabled for an inter-coded block, implicit MTS can also be enabled. 2.26. Sub-Block Transform (SBT) In VTM, sub-block transform is introduced for inter-frame predicted CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When the inter-frame predicted CU has cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded and decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded using the inferred adaptive transform, and the other part of the residual block is zeroed. When SBT is used for inter-coded CU, SBT type and SBT location information are signaled in the bitstream. Figure 35 As shown, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is like a binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is like an asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 in luminance samples, the 1:3 / 3:1 partitioning along that dimension is prohibited. There are up to 8 SBT modes for a CU. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are selected in Figure 35 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DCT-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slice, cbf, and horizontal and vertical core transform types of the residual block. SBT is not applied to CUs coded in inter-intra combination mode. 2.27. Adaptive Merge Candidate Reordering Based on Template Matching To improve encoding and decoding efficiency, after building the merge candidate list, the order of each merge candidate is adjusted based on the template matching cost. Merge candidates are arranged in the list in ascending order of template matching cost. They are operated on in subgroups. The template matching cost is measured by the SAD (Sum of Absolute Difference) between the neighboring samples of the current CU and its corresponding reference samples. If the Merge candidate includes bidirectionally predicted motion information, then Figure 36 As shown in , the corresponding reference sample is the average of the corresponding reference sample in reference list 0 and the corresponding reference sample in reference list 1. If the Merge candidate includes sub-CU level motion information, then Figure 37 As shown, the corresponding reference sample is composed of the neighboring samples of the corresponding reference sub-block. like Figure 38 As shown, the sorting process is performed in the form of subgroups. The first three merge candidates are sorted together. The next three merge candidates are sorted together. The template size (width of the left template or height of the upper template) is 1. The subgroup size is 3. 2.28. Adaptive Merge Candidate List It can be assumed that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the following 3 Merge candidates are taken as the second subgroup (ie, the last subgroup). For the encoder, after building the Merge candidate list, such as Figure 39 As shown, some merge candidates are adaptively reordered in ascending order of merge candidate cost. More specifically, the template matching costs of the Merge candidates in all subgroups except the last subgroup are calculated; then, except for the last subgroup, the Merge candidates in the own subgroup are reordered; finally, the final Merge candidate list is obtained. For the decoder, after building the Merge candidate list, such as Figure 40 As shown in , some / no Merge candidates are adaptively reordered in ascending order of Merge candidate cost. Figure 40 In , the subgroup where the selected (signaled) Merge candidate is located is called the selected subgroup. More specifically, if the selected Merge candidate is located in the last subgroup, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed, and the Merge candidate list is not changed; otherwise, the execution process is as follows: After all Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching costs of the Merge candidates in the selected subgroup are calculated; the Merge candidates in the selected subgroup are reordered; and finally, a new Merge candidate list is obtained. For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template. When deriving the reference samples of the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision. The reference samples (RT) of the template used for bidirectional prediction are derived by taking a weighted average of the reference samples (RT0) of the template in reference list 0 and the reference samples (RT1) of the template in reference list 1 as follows. RT=((8-w)*RT0+w*RT1+4)>>3 (2-32) The weights (8-w) of the reference templates in reference list 0 and the weights (w) of the reference templates in reference list 1 are determined by the BCW index of the merge candidate. BCW indices equal to {0, 1, 2, 3, 4} correspond to w equal to {-2, 3, 4, 5, 10}, respectively. If the local illumination compensation (LIC) flag of the Merge candidate is true, the LIC method is used to derive the reference samples of the template. The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT. The template size is 1. This means that the width of the left template and / or the height of the upper template is 1. If the codec mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered. If the coding mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered. 2.29. IBC with extended reference area An IBC reference area design that does not increase the current storage area required by ECM-3 is proposed and its performance is tested. Figure 41 The design is shown in Figure 1. In the figure, the blue square represents the current CTU and the green square represents the CTU that can be used by IBC reference. Specifically, assuming that W represents the maximum horizontal CTU index and the current CTU index is (m, n), for the codec unit in the current CTU, the CTUs with indices (0, n)...(m, n) and (m-1, n)...(W, n) define the reference region that can be used by IBC. One reason for such a design is that in the current ECM, the CTUs to the left, above, and above-left are being used and therefore need to be preserved. To achieve this, all CTUs to the right of the upper CTU in the upper CTU row (for CTUs to be encoded in the current CTU row) and all CTUs to the left of the current CTU in the current CTU row (for CTUs to be encoded in the next CTU row) must be preserved. This means that such a design does not increase the buffer size required for the current ECM. 2.30. IBC with Template Matching It is proposed to use template matching with IBC for both IBC Merge mode and IBCAMVP mode. Compared to the list used by the regular IBC Merge mode, the IBC-TM Merge list has been modified so that candidates are selected based on a deduplication method with the same motion distance between candidates as in the regular TM Merge mode. The ending zero motion satisfy (which is meaningless for intra codecs) has been replaced by the motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H) CUs, and then the list is satisfied with the one on the left without deduplication if necessary. In IBC-TM Merge mode, the selected candidates are refined using template matching methods before the RDO or decoding process.The IBC-TM Merge mode has been competed with the regular IBC Merge mode, and the TM-Merge flag is signaled. In IBC-TM AMVP mode, up to three candidates are selected from the IBC Merge list. Each of these three selected candidates is refined using a template matching method and ranked according to their resulting template matching cost. Then, typically only the top two are considered in the motion estimation process. Since IBC motion vectors are constrained to be integers and are Figure 42 The reference region shown is within the template matching refinement for both IBC-TM Merge and AMVP modes, so the template matching refinement is quite simple. Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, while in IBC-TM AMVP mode, they are performed with integer precision or 4-pixel precision. In both cases, the refined motion vectors in each refinement step must obey the constraints of the reference region. 2.31. Reconstruction of Reordered IBC (RR-IBC) Screen content codecs such as Intra Block Copy (IBC) generate prediction blocks by directly copying previously coded reference areas in the same picture. Symmetry is often observed in video content, especially in text character areas and computer-generated graphics in screen content sequences, such as Figure 43 Therefore, a specific screen content codec that takes symmetry into account will effectively compress such video content. The Reconstruction Reordered Inter-Band Codec (RR-IBC) mode is proposed for screen content video coding and decoding. When applied, the samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped back to restore the original block. For blocks decoded by RR-IBC, two flipping methods are supported, namely horizontal flipping and vertical flipping. First, for blocks decoded by IBC AMVP, a syntax flag is transmitted by signaling to indicate whether the reconstruction is flipped, and if it is flipped, another flag is further transmitted by signaling to specify the flip type. For IBC Merge, the flip type is inherited from the adjacent block without syntax signaling. Taking into account horizontal symmetry or vertical symmetry, the current block and the reference block are usually aligned horizontally or vertically. Therefore, when horizontal flipping is applied, the vertical component of BV is not transmitted by signaling and is presumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of BV is not transmitted by signaling and is presumed to be equal to 0. In order to better exploit the symmetry, a flip-aware BV adjustment method is applied to refine the block vector candidates. Figure 44A and Figure 44B As shown, (x nbr ,y nbr ) and (x cur ,y cur ) represent the coordinates of the center sample points of the neighboring blocks and the current block, BV nbr and BV cur Represents the BV of the neighboring block and the current block respectively. In the case where the neighboring block is horizontally flipped, BV cur The horizontal component of is not inherited directly from the neighboring block BV, but by nbr The horizontal component (expressed as BV nbr h ) adds motion shift to calculate, i.e., BV cur h =2(x nbr -x cur )+BV nbr h Similarly, in the case where the adjacent block is vertically flipped, BV cur The vertical component of the nbr The vertical component (expressed as BV nbr v ) adds motion shift to calculate, i.e., BV cur v =2(y nbr -y cur )+BV nbr v . 2.32. Intra-frame template matching prediction Intra Template Matching Prediction (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. The encoder searches for the template most similar to the current template in the reconstructed portion of the current frame for a predefined search range and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is obtained by combining the L-shaped causal neighbors of the current block with Figure 45 is generated by matching another block in a predefined search area in the , the predefined search area including: R1: current CTU; R2: upper left CTU; R3: upper CTU; R4: left CTU. SAD is used as the cost function. In each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW SearchRange_h=a*BlkH Where "a" is a constant that controls the gain / complexity tradeoff. In practice, "a" is equal to 5. The intra template matching tool is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.33 Intra-frame prediction fusion Intra prediction fusion methods use multiple prediction values generated from different modes / reference lines. In subtest a, multiple intra prediction values are generated and then fused by weighted averaging. The process of deriving the prediction value to be used in the fusion process is described as follows: 1) For the angular intra prediction mode in the single mode case including TIMD and DIMD, the proposed method is implemented by transforming the angular intra prediction mode from the angular intra prediction mode represented as p fusion =w0p line +w1p line+ The intra prediction is derived by weighting the intra prediction obtained from multiple reference rows, where p line is intra prediction from the default reference line, and p line+ is the prediction from the row above the default reference row. The weights are set to w0=3 / 4 and w1=1 / 4. 2) For TIMD mode with mixing, p line is used in the first mode (w0=1, w1=0), and P line+ is used in the second mode (w0=0, w1=1). 3) For DIMD mode with hybrid, the number of prediction values selected for weighted averaging increases from 3 to 6. In subtest b, intra prediction fusion is performed on reference lines instead of prediction blocks. Two reference lines (called r line and r line+1) is used for intra prediction fusion. The corresponding intra prediction angle DeltaInt is considered in the fusion process. Each value (r fusion [i]) is derived from: r fusion [i]=(3·r line [i]+r line+1 [i+DeltaInt])>>2. When angular intra modes have non-integer slopes (required reference sample interpolation) and block size is greater than 16, the proposed intra prediction fusion is applied to luma blocks, which is used together with MRL, but not to ISP-coded blocks. In the method studied in subtest a, PDPC is applied for the intra prediction mode that uses the reference line closest to the current block. 2.34 Template-based Multi-reference Line Intra Prediction (TMRL) The proposed TMRL model includes the following aspects: a) Extended reference row candidate list and intra prediction mode candidate list. The extended reference row candidate list used in this proposal is {1, 3, 5, 7, 12}. The restriction on the top CTU row remains unchanged. The size of the intra prediction mode candidate list is 10. The construction of the intra prediction mode candidate list is similar to MPM. The differences are: Planar mode is excluded from the proposed intra prediction mode candidate list. The DC mode is added after the 5 neighboring PUs' mode and the DIMD mode if it is not already included. Angular modes with incremental angles from ±1 to ±4 (compared to existing angular modes in the intra prediction mode candidate list) are added. b) Construction of TMRL candidate list. For a block, there are 5×10=50 combinations of extended reference lines and allowed intra prediction modes. Since the extended reference lines start from reference line 1, the area covered by reference line 0 is used for template matching. For the template area (see Figure 46 ) is calculated between the prediction (generated by the 50 combinations) and the reconstruction. The 20 combinations with the smallest SAD costs are selected in ascending order to form the TMRL candidate list. c) TMRL signaling Instead of directly encoding and decoding the reference line and intra mode, an index into the TMRL candidate list is encoded to indicate which combination of reference line and prediction mode is used to encode and decode the current block. In the proposed TMRL mode, a truncated Golomb-Rice codec with a divisor of 4 is used to encode and decode the selected combination from the combination list. The binarization process and codewords are shown in Table 2-8. Table 2-8-TMRL index binarization process d) Modification on the encoder side Encoder side modifications are tested to further improve the codec efficiency. For intra blocks larger than 8×8, if no TMRL mode is selected by SATD comparison, an additional TMRL RDO is added. 2.35 Convolutional Cross-Component Model (CCCM) for Intra Prediction We propose to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit to what the current CCLM mode does. Like CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. In addition, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luminance reference value, and the other model is for the remaining samples (following the spirit of CCLM design). For PUs with at least 128 available reference samples, the multi-model CCCM mode can be selected. 2.35.1 Convolutional Filters The proposed convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a nonlinear term, and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) luminance sample co-located with the chrominance sample to be predicted and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 47 shown. The nonlinear term P is expressed as the square of the center luminance sample C and is scaled to the sample value range of the content: P=(C*C+midVal)>>bitDepth. That is, for 10-bit content, it is calculated as: P=( C*C+512)>>10. The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B. 2.35.2 Calculation of filter coefficients Filter coefficient c i It is calculated by minimizing the MSE between the predicted chrominance samples and the reconstructed chrominance samples in the reference region. Figure 48 A reference region consisting of six rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is required to support the "side samples" of the plus-shaped spatial filter and is filled in when not in the available region. MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-decomposed, and the final filter coefficients are calculated using inverse substitution. This process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer arithmetic. 2.35.3 Bitstream Signaling The use of this mode is signaled via a PU-level flag via the CABAC codec. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a submode of CCLM. That is, the CCCM flag is signaled only when the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM). 2.36 Gradient Linear Model (GLM) Compared to CCLM, GLM uses the gradient of luma samples to infer the linear model instead of the downsampled luma values. Specifically, when GLM is applied, the input of the CCLM process (i.e., the downsampled luma samples L) is replaced by the luma sample gradient G. The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β For signaling, when CCLM mode is enabled for the current CU, two flags are separately transmitted by signal for the Cb component and the Cr component to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further transmitted by signal to select one of the four gradient filters for gradient calculation. Enable four gradient filters for GLM, such as Figure 49 , which shows four Soble-based gradient modes for GLM. 3. Question In the current design of intra-frame TMP, the best prediction block is copied from the reconstructed part of the current frame, whose L-shaped template matches the current template. However, due to template inaccuracy in some cases, the copied prediction block may not always be selected after rate-distortion optimization. The codec performance of intra-frame TMP can be improved by fusing it with other codec tools (e.g., intra prediction). 4. Specific Implementation Methods The following embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow sense. In addition, these embodiments can be combined in any way. In the present disclosure, intra-frame TMP may not be limited to the current intra-frame TMP technology, but may be interpreted as a technology that obtains a reference (or prediction) block using samples in the current slice / slice / sub-picture / picture / other video unit (e.g., CTU row) in addition to conventional intra-frame prediction methods. In the following discussion, Intra TMP may be replaced by other codec tools that rely on coded / decoded / reconstructed information within the same region, eg, palette, intra block copy (IBC). Fusion of intra-frame template matching prediction and intra-frame prediction 1. We propose a fusion of intra-frame template matching prediction and codec tools to derive the prediction / reconstruction of video units. We denote the fusion method as IntraTMP_fusion mode. a. In one example, the codec tool may refer to inter-frame prediction, or palette, or IBC, or BDPCM. b. In one example, the codec tool may refer to intra prediction. i. In one example, intra prediction may refer to a conventional intra prediction method (e.g., intra prediction using 35 intra prediction modes in HEVC or 67 intra prediction modes in VVC) or other intra prediction methods that utilize samples in the current slice / slice / sub-picture / picture / other video units (e.g., CU, PU, TU, CTU, CTU row) other than the intra TMP to obtain a prediction block. ii. In one example, intra prediction may refer to DIMD, TIMD, ISP, MIP, MRL, PDPC / gradient PDPC, intra prediction fusion, and TMRL. iii. In one example, intra prediction may refer to cross-component prediction (CCLM), multi-model CCLM, left CCLM, top CCLM, CCCM, left / top CCCM, GLM, or variants thereof, etc. c. In one example, intra-TMP and more than one codec tool can be fused. d. In one example, intra TMP and codecs with more than one prediction signal can be fused. e. In one example, P(x, y) = w IP1 *IP1(x,y)+w IP2 *IP2(x,y)+...+w IPn *IP n (x, y) + w TMP1 *IntraTMP1(x,y)+w TMP2 *IntraTMP2(x, y)+...+w TMPm *IntraTMP m (x, y), where P(x, y) is the generated prediction value, IP k (x, y) is the prediction generated by the k-th intra prediction, IntraTMP j (x, y) is the prediction generated by the jth intra TMP, and w IPk and w TMPj is the corresponding weighted value. i. In one example, n=1 or 2 or 3 and m=1. ii. In one example, n=1 or 2 or 3 and m=2. iii. In one example, n=1 and m=1 or 2 or 3. iv. In one example, n=2 and m=1 or 2 or 3. v. In one example, n=0 and m=or 2 or 3 or 4 or 5. vi. In one example, n=2 or 3 or 4 or 5 and m=0. f. In one example, one or more intra TMP candidates may be used to generate an intra TMP prediction signal. i. In one example, the derivation of one or more intra TMP candidates can be the same as the derivation of the intra TMP. ii. Alternatively, the derivation of one or more intra TMP candidates may be different from the derivation of the intra TMP. 1) In one example, the current template used to derive the intra TMP candidate may be different. a) In one example, the shape / size of the current template may be different. b) In one example, the current template may be downsampled. c) In one example, one or more samples in the current template may be modified before being used to derive intra TMP candidates. i. In one example, the prediction signal of the current template can be used to modify the current template. The current template, the prediction signal of the current template, and the modified template are represented as T, T p and T'. 1. In one example, the prediction signal may be generated using the same intra prediction method as used to fuse the intra TMP for the current block. 2. In one example, T' = (a*Tb*T p ) / c. a. In one example, a=2, b=1, and c=1. b. In one example, a=4, b=3, and c=1. c. In one example, a=8, b=7, and c=1. d. In one example, a=16, b=15, and c=1. 3. In one example, T' = (a*Tb*T p +offset)>>shift. ii. In one example, the top left sample point of the current template may not be modified. 2) In one example, the search area, search region, or search method may be different. iii. In one example, more than one intra TMP candidate may be derived. 1) In one example, more than one intra TMP candidate can be derived from different search ranges. a) Alternatively, at least two intra TMP candidates may be derived from the same search range. 2) In one example, which intra TMP candidate is used for fusion may be predefined, signaled using syntax elements, or derived. 3) In one example, the number of intra-frame TMP candidates may be predefined, signaled, or derived. iv. In one example, template matching can be used to derive / refine one or more intra TMP candidates. g. In one example, one or more intra prediction modes (IPMs) may be used to generate an intra prediction signal. i. In one example, the IPM may be predefined, signaled, or derived. 1) In one example, the predefined IPM may refer to PLANAR mode, DC mode, horizontal mode, and vertical mode. ii. In one example, the block offset of the intra TMP candidate can be used to derive the IPM. iii. In one example, IPM can be derived using neighboring samples, such as DIMD and / or TIMD. 1) In one example, IPM is derived using DIMD and TIMD. a) In one example, the intra prediction mode derived using DIMD can be used in TIMD to derive the intra prediction mode for TIMD. iv. In one example, an IPM candidate list is constructed, and one or more IPMs in the list may be used to generate an intra prediction signal. 1) In one example, which IPM is used to generate the intra prediction signal may be predefined, or signaled using syntax elements, or derived. v. In one example, at least one codec tool may be different from traditional intra prediction. 1) In one example, the codec tool may refer to how to fill the reference samples, or whether and / or how to filter the reference samples, or whether and / or how to apply a filtering process (e.g., PDPC / gradient PDPC), or whether and / or how to use an interpolation filter. 2) Alternatively, the coding tools used to obtain the intra prediction signal can be the same as conventional intra prediction. h. In one example, one or more intra TMP candidates and / or one or more IPMs may be reordered before being used to generate an intra TMP prediction signal or an intra prediction signal. i. In one example, template matching or bilateral matching costs can be used for re-ranking. ii. In one example, reordering can be used for intra TMP candidates. iii. In one example, reordering can be used for IPM. iv. In one example, reordering can be used for the combination of intra TMP candidates and IPM. i. In one example, the intra prediction signal and / or the fused / final prediction signal may be refined through a filtering process. i. In one example, the filtering process may refer to PDPC or gradient PDPC. 2. In one example, the weighting parameters used to fuse the intra-frame TMP prediction signal and the intra-frame prediction signal may be predefined, signaled, or derived. a. In one example, the weighting parameters may be predefined. b. In one example, an intra prediction signal and an intra TMP prediction signal are used in IntraTMP_fusion. i. In one example, P(x, y) = w IP *IP(x,y)+w TMP *IntraTMP(x, y), where w IP +w TMP =1. ii. In one example, P(x, y) = (w IP *IP(x,y)+w TMP *IntraTMP(x, y)+offset)>>shift, where w IP +w TMP =(1< <shift)。 1) In one example, offset=0. 2) In one example, offset=1<<(shift-1). 3) In one example, w IP =1,shift=1. 4) In one example, w IP =1 / 2 / 3, shift=2. 5) In one example, w IP =1 / 2 / 3 / 4 / 5 / 6 / 7, shift=3. 6) In one example, w IP =1 / 2 / 3 / 4 / 5 / 6 / 7 / 8 / 9 / 10 / 11 / 12 / 13 / 14 / 15, shift=4. 7) In one example, shift = 5 / 6 / 7 / 8. c. In one example, weighting parameters may be transmitted via signal. i. In one example, a set of weighting parameters is constructed and an index indicating the weighting parameter may be signaled. d. In one example, weighting parameters can be derived using codec information. i. In one example, the codec information may refer to the codec mode of the neighboring unit. 1) In one example, the weighting parameters may depend on whether one or more neighboring units are coded using intra prediction or IBC mode. ii. In one example, the codec information may refer to an intra prediction mode used to obtain an intra prediction signal. iii. In one example, the codec information may refer to the block size or block dimension of the current video unit and / or the neighboring video units. iv. In one example, the weighting parameters may be derived using a template matching method (eg, with a minimum template matching cost). v. In one example, weighting parameters may be derived using codec information generated when searching for intra TMP candidates. 1) In one example, a search cost (eg, SAD or MRSAD) may be used. vi. In one example, a template matching based approach can be used to derive weighting parameters. vii. In one example, the weighting parameters may be derived using the LDL method or Gaussian elimination (eg, the method used in CCCM). 1) In one example, one or more samples in the current template and / or a reference of the current template indicated by an intra TMP candidate and / or a prediction signal of the current template may be used. e. In one example, the weighting parameters may depend on the video content. i. In one example, the weighting parameters may be different for natural sequences and screen content sequences. f. In one example, weighted values may be generated in the same or similar manner as a GPM or SGPM generates weighted values. i. In one example, at least one index or syntax element may be signaled to indicate which set of weight values to use. ii. In one example, without signaling which set of weight values to use, they are derived at the decoder. 3. In one example, the intra TMP prediction signal and the second prediction signal (such as the intra prediction signal) can be fused by directly combining the two predictions based on position. a. For at least one position, intra TMP prediction is applied, and for at least one position, second prediction is applied. 4. In one example, more than one fusion method can be applied to an intra-TMP. a. In one example, which fusion method is used can be predefined, signaled, or derived. 5. Whether and / or how to apply the IntraTMP_fusion mode for a video unit may depend on codec information, which may refer to: a. Whether to allow intra-frame TMP or intra-frame prediction method b. Block dimensions and / or block size c. Block depth d. Slice / picture type and / or partition tree type (single tree, dual tree, or partial dual tree) i. In one example, IntraTMP_fusion may only be applied to I slices / pictures. 1) Alternatively, IntraTMP_fusion can be applied to all slice / picture types. e. Time domain layer identification f. Block location g. Color component. 6. In one example, the codec information used in IntraTMP_fusion can be used to encode and decode subsequent video units. a. In one example, the block vector used to generate the intra TMP prediction signal can be regarded as the intra TMP and and / or block vector for normal IBC. i. In one example, the block vector may be added to the IBC HMVP table. ii. In one example, the block vector may be used to construct an IBC AMVP / Merge candidate list for subsequent video units. b. In one example, the IPM used to generate the intra prediction signal can be regarded as an IPM for normal intra prediction. i. In one example, the IPM can be used to construct the MPM list for subsequent video units. ii. In one example, IPM can be used for chroma prediction. iii. In one example, IPM can be propagated for non-intra-coded video units. c. Alternatively, the codec information may not be used for subsequent video units. i. In one example, instead of the IPM used for the current video unit, a predefined IPM (eg, DC or Planar) may be used for subsequent video units. 7. In one example, how to template match a block may depend on whether the block is to be fused via intra TMP and secondary prediction. a. In one example, the second prediction should be taken into account (such as subtracted before) when calculating the template cost. 8. In one example, whether and / or how IntraTMP_fusion is applied may depend on the color format and / or color components. a. In one example, IntraTMP_fusion can be applied to all color components. b. In one example, when IntraTMP_fusion is applied to chroma components, the derivation of intra prediction may be different from that for luma components. i. In one example, intra prediction can be obtained using CCLM or MMLM or CCCM or Chroma-DIMD or Chroma-TIMD or a fusion of CCLM / MMLM / CCCM with angular mode. c. In one example, whether and / or how IntraTMP_fusion is applied to the first component may depend on whether IntraTMP_fusion is applied to the second component. i. In one example, the first component may refer to a chrominance component (eg, Cb and / or Cr), and the second component may refer to a luma component (eg, Y). ii. In one example, the manner in which IntraTMP_fusion is applied to the first component may be the same as that of the second component. 1) Alternatively, the way IntraTMP_fusion is applied to the first component may be different from that to the second component. a) In one example, the weighting parameters may be different. d. In one example, IntraTMP_fusion may be applied to the luma component but not to the chroma components. i. In one example, the luma component may refer to Y in the YCbCr color space or G in the RGB color space. ii. In one example, the chrominance components may refer to Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space. IntraTMP_ f usion signaling 9. Indication of IntraTMP_fusion mode can be derived on the fly. 10. The indication of IntraTMP_fusion mode may be conditionally signaled, where the conditions may include: a. Whether to allow intra-frame TMP or intra-frame prediction method b. Block dimensions and / or block size c. Block depth d. Slice / picture type and / or partition tree type (single tree, dual tree, or partial dual tree) i. In one example, the indication of IntraTMP_fusion mode may be signaled only for I slices / pictures. 1) Alternatively, an indication of the IntraTMP_fusion mode may be signaled for all slice / picture types. e. Time domain layer identification f. Block location i. In one example, for the block located in the upper left corner of the slice / picture, the indication of IntraTMP_fusion mode is not signaled. g. Color component h. In one example, if the indication of IntraTMP_fusion mode is not signaled, it can be inferred as a default value. i. In one example, if the indication of IntraTMP_fusion mode is not signaled, it can be presumed to be false. i. In one example, if an indication of IntraTMP_fusion mode is not signaled, it may be presumed to be true. 11. Whether the current block is coded in IntraTMP_fusion mode may be signaled using one or more syntax elements. a. In one example, the syntax element may be binarized using a fixed length codec, a truncated unary codec, a unary codec, or an EG codec, or may be coded as a flag. b. In one example, syntax elements may be bypass coded or context coded. i. The context may depend on coded information such as block dimension and / or block size and / or slice / picture type and / or information of neighboring blocks (adjacent or non-adjacent) and / or information of other codec tools used for the current block and / or information of the temporal layer. c. In one example, when the current video unit is intra TMP codec, an indication of IntraTMP_fusion mode may be signaled. d. In one example, one or more syntax elements can be in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header are transmitted through signals. e. In one example, syntax elements may be coded in a predictive manner. f. In one example, syntax elements of the current block can be predicted from syntax elements of neighboring blocks. General Items 12. In the above examples, a video unit may refer to a color component / sub-picture / slice / slice / codec tree unit (CTU) / CTU row / CTU group / codec unit (CU) / prediction unit (PU) / transform unit (TU) / codec tree block (CTB) / codec block (CB) / prediction block (PB) / transform block (TB) / block / subblock of a block / subregion within a block / any other region containing more than one sample or pixel. 13. Whether and / or how to apply the methods disclosed above can be transmitted through signals at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 14. Whether and / or how to apply the above methods may depend on the following information: a. Messages transmitted by signaling in DPS / SPS / VPS / PPS / APS / picture header / slice header / slice group header / largest codec unit (LCU) / codec unit (CU) / LCU row / LCU group / TU / PU block / video codec unit b. Location of CU / PU / TU / block / video codec unit c. Block dimensions of the current block and / or its neighboring blocks d. Block shape of the current block and / or its neighboring blocks e. Block codec mode, for example, IBC or non-IBC inter-frame mode or non-IBC sub-block mode f. Indication of color format (such as 4:2:0, 4:4:4) g. Codec tree structure h. Slice / slice group type and / or picture type i. Color components (e.g., may be applied only to chroma components or luma components) j. Time domain layer ID k. Standard grade / level / tier.
[0108] As used herein, the term "video unit" or "video block" may be a sequence, a picture, a slice, a tile, a sub-picture, a codec tree unit (CTU) / codec tree block (CTB), a CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTUs / CTBs, one or more virtual pipeline data units (VPDUs), or a sub-region within a picture / slice / slice / tile. The term "reference row" may refer to a row and / or column of reconstructed samples that are adjacent or non-adjacent to a current block and are used to derive intra-prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by an intra-prediction mode (e.g., conventional intra-prediction with an intra-prediction mode), or by weighting reference samples of a reference row using a matrix or vector to derive intra-prediction of the current video unit (e.g., MIP).
[0109] Figure 50 FIG. 5 is a flow chart of a method 5000 for video processing according to an embodiment of the present disclosure. The method 5000 is implemented during conversion between a video unit of a video and a bitstream of the video.
[0110] At block 5010, a fusion of an intra-template matching prediction (Intra-TMP) mode and codec tools is determined for conversion between video units of a video and a bitstream of the video units. For example, intra-template matching prediction and codec tools may be fused to derive prediction / reconstruction of the video units. The fusion method may be denoted as IntraTMP_fusion mode.
[0111] At block 5020, prediction or reconstruction of a video unit is derived based on a fusion of an intra TMP mode and codec tools. In some embodiments, a video unit comprises at least one of the following: a color component, a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec unit (CU), a codec tree unit (CTU), a CTU row, a CTU group, a slice, a sub-picture, a block, a sub-region within a block, or a region containing more than one sample or pixel.
[0112] At block 5030, conversion is performed based on the prediction or reconstruction of the video unit. In some embodiments, conversion may include encoding the video unit into a bitstream. Alternatively or additionally, conversion may include decoding the video unit from the bitstream. In this way, the codec efficiency and codec performance of Intra-TMP can be improved by combining Intra-TMP with other codec tools.
[0113] In some embodiments, the codec tool is one of the following: inter prediction mode, palette mode, intra block copy (IBC) mode, or block-based incremental pulse coding modulation (BDPCM) mode. In some other embodiments, the codec tool is an intra prediction mode. For example, the intra prediction mode includes one of the following: a conventional intra prediction mode (e.g., intra prediction using 35 intra prediction modes or 67 intra prediction modes in HEVC) or other intra prediction modes that obtain a prediction block using samples in one of the following other than intra template matching prediction: a current slice, a current slice, a current sub-picture, a current picture (e.g., a CU, a PU, a TU, a CTU, a CTU row), or other video units.
[0114] In some embodiments, the intra prediction method includes one of the following: decoder-side intra mode derivation (DIMD), template-based intra mode derivation (TIMD), intra sub-partitioning (ISP), matrix-weighted intra prediction (MIP), multiple reference lines (MRL), position-dependent intra prediction combination (PDPC), gradient PDPC, intra prediction fusion, or template-based multiple reference line intra prediction (TMRL). In some other embodiments, the intra prediction method includes one of the following: cross-component linear model (CCLM), a variant of CCLM, multi-model CCLM, left CCLM, upper CCLM, convolutional cross-component model (CCCM), a variant of CCCM, left CCCM, upper CCCM, gradient linear model (GLM), or a variant of GLM.
[0115] In some embodiments, one or more intra TMP candidates are used to generate an intra TMP prediction signal. In some embodiments, the derivation of the one or more intra TMP candidates is the same as the derivation of the intra TMP.
[0116] Alternatively, the derivation of one or more intra TMP candidates is different from the derivation of the intra TMP, for example, the current template used to derive the intra TMP candidates is different.
[0117] In some embodiments, the current template is of a different shape or size. In some other embodiments, the current template is downsampled.
[0118] In some embodiments, one or more samples in the current template are modified before being used to derive an intra TMP candidate. For example, the prediction signal of the current template is used to modify the current template. For example, the current template, the prediction signal of the current template, and the modified template can be denoted as T, Tp, and T', respectively.
[0119] In some embodiments, the prediction signal is generated using the same intra prediction method as that used to fuse the intra TMP of the current block. For example, T'=(a*Tb*Tp) / c, where T' represents the modified template, T represents the current template, Tp represents the prediction signal, and a, b, and c are integers. In some embodiments, a=2, b=1, and c=1. In some other embodiments, a=4, b=3, and c=1. Alternatively, a=8, b=7, and c=1. As another example, a=16, b=15, and c=1. In some other embodiments, T'=(a*Tb*Tp+offset)>>shift, where T' represents the modified template, T represents the current template, Tp represents the prediction signal, shift is a parameter, and a, b, and c are integers. In some embodiments, the upper left sample of the current template is not modified. In some other embodiments, at least one of the following items for the derivation of one or more intra TMP candidates is different from the intra TMP: search area, search region, or search method.
[0120] In some embodiments, multiple intra TMP candidates are derived. In some embodiments, multiple intra TMP candidates are derived from different search regions. In some other embodiments, at least two intra TMP candidates are derived from the same search region.
[0121] In some embodiments, which intra TMP candidate is used for fusion is predefined. Alternatively, which intra TMP candidate is used for fusion is indicated using a syntax element. In some other embodiments, which intra TMP candidate is used for fusion is derived.
[0122] In some embodiments, the number of intra TMP candidates is predefined. Alternatively, the number of intra TMP candidates is indicated. In some other embodiments, the number of intra TMP candidates is derived.
[0123] In some embodiments, template matching is used to derive one or more intra TMP candidates.Alternatively or additionally, template matching is used to refine one or more intra TMP candidates.
[0124] In some embodiments, one or more intra prediction modes (IPMs) are used to generate an intra prediction signal. For example, one or more IPMs are predefined. Alternatively, one or more IPMs are indicated. In some other embodiments, one or more IPMs are derived. In some embodiments, the one or more predefined IPMs include at least one of the following: a planar mode, a direct current (DC) mode, a horizontal mode, or a vertical mode.
[0125] In some embodiments, the block offset of the intra TMP candidate is used to derive one or more IPMs. In some embodiments, the one or more IPMs are derived using neighboring samples. For example, the one or more IPMs are derived using at least one of: DIMD or TIMD. In some embodiments, the intra prediction mode derived using DIMD is used in TIMD to derive the intra prediction mode for TIMD.
[0126] In some embodiments, an IPM candidate list is constructed, and one or more IPMs in the IPM candidate list are used to generate an intra-frame prediction signal. For example, which IPM is used to generate the intra-frame prediction signal is predefined. Alternatively, which IPM is used to generate the intra-frame prediction signal is indicated using a syntax element. In some other embodiments, which IPM is used to generate the intra-frame prediction signal is derived.
[0127] In some embodiments, at least one codec tool is different from traditional intra-frame prediction. For example, at least one codec tool refers to at least one of the following: how to fill reference samples, whether to filter reference samples and / or how to filter reference samples, whether to apply a filtering process (e.g., PDPC / gradient PDPC) and / or how to apply the filtering process (e.g., PDPC / gradient PDPC), or whether to use an interpolation filter and / or how to use the interpolation filter. In some other embodiments, at least one codec tool used to obtain the intra-frame prediction signal is the same as traditional intra-frame prediction.
[0128] In some embodiments, one or more intra TMP candidates are reordered before being used to generate the intra TMP prediction signal or the intra prediction signal. Alternatively or additionally, one or more IPMs are reordered before being used to generate the intra TMP prediction signal or the intra prediction signal.
[0129] In some embodiments, the template matching cost is used for reordering, or the bilateral matching cost is used for reordering. For example, reordering is used for one or more intra-frame TMP candidates. Alternatively, reordering is used for one or more IPMs. In some other embodiments, reordering is used for a combination of one or more intra-frame TMP candidates and one or more IPMs.
[0130] In some embodiments, at least one of the intra prediction signal or the final prediction signal is refined by a filtering process. Alternatively, at least one of the intra prediction signal or the fused prediction signal is refined by a filtering process. For example, the filtering process may be PDPC or gradient PDPC.
[0131] In some embodiments, the intra TMP mode is combined with multiple codecs. Alternatively, the intra TMP mode is combined with a codec with multiple prediction signals.
[0132] In some embodiments, P(x, y) = w IP1 *IP1(x,y)+w IP2 *IP2(x,y)+...+w IPn *IP, where P(x, y) represents the prediction of the video unit, IP k (x, y) represents the prediction signal generated by the k-th intra prediction, IntraTMP j (x, y) represents the prediction signal generated by the j-th intra-frame TMP, w IPk represents the weighting parameter corresponding to the prediction signal generated by the k-th intra prediction, w TMPj represents the weighting parameters corresponding to the prediction signal generated by the j-th intra-frame TMP, and i, j, n, and m are integers. In some embodiments, n = 1 or 2 or 3 and m = 1. Alternatively, n = 1 or 2 or 3 and m = 2. In some other embodiments, n = 1 and m = 1 or 2 or 3. In some embodiments, n = 2 and m = 1 or 2 or 3. Alternatively, n = 0 and m = 2 or 3 or 4 or 5. In some other embodiments, n = 2 or 3 or 4 or 5 and m = 0.
[0133] In some embodiments, a set of weighting parameters used to fuse the intra-frame TMP prediction signal and the intra-frame prediction signal is predefined. Alternatively, the set of weighting parameters is indicated. In some other embodiments, the set of weighting parameters is derived.
[0134] In some embodiments, the intra prediction signal and the intra TMP prediction signal are used in the fusion of the intra TMP mode with the codec tool. For example, P(x, y) = w IP *IP(x,y)+w TMP *IntraTMP(x, y), where P(x, y) represents the prediction of the video unit, IP(x, y) represents the intra-frame prediction signal, IntraTMP(x, y) represents the intra-frame TMP prediction signal, and w IP represents the weighting parameter corresponding to the intra-frame prediction signal n, w TMP represents the weighting parameter corresponding to the intra-frame TMP prediction signal, and w IP +w TMP =1.
[0135] In some other embodiments, P(x, y)=(w IP *IP(x,y)+w TMP *IntraTMP(x, y)+offset)>>shift, where P(x, y) represents the prediction of the video unit, IP(x, y) represents the intra-frame prediction signal, IntraTMP(x, y) represents the intra-frame TMP prediction signal, wIP represents the weighted parameter corresponding to the intra prediction signal n, w TMP represents the weighted parameter corresponding to the intra TMP prediction signal, w IP +w TMP =(1 << <shift>), and <shift> represents a parameter. For example, offset = 0. As another example, offset = 1 << (<shift> - 1). In some other embodiments, <shift> = 5 or 6 or 7 or 8.
[0136] In some embodiments, w IP = 1, <shift> = 1. In some other embodiments, w IP = 1 or 2 or 3, <shift> = 2. Alternatively, w IP = 1 or 2 or 3 or 4 or 5 or 6 or 7, <shift> = 3. As another example, w IP = 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15, <shift> = 4.
[0137] In some embodiments, the set of weighted parameters is signaled. For example, the set of weighted parameters is constructed and an index indicating the set of weighted parameters is indicated.
[0138] In some embodiments, the set of weighted parameters is derived using codec information. For example, the codec information includes the codec mode of neighboring video units. In an example embodiment, the set of weighted parameters depends on whether one or more neighboring video units are coded using intra prediction or IBC mode.
[0139] In some embodiments, the codec information includes the intra prediction mode used to obtain the intra prediction signal. In some other embodiments, the codec information includes at least one of the following: the block size of the video unit, the block dimension of the video unit, the block size of neighboring video units, or the block dimension of neighboring video units.
[0140] In some embodiments, the set of weighted parameters is derived using a template matching method. For example, the set of weighted parameters can be derived using the minimum template matching cost.
[0141] In some embodiments, the set of weighted parameters is derived using the codec information generated during the search for intra TMP candidates. For example, the codec information includes the search cost. For example, the search cost (e.g., SAD or MRSAD) can be used.
[0142] In some embodiments, a template matching-based method is used to derive the set of weighting parameters. In some other embodiments, the set of weighting parameters is derived using an LDL method or Gaussian elimination. For example, one or more samples in the current template indicated by the intra-frame TMP candidate are used. Alternatively or additionally, one or more samples in a reference template of the current template indicated by the intra-frame TMP candidate are used. In some other embodiments, a prediction signal of the current template is used.
[0143] In some embodiments, the set of weighting parameters depends on the video content of the video unit. For example, the weighting parameters are different for natural sequences and screen content sequences.
[0144] In some embodiments, the set of weighting parameters is generated in the same or similar manner as that used for geometric partitioning mode (GPM) or spatial GPM (SGPM). In some embodiments, at least one index or syntax element is indicated to indicate which set of weighting parameters will be used. In some other embodiments, the set of weighting parameters is derived at the decoder without signaling which set of weighting parameters will be used.
[0145] In some embodiments, the codec information used for the fusion of the intra TMP mode and the codec tool is used to encode and decode subsequent video units of the video unit. For example, the block vector used to generate the intra TMP prediction signal is considered to be a block vector of at least one of the following: intra TMP or normal IBC.
[0146] In some embodiments, the block vector is added to an IBC history-based motion vector prediction (HMVP) table. In some other embodiments, the block vector is used to construct an IBC advanced motion vector prediction (AMVP) candidate list for subsequent video units. Alternatively or additionally, the block is used to construct an IBC Merge candidate list for subsequent video units.
[0147] In some embodiments, the IPM used to generate the intra prediction signal is treated as an IPM for normal intra prediction. For example, the IPM is used to construct a most probable mode (MPM) list for subsequent video units. In some embodiments, the IPM is used for chroma prediction. In some other embodiments, the IPM is propagated for video units that are not intra-coded.
[0148] In some embodiments, the codec information used for the fusion of the intra-frame TMP mode and the codec tool is not used to encode and decode subsequent video units of the video unit. For example, instead of the IPM used for the current video unit, a predefined IPM is used for the subsequent video unit.
[0149] In some embodiments, the intra TMP prediction signal and the second predicted signal are fused by directly combining the intra TMP prediction signal and the second predicted signal based on position. For example, intra TMP prediction is applied, and for at least one position, the second prediction is applied.
[0150] In some embodiments, multiple fusion methods are applied to the intra-frame TMP. For example, which fusion method is used is predefined. Alternatively, which fusion method is used is indicated. In some other embodiments, which fusion method is used is derived.
[0151] In some embodiments, whether to apply the intra TMP mode for a video unit in combination with a codec tool and / or the manner in which the intra TMP mode for a video unit in combination with a codec tool is applied depends on codec information, for example, the codec information includes at least one of the following: whether intra TMP or intra prediction method is allowed, block dimension, block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0152] In some embodiments, the intra TMP mode is combined with the codec tools and is applied to I slices or I pictures. Alternatively, the intra TMP mode is combined with the codec tools and is applied to all slice types or all picture types.
[0153] In some embodiments, the way template matching is performed for a block depends on whether the block is to be fused by an intra TMP and a second prediction. For example, the second prediction is taken into account during calculation of the template cost.
[0154] In some embodiments, whether to apply the intra-frame TMP mode and the codec fusion and / or the manner in which the intra-frame TMP mode and the codec fusion is applied depends on at least one of the following: color format or color component. For example, the intra-frame TMP mode and the codec fusion is applied to all color components.
[0155] In some embodiments, if a fusion of the intra TMP mode and the codec is applied to the chroma components, the derivation of intra prediction is different from the derivation for the luma component. In some embodiments, intra prediction is obtained using one of the following: CCLM, Multi-Model CCLM (MMLM), CCCM, Chroma-DIMD, Chroma-TIMD, CCLM fusion with angular mode, MMLM fusion with angular mode, or CCCM fusion with angular mode.
[0156] In some embodiments, whether to apply the intra-frame TMP mode in combination with the codec tool to the first component and / or the manner in which the intra-frame TMP mode in combination with the codec tool is applied to the first component depends on whether the intra-frame TMP mode in combination with the codec tool is applied to the second component. For example, the first component includes a chroma component and the second component includes a luma component.
[0157] In some embodiments, the fusion of the intra-frame TMP mode and the codec tool is applied to the first component in the same manner as the second component. In some other embodiments, the fusion of the intra-frame TMP mode and the codec tool is applied to the first component in a different manner than the second component. In some embodiments, the weighting parameters of the fusion are different.
[0158] In some embodiments, the intra-frame TMP mode is combined with the codec tool to be applied to the luma component but not to the chroma components. For example, the luma component includes Y in the YCbCr color space or green (G) in the red, green, and blue (RGB) color space.
[0159] In some embodiments, the chrominance component includes at least one of the following: Cb or Cr in the YCbCr color space. Alternatively, the chrominance component includes at least one of the following: R or B in the RGB color space.
[0160] In some embodiments, the indication of the integration of the intra TMP mode with the codec is dynamically derived. In some other embodiments, the indication of the integration of the intra TMP mode with the codec is signaled based on a condition. For example, the condition includes at least one of: whether intra TMP or intra prediction method is allowed, block dimension, block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0161] In some embodiments, the intra TMP mode is indicated for I slices or I pictures, and in some other embodiments, the intra TMP mode is indicated for all slice types or all picture types.
[0162] In some embodiments, for a block located at the top left corner of a slice, the fusion of the intra TMP mode with the codec tool is not signaled.Alternatively or additionally, for a block located at the top left corner of a picture, the fusion of the intra TMP mode with the codec tool is not signaled.
[0163] In some embodiments, if the indication of the integration of the intra-frame TMP mode with the codec tool is not signaled, the indication is presumed to be a default value. In some embodiments, if the indication of the integration of the intra-frame TMP mode with the codec tool is not signaled, the indication is presumed to be false. Alternatively, if the indication of the integration of the intra-frame TMP mode with the codec tool is not signaled, the indication is presumed to be true.
[0164] In some embodiments, whether the current block is encoded or decoded using a fusion of an intra TMP mode and a codec is signaled using one or more syntax elements. In some embodiments, the one or more syntax elements are binarized into one of the following: a flag, a fixed-length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code.
[0165] In some embodiments, one or more syntax elements are context coded. Alternatively, one or more syntax elements are bypass coded. In some embodiments, the context depends on coded information. In some embodiments, the coded information includes at least one of the following: block dimensions, block size, slice type, picture type, information about neighboring blocks, information about other codecs used for the current block, or information about temporal layers.
[0166] In some embodiments, if the video unit is intra TMP coded, the fusion of the intra TMP mode with the codec tool is signaled. In some embodiments, one or more syntax elements are indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0167] In some embodiments, one or more syntax elements are encoded or decoded in a predictive manner. In some embodiments, one or more syntax elements are predicted from one or more syntax elements of neighboring blocks.
[0168] In some embodiments, an indication of whether and / or how to derive prediction or reconstruction of a video unit based on a fusion of an intra TMP mode with a codec is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level. In some embodiments, an indication of whether and / or how to derive prediction or reconstruction of a video unit based on a fusion of an intra TMP mode with a codec is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header. In some embodiments, method 5000 further comprises: determining whether and / or how to derive prediction or reconstruction of a video unit based on fusion of an intra-frame TMP mode with a codec tool based on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, TU, PU block, video codec unit, position of one of the following: CU, PU, TU, block, video codec unit, block dimensions of the current block and / or neighboring blocks of the current block, block shape of the current block and / or neighboring blocks of the current block, codec mode of the video unit, indication of a color format, codec tree structure, slice type, slice group type, picture type, color component, temporal layer identifier, profile or level or layer of the standard.
[0169] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode and the codec tool; and generating a bitstream based on the prediction or reconstruction of the video unit.
[0170] According to some further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode and the codec tool; generating a bitstream based on the prediction or reconstruction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0171] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0172] Item 1. A method for video processing, comprising: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool for conversion between a video unit of a video and a bitstream of the video unit; deriving a prediction or reconstruction of the video unit based on the fusion of the intra-frame TPM mode and the codec tool; and performing the conversion based on the prediction or reconstruction of the video unit.
[0173] Item 2. The method of Item 1, wherein the codec tool is one of: an inter-frame prediction mode, a palette mode, an intra-block copy (IBC) mode, or a block-based delta pulse codec modulation (BDPCM) mode.
[0174] Clause 3. The method of clause 1, wherein the coding tool is an intra prediction mode.
[0175] Item 4. A method according to Item 3, wherein the intra-frame prediction mode includes one of the following: a conventional intra-frame prediction mode or another intra-frame prediction mode, and the other intra-frame prediction mode obtains a prediction block using samples from one of the following other than intra-frame template matching prediction: a current slice, a current slice, a current sub-picture, a current picture, or another video unit.
[0176] Item 5. A method according to Item 3, wherein the intra-frame prediction method includes one of the following: decoder-side intra-frame mode derivation (DIMD), template-based intra-frame mode derivation (TIMD), intra-frame sub-partitioning (ISP), matrix-weighted intra-frame prediction (MIP), multiple reference lines (MRL), position-dependent intra-frame prediction combination (PDPC), gradient PDPC, intra-frame prediction fusion or template-based multiple reference line intra-frame prediction (TMRL).
[0177] Item 6. A method according to Item 3, wherein the intra-frame prediction method includes one of the following: a cross-component linear model (CCLM), a variant of CCLM, a multi-model CCLM, a left CCLM, an upper CCLM, a convolutional cross-component model (CCCM), a variant of CCCM, a left CCCM, an upper CCCM, a gradient linear model (GLM) or a variant of GLM.
[0178] Clause 7. The method of clause 1, wherein one or more intra TMP candidates are used to generate the intra TMP prediction signal.
[0179] Item 8. The method of Item 7, wherein the derivation of the one or more intra-frame TMP candidates is different from the derivation of the intra-frame TMP, or wherein the derivation of the one or more intra-frame TMP candidates is the same as the derivation of the intra-frame TMP.
[0180] Item 9. The method of Item 8, wherein the current template used to derive the intra TMP candidate is different.
[0181] Item 10. The method of Item 9, wherein the current template is of a different shape or size.
[0182] Clause 11. The method of clause 9, wherein the current template is downsampled.
[0183] Clause 12. The method of clause 9, wherein one or more samples in the current template are modified before being used to derive the intra TMP candidate.
[0184] Clause 13. The method of clause 12, wherein the prediction signal of the current template is used to modify the current template.
[0185] Item 14. The method of Item 13, wherein the prediction signal is generated using the same intra prediction method used to fuse the intra TMP of the current block.
[0186] Item 15. The method of Item 13, wherein T' = (a*Tb*Tp) / c, wherein T' represents the modified template, T represents the current template, Tp represents the prediction signal, and a, b, and c are integers.
[0187] Item 16. The method of Item 15, wherein a=2, b=1, and c=1, or wherein a=4, b=3, and c=1, wherein a=8, b=7, and c=1, or wherein a=16, b=15, and c=1.
[0188] Item 17. The method according to Item 13, wherein T'=(a*Tb*Tp+offset)>>shift, wherein T' represents the modified template, T represents the current template, Tp represents the prediction signal, shift is a parameter, and a, b and c are integers.
[0189] Clause 18. The method of clause 12, wherein the top left sample point of the current template is not modified.
[0190] Clause 19. The method of clause 8, wherein at least one of the following for the derivation of the one or more intra TMP candidates is different from an intra TMP: a search area, a search region, or a search method.
[0191] Item 20. The method of Item 7, wherein a plurality of intra TMP candidates are derived.
[0192] Item 21. The method of Item 20, wherein the plurality of intra TMP candidates are derived from different search results.
[0193] Item 22. The method of Item 20, wherein at least two intra TMP candidates are derived from the same search area.
[0194] Item 23. A method according to item 20, wherein which intra-frame TMP candidate is used for the fusion is predefined, or wherein which intra-frame TMP candidate is used for the fusion is indicated using a syntax element, or wherein which intra-frame TMP candidate is used for the fusion is derived.
[0195] Item 24. The method of Item 20, wherein the number of intra TMP candidates is predefined, wherein the number of intra TMP candidates is indicated, or wherein the number of intra TMP candidates is derived.
[0196] Item 25. The method of Item 7, wherein template matching is used to derive the one or more intra-frame TMP candidates, and / or wherein the template matching is used to refine the one or more intra-frame TMP candidates.
[0197] Clause 26. The method of clause 1, wherein one or more intra prediction modes (IPMs) are used to generate the intra prediction signal.
[0198] Clause 27. The method of clause 26, wherein the one or more IPMs are predefined, or wherein the one or more IPMs are indicated, or wherein the one or more IPMs are derived.
[0199] Item 28. The method of Item 27, wherein the one or more predefined IPMs include at least one of: a planar mode, a direct current (DC) mode, a horizontal mode, or a vertical mode.
[0200] Item 29. The method of Item 26, wherein block offsets of intra-frame TMP candidates are used to derive the one or more IPMs.
[0201] Item 30. The method of Item 26, wherein the one or more IPMs are derived using neighboring samples.
[0202] Item 31. The method of Item 30, wherein the one or more IPMs are derived using at least one of: DIMD or TIMD.
[0203] Item 32. The method of Item 30, wherein the derived intra prediction mode using DIMD is used in TIMD to derive the intra prediction mode of TIMD.
[0204] Item 33. The method of Item 26, wherein an IPM candidate list is constructed, and one or more IPMs in the IPM candidate list are used to generate the intra prediction signal.
[0205] Item 34. A method according to item 33, wherein which IPM is used to generate the intra-frame prediction signal is predefined, or wherein which IPM is used to generate the intra-frame prediction signal is indicated using a syntax element, or which IPM is used to generate the intra-frame prediction signal is derived.
[0206] Item 35. The method of Item 26, wherein at least one codec tool is different from conventional intra prediction.
[0207] Item 36. A method according to Item 35, wherein the at least one coding tool refers to at least one of the following: a way of filling reference samples, whether to filter the reference samples and / or a way of filtering the reference samples, whether to apply a filtering process and / or a way of applying the filtering process, or whether to use an interpolation filter and / or a way of using an interpolation filter.
[0208] Item 37. The method of Item 35, wherein the at least one codec tool used to obtain the intra prediction signal is the same as conventional intra prediction.
[0209] Item 38. A method according to item 1, wherein one or more intra-frame TMP candidates are reordered before being used to generate an intra-frame TMP prediction signal or an intra-frame prediction signal, and / or wherein one or more IPMs are reordered before being used to generate the intra-frame TMP prediction signal or the intra-frame prediction signal.
[0210] Item 39. The method of Item 38, wherein a template matching cost is used for the reordering, or wherein a bilateral matching cost is used for the reordering.
[0211] Item 40. The method of Item 38, wherein the reordering is applied to the one or more intra-frame TMP candidates.
[0212] Clause 41. The method of clause 38, wherein the reordering is applied to one or more IPMs.
[0213] Item 42. The method of Item 38, wherein the reordering is applied to a combination of the one or more intra TMP candidates and the one or more IPMs.
[0214] Item 43. A method according to item 1, wherein at least one of the intra-frame prediction signal or the final prediction signal is refined by a filtering process, or wherein at least one of the intra-frame prediction signal or the fused prediction signal is refined by the filtering process.
[0215] Item 44. The method of Item 43, wherein the filtering process comprises one of: position-dependent intra prediction combining (PDPC) or gradient PDPC.
[0216] Item 45. The method of Item 1, wherein the intra-frame TMP mode and multiple codec tools are fused.
[0217] Item 46. The method of Item 1, wherein the intra TMP mode and the codec with multiple prediction signals are fused.
[0218] Item 47. The method of Item 1, wherein P(x, y) = w IP1 *IP1(x,y)+w IP2 *IP2(x,y)+...+w IPn *IP n (x, y) + w TMP1 *IntraTMP1(x,y)+w TMP2 *IntraTMP2(x, y)+...+w TMPm *IntraTMP m (x, y), where P(x, y) represents the prediction of the video unit, IP k (x, y) represents the prediction signal generated by the k-th intra prediction, IntraTMP j (x, y) represents the prediction signal generated by the j-th intra-frame TMP, w IPk represents the weighting parameter corresponding to the prediction signal generated by the k-th intra-frame prediction, w TMPj represents a weighting parameter corresponding to the prediction signal generated by the j-th intra-frame TMP, and i, j, n and m are integers.
[0219] Item 48. The method of Item 47, wherein n=1 or 2 or 3 and m=1, or wherein n=1 or 2 or 3 and m=2, wherein n=1 and m=1 or 2 or 3, wherein n=2 and m=1 or 2 or 3, wherein n=0 and m=2 or 3 or 4 or 5, or wherein n=2 or 3 or 4 or 5 and m=0.
[0220] Item 49. A method according to item 1, wherein a set of weighting parameters used to fuse the intra-frame TMP prediction signal and the intra-frame prediction signal is predefined, or wherein the set of weighting parameters is indicated, or wherein the set of weighting parameters is derived.
[0221] Item 50. The method of Item 49, wherein an intra prediction signal and an intra TMP prediction signal are used in the fusion of the intra TMP mode with the codec tool.
[0222] Item 51. The method of Item 50, wherein P(x, y) = w IP *IP(x,y)+w TMP *IntraTMP(x,y), and wherein P(x,y) represents the prediction of the video unit, IP(x,y) represents the intra prediction signal, IntraTMP(x,y) represents the intra TMP prediction signal, w IP represents the weighting parameter corresponding to the intra-frame prediction signal n, w TMP represents the weighting parameter corresponding to the intra-frame TMP prediction signal, and w IP +w TMP =1.
[0223] Item 52. The method of Item 50, wherein P(x, y) = (w IP *IP(x,y)+w TMP *IntraTMP(x,y)+offset)>>shift, and wherein P(x,y) represents the prediction of the video unit, IP(x,y) represents the intra prediction signal, IntraTMP(x,y) represents the intra TMP prediction signal, w IP represents the weighting parameter corresponding to the intra-frame prediction signal n, w TMP represents the weighting parameter corresponding to the intra-frame TMP prediction signal, w IP +w TMP =(1<<shift), and shift represents a parameter.
[0224] Item 53. The method of Item 52, wherein offset = 0, or wherein offset = 1 << (shift - 1), or wherein shift = 5 or 6 or 7 or 8.
[0225] Item 54. The method of Item 52, wherein w IP =1, shift=1, or w IP =1 or 2 or 3, shift=2, or w IP =1 or 2 or 3 or 4 or 5 or 6 or 7, shift=3, or w IP =1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15, shift=4.
[0226] Item 55. The method of Item 49, wherein the set of weighting parameters is transmitted via a signal.
[0227] Clause 56. The method of clause 55, wherein the set of weighting parameters is constructed and an index indicating the set of weighting parameters is indicated.
[0228] Item 57. The method of Item 49, wherein the set of weighting parameters is derived using codec information.
[0229] Item 58. The method of Item 57, wherein the codec information comprises a codec mode of a neighboring video unit.
[0230] Item 59. The method of Item 58, wherein the set of weighting parameters depends on whether one or more neighboring video units are encoded using intra prediction or IBC mode.
[0231] Item 60. The method of Item 57, wherein the codec information includes an intra prediction mode used to obtain an intra prediction signal.
[0232] Item 61. The method of Item 57, wherein the codec information comprises at least one of: a block size of the video unit, a block dimension of the video unit, a block size of a neighboring video unit, or a block dimension of a neighboring video unit.
[0233] Item 62. The method of Item 57, wherein the set of weighting parameters is derived using a template matching method.
[0234] Item 63. The method of Item 57, wherein the set of weighting parameters is derived using the codec information generated during the search for intra-frame TMP candidates.
[0235] Item 64. The method of Item 63, wherein the codec information includes a search cost.
[0236] Item 65. The method of Item 57, wherein a template matching based method is used to derive the set of weighting parameters.
[0237] Item 66. The method of Item 57, wherein the set of weighting parameters is derived using the LDL method or Gaussian elimination.
[0238] Item 67. A method according to item 66, wherein one or more samples in the current template indicated by the intra-frame TMP candidate are used, and / or wherein one or more samples in the reference template of the current template indicated by the intra-frame TMP candidate are used, and / or wherein the prediction signal of the current template is used.
[0239] Item 68. The method of Item 49, wherein the set of weighting parameters depends on the video content of the video unit.
[0240] Item 69. The method of Item 68, wherein the weighting parameters are different for natural sequences and screen content sequences.
[0241] Item 70. The method of Item 49, wherein the set of weighting parameters is generated in the same or similar manner as the weighting parameters generated by a geometric partitioning mode (GPM) or a spatial GPM (SGPM).
[0242] Item 71. The method of item 70, wherein at least one index or syntax element is indicated to indicate which set of weighting parameters is to be used.
[0243] Item 72. The method of Item 70, wherein the set of weighting parameters is derived at the decoder without signaling which set of weighting parameters is to be used.
[0244] Item 73. The method of Item 1, wherein the codec information used for the fusion of the intra TPM mode and the codec tool is used to encode and decode a subsequent video unit of the video unit.
[0245] Item 74. The method of Item 73, wherein the block vector used to generate the intra TMP prediction signal is considered to be the block vector of at least the last of: intra TMP or normal IBC.
[0246] Item 75. The method of Item 74, wherein the block vector is added to an IBC history based motion vector prediction (HMVP) table.
[0247] Item 76. The method of Item 74, wherein the block vector is used to construct an IBC Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video unit, and / or wherein the block is used to construct an IBC Merge candidate list for the subsequent video unit.
[0248] Item 77. The method of Item 73, wherein the IPM used to generate the intra prediction signal is regarded as an IPM for normal intra prediction.
[0249] Item 78. The method of Item 77, wherein the IPM is used to construct a most probable mode (MPM) list for subsequent video units.
[0250] Item 79. The method of Item 77, wherein the IPM is used for chroma prediction.
[0251] Item 80. The method of Item 77, wherein the IPM is propagated for non-intra-coded video units.
[0252] Item 81. The method of Item 1, wherein codec information used for the fusion of the intra TPM mode with the codec tool is not used for encoding and decoding subsequent video units of the video unit.
[0253] Item 82. The method of Item 81, wherein instead of the IPM used for the current video unit, a predefined IPM is used for the subsequent video unit.
[0254] Item 83. The method of Item 1, wherein the intra TMP prediction signal and the second predicted signal are fused by directly combining the intra TMP prediction signal and the second predicted signal based on position.
[0255] Item 84. The method of Item 83, wherein the intra TMP prediction is applied and for at least one position the second prediction is applied.
[0256] Item 85. The method of Item 1, wherein multiple fusion methods are applied to the intra-frame TMP.
[0257] Item 86. A method according to Item 85, wherein which fusion method is used is predefined, or wherein which fusion method is used is indicated, or wherein which fusion method is used is derived.
[0258] Item 87. A method according to Item 1, wherein whether to apply the fusion of the intra-frame TPM mode for the video unit and the codec tool and / or the manner in which the fusion of the intra-frame TPM mode for the video unit and the codec tool is applied depends on codec information.
[0259] Item 88. A method according to item 87, wherein the coding and decoding information includes at least one of the following: whether intra-frame TMP or the intra-frame prediction method is allowed, block dimension, block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position or color component.
[0260] Item 89. The method of Item 88, wherein the fusion of the intra TPM mode with the codec tool is applied to an I slice or an I picture.
[0261] Item 90. The method of Item 88, wherein the fusion of the intra TPM mode with the codec tool is applied to all slice types or all picture types.
[0262] Item 91. The method of Item 1, wherein the manner in which template matching is performed for a block depends on whether the block is to be fused by the intra TMP and a second prediction.
[0263] Item 92. The method of Item 91, wherein the second prediction is taken into account during calculation of the template cost.
[0264] Item 93. A method according to Item 1, wherein whether to apply the fusion of the intra-frame TPM mode and the codec tool and / or the manner in which the fusion of the intra-frame TPM mode and the codec tool is applied depends on at least one of the following: color format or color component.
[0265] Item 94. The method of Item 93, wherein the fusion of the intra-TPM mode with the codec tool is applied to all color components.
[0266] Item 95. The method of Item 93, wherein if the fusion of the intra TPM mode with the codec tool is applied to chroma components, the derivation of intra prediction is different from the derivation for luma components.
[0267] Item 96. A method according to Item 95, wherein the intra-frame prediction is obtained using one of the following: CCLM, multi-model CCLM (MMLM), CCCM, chroma-DIMD, chroma-TIMD, a fusion of CCLM and angular mode, a fusion of MMLM and angular mode, or a fusion of CCCM and angular mode.
[0268] Item 97. A method according to Item 93, wherein whether the fusion of the intra-frame TPM mode and the codec tool is applied to the first component and / or the manner in which the fusion of the intra-frame TPM mode and the codec tool is applied to the first component depends on whether the fusion of the intra-frame TPM mode and the codec tool is applied to the second component.
[0269] Item 98. The method of Item 97, wherein the first component comprises a chrominance component and the second component comprises a luma component.
[0270] Item 99. The method of Item 97, wherein the fusion of the intra-TPM mode with the codec tool is applied to the first component in the same manner as applied to the second component.
[0271] Item 100. The method of Item 97, wherein the fusion of the intra-TPM mode with the codec tool is applied to the first component differently than to the second component.
[0272] Item 101. A method according to Item 100, wherein the weighting parameters of the fusion are different.
[0273] Item 102. The method of Item 93, wherein the fusion of the intra TPM mode with the codec tool is applied to the luma component but not to the chroma components.
[0274] Item 103. The method of Item 102, wherein the luma component comprises Y in a YCbCr color space or green (G) in a red-green-blue (RGB) color space.
[0275] Item 104. The method of Item 102, wherein the chrominance component comprises at least one of: Cb or Cr in a YCbCr color space, or wherein the chrominance component comprises at least one of: R or B in an RGB color space.
[0276] Item 105. The method of any one of Items 1-104, wherein the indication of the fusion of the intra-TPM mode with the codec tool is derived dynamically.
[0277] Item 106. The method of any one of Items 1-104, wherein the indication of the fusion of the intra-TPM mode with the codec tool is signaled based on a condition.
[0278] Item 107. A method according to item 106, wherein the condition includes at least one of the following: whether intra-frame TMP or the intra-frame prediction method is allowed, block dimension, block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position or color component.
[0279] Item 108. The method of Item 107, wherein for an I slice or an I picture, the fusion of an intra TPM mode with the codec tool is indicated.
[0280] Item 109. The method of Item 107, wherein the fusion of intra TPM mode with the codec tool is indicated for all slice types or all picture types.
[0281] Item 110. The method of Item 107, wherein for a block located at the top left corner of a slice or a top left corner of a picture, the fusion of the intra TPM mode with the codec tool is not signaled.
[0282] Item 111. The method of Item 106, wherein if the indication of the fusion of intra-TPM mode with the codec tool is not signaled, the indication is inferred to be a default value.
[0283] Item 112. The method of Item 106, wherein if the indication of the fusion of intra TPM mode with the codec tool is not signaled, the indication is presumed to be false.
[0284] Item 113. The method of Item 106, wherein if the indication of the fusion of intra TPM mode with the codec tool is not signaled, the indication is presumed to be true.
[0285] Item 114. The method of any one of Items 1-106, wherein whether the current block is coded using the fusion of the intra TPM mode with the codec tool is signaled using one or more syntax elements.
[0286] Item 115. The method of Item 114, wherein the one or more syntax elements are binarized as one of: a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code.
[0287] Item 116. The method of Item 114, wherein the one or more syntax elements are context coded, or wherein the one or more syntax elements are bypass coded.
[0288] Item 117. The method of Item 116, wherein the context depends on encoded information.
[0289] Item 118. A method according to item 117, wherein the coded information includes at least one of the following: block dimension, block size, slice type, picture type, information of neighboring blocks, information of other coding tools used for the current block, or information of time domain layers.
[0290] Item 119. The method of Item 114, wherein if the video unit is intra-TPM coded, the fusion of intra-TPM mode with the codec tool is transmitted via a signal.
[0291] Item 120. A method according to item 114, wherein the one or more syntax elements are indicated at one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.
[0292] Item 121. The method of Item 114, wherein the one or more syntax elements are encoded in a predictive manner.
[0293] Item 122. The method of Item 114, wherein the one or more syntax elements are predicted by one or more syntax elements of a neighboring block.
[0294] Item 123. A method according to any one of Items 1-122, wherein the video unit includes at least one of the following: a color component, a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec unit (CU), a codec tree unit (CTU), a CTU row, a CTU group, a slice, a sub-picture, a block, a sub-region within a block, or a region including more than one sample or pixel.
[0295] Item 124. A method according to any one of items 1-123, wherein an indication of whether and / or how the prediction or reconstruction of the video unit is derived based on the fusion of the intra-frame TPM mode with the codec tool is indicated at one of the following: sequence level, picture group level, picture level, slice level or slice group level.
[0296] Item 125. A method according to any one of Items 1-123, wherein an indication of whether and / or how the prediction or reconstruction of the video unit is derived based on the fusion of the intra-frame TPM mode with the codec tool is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.
[0297] Item 126. The method according to any one of items 1-123 further includes: determining whether and / or how to derive the prediction or reconstruction of the video unit based on the fusion of the intra-frame TPM mode and the codec tool based on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, TU, PU block, video codec unit, the position of one of the following: CU, PU, TU, block, video codec unit, block dimensions of the current block and / or the neighboring blocks of the current block, block shape of the current block and / or the neighboring blocks of the current block, the codec mode of the video unit, an indication of the color format, the codec tree structure, slice type, slice group type, picture type, color component, temporal layer identifier, standard grade or level or layer.
[0298] Item 127. The method of any one of Items 1-126, wherein the converting comprises encoding the video unit into the bitstream.
[0299] Item 128. The method of any one of Items 1-126, wherein the converting comprises decoding the video unit from the bitstream.
[0300] Item 129. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1-128.
[0301] Item 130. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of Items 1-128.
[0302] Item 131. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode and the codec tool; and generating the bitstream based on the prediction or reconstruction of the video unit.
[0303] Item 132. A method for storing a bitstream of a video, comprising: determining a fusion of an intra-frame template matching prediction (intra-frame TMP) mode and a codec tool; deriving a prediction or reconstruction of a video unit of the video based on the fusion of the intra-frame TMP mode and the codec tool; generating the bitstream based on the prediction or reconstruction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0304] Figure 51 A block diagram of a computing device 5100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 5100 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0305] It should be understood that Figure 51 The computing device 5100 shown in FIG. 5 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0306] like Figure 51As shown, computing device 5100 includes a general computing device 5100. Computing device 5100 may include at least one or more processors or processing units 5110, memory 5120, storage unit 5130, one or more communication units 5140, one or more input devices 5150, and one or more output devices 5160.
[0307] In some embodiments, the computing device 5100 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, and including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 5100 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0308] The processing unit 5110 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 5120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of the computing device 5100. The processing unit 5110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0309] The computing device 5100 typically includes various computer storage media. Such media can be any media accessible by the computing device 5100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 5120 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 5130 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 5100.
[0310] The computing device 5100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 51 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0311] The communication unit 5140 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 5100 can be implemented by a single computing cluster or multiple computing machines communicating via a communication connection. Thus, the computing device 5100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0312] The input device 5150 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. The output device 5160 may be one or more of various output devices, such as a display, speaker, printer, etc. With the help of the communication unit 5140, the computing device 5100 may also communicate with one or more external devices (not shown), such as storage devices and display devices. The computing device 5100 may also communicate with one or more devices that enable a user to interact with the computing device 5100, or any device that enables the computing device 5100 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.), if necessary. Such communication may be carried out via an input / output (I / O) interface (not shown).
[0313] In some embodiments, some or all components of the computing device 5100 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on servers in a remote location. Computing resources in a cloud computing environment can be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although they appear to be a single access point for users. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, they can be provided from a conventional server or installed directly or otherwise on a client device.
[0314] In embodiments of the present disclosure, the computing device 5100 may be used to implement video encoding / decoding. The memory 5120 may include one or more video encoding / decoding modules 5125 having one or more program instructions. These modules can be accessed and executed by the processing unit 5110 to perform the functions of the various embodiments described herein.
[0315] In an example embodiment performing video encoding, an input device 5150 may receive video data as input to be encoded 5170. The video data may be processed, for example, by a video codec module 5125 to generate an encoded bitstream. The encoded bitstream may be provided as output 5180 via an output device 5160.
[0316] In an example embodiment performing video decoding, an input device 5150 may receive an encoded bitstream as input 5170. The encoded bitstream may be processed, for example, by a video codec module 5125 to generate decoded video data. The decoded video data may be provided as output 5180 via an output device 5160.
[0317] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such changes are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A video processing method, comprising: Determining the integration of an intra-frame template matching prediction (intra-frame TMP) mode with codec tools for conversion between video units of a video and a bitstream of the video units; deriving a prediction or reconstruction of the video unit based on the fusion of the intra TPM mode and the codec tool; and The conversion is performed based on the prediction or reconstruction of the video unit.
2. The method of claim 1, wherein the codec tool is one of: an inter-frame prediction mode, a palette mode, an intra-block copy (IBC) mode, or a block-based delta pulse codec modulation (BDPCM) mode. The method of claim 1 , wherein the coding tool is an intra prediction mode.
4. The method according to claim 3, wherein the intra prediction mode comprises one of the following: Regular intra prediction mode, or Other intra prediction modes, wherein the other intra prediction modes obtain a prediction block by using samples in one of the following except for intra template matching prediction: current slice, current slice, current sub-picture, current picture, or other video units.
5. The method of claim 3, wherein the intra prediction method comprises one of the following: decoder-side intra mode derivation (DIMD), template-based intra mode derivation (TIMD), intra sub-partitioning (ISP), matrix-weighted intra prediction (MIP), multiple reference lines (MRL), position-dependent intra prediction combination (PDPC), gradient PDPC, intra prediction fusion, or template-based multiple reference line intra prediction (TMRL).
6. The method of claim 3 , wherein the intra prediction method comprises one of the following: a cross-component linear model (CCLM), a variant of CCLM, a multi-model CCLM, a left CCLM, an upper CCLM, a convolutional cross-component model (CCCM), a variant of CCCM, a left CCCM, an upper CCCM, a gradient linear model (GLM), or a variant of GLM. The method of claim 1 , wherein one or more intra TMP candidates are used to generate an intra TMP prediction signal.
8. The method of claim 7, wherein the derivation of the one or more intra TMP candidates is different from the derivation of the intra TMP, or The derivation of the one or more intra TMP candidates is the same as the derivation of the intra TMP.
9. The method of claim 8, wherein the current templates used to derive the intra TMP candidates are different.
10. The method of claim 9, wherein the current template has a different shape or size. The method of claim 9 , wherein the current template is downsampled.
12. The method of claim 9, wherein one or more samples in the current template are modified before being used to derive the intra TMP candidate.
13. The method of claim 12, wherein the prediction signal of the current template is used to modify the current template. 14 . The method of claim 13 , wherein the prediction signal is generated using the same intra prediction method as that used to fuse the intra TMP of the current block.
15. The method of claim 13, wherein T'=(a*Tb*Tp) / c, wherein T' represents the modified template, T represents the current template, Tp represents the prediction signal, and a, b, and c are integers.
16. The method of claim 15, wherein a=2, b=1, and c=1, or Where a=4, b=3, and c=1, where a=8, b=7, and c=1, or Where a=16, b=15, and c=1.
17. The method of claim 13, wherein T'=(a*Tb*Tp+offset)>>shift, wherein T' represents the modified template, T represents the current template, Tp represents the prediction signal, shift is a parameter, and a, b and c are integers. The method according to claim 12 , wherein the upper left sample point of the current template is not modified.
19. The method of claim 8, wherein at least one of the following for the derived one or more intra TMP candidates is different from an intra TMP: Search area, Search area, or Search method.
20. The method of claim 7, wherein a plurality of intra TMP candidates are derived. The method of claim 20 , wherein the plurality of intra TMP candidates are derived from different search areas.
22. The method of claim 20, wherein at least two intra TMP candidates are derived from the same search area.
23. The method of claim 20, wherein which intra-frame TMP candidate is used for the fusion is predefined, or where which intra TMP candidate is used for the fusion is indicated using a syntax element, or Which intra TMP candidate is used for the fusion is derived.
24. The method of claim 20, wherein the number of intra-frame TMP candidates is predefined, wherein said number of intra-frame TMP candidates is indicated, or wherein the number of intra TMP candidates is derived.
25. The method of claim 7, wherein template matching is used to derive the one or more intra-frame TMP candidates, and / or The template matching is used to refine the one or more intra-frame TMP candidates.
26. The method of claim 1, wherein one or more intra prediction modes (IPMs) are used to generate the intra prediction signal.
27. The method of claim 26, wherein the one or more IPMs are predefined, or wherein the one or more IPMs are indicated, or wherein the one or more IPMs are derived.
28. The method of claim 27, wherein the one or more predefined IPMs include at least one of the following: Plane mode, Direct current (DC) mode, horizontal mode, or Vertical mode.
29. The method of claim 26, wherein block offsets of intra-frame TMP candidates are used to derive the one or more IPMs.
30. The method of claim 26, wherein the one or more IPMs are derived using neighboring samples.
31. The method of claim 30, wherein the one or more IPMs are derived using at least one of: DIMD or TIMD.
32. The method of claim 30, wherein the derived intra prediction mode using DIMD is used in TIMD to derive the intra prediction mode of TIMD.
33. The method of claim 26, wherein an IPM candidate list is constructed, and one or more IPMs in the IPM candidate list are used to generate the intra prediction signal.
34. The method according to claim 33, wherein which IPM is used to generate the intra prediction signal is predefined, or where which IPM is used to generate the intra prediction signal is indicated using a syntax element, or Which IPM is used to generate the intra prediction signal is derived.
35. The method of claim 26, wherein at least one codec tool is different from conventional intra prediction.
36. The method of claim 35, wherein the at least one codec tool is at least one of: The method of filling reference samples, whether to filter the reference samples and / or the manner in which the reference samples are filtered, whether and / or how filtering is applied, or Whether to use interpolation filters and / or how to use interpolation filters.
37. The method of claim 35, wherein the at least one codec tool used to obtain the intra prediction signal is the same as conventional intra prediction.
38. The method of claim 1, wherein one or more intra TMP candidates are reordered before being used to generate an intra TMP prediction signal or an intra prediction signal, and / or One or more IPMs are reordered before being used to generate the intra-frame TMP prediction signal or the intra-frame prediction signal.
39. The method of claim 38, wherein template matching costs are used for the reordering, or The bilateral matching cost is used for the reordering.
40. The method of claim 38, wherein the reordering is used for the one or more intra-frame TMP candidates.
41. The method of claim 38, wherein the reordering is used for one or more IPMs.
42. The method of claim 38, wherein the reordering is used for a combination of the one or more intra TMP candidates and the one or more IPMs.
43. The method of claim 1, wherein at least one of the intra prediction signal or the final prediction signal is refined by a filtering process, or At least one of the intra prediction signal or the fusion prediction signal is refined through the filtering process.
44. The method of claim 43, wherein the filtering process comprises one of: position-dependent intra prediction combining (PDPC) or gradient PDPC.
45. The method of claim 1, wherein the intra-frame TMP mode and multiple codec tools are fused.
46. The method of claim 1, wherein the intra TMP mode and the codec with multiple prediction signals are fused.
47. The method of claim 1, wherein P(x, y) = w IP1 *IP1(x,y)+w IP2 *IP2(x,y)+…+w IPn *IP n (x, y) + w TMP1 *IntraTMP1(x,y)+w TMP2 *IntraTMP2(x, y)+…+w TMPm *IntraTMP m (x, y), Where P(x, y) represents the prediction of the video unit, IP k (x, y) represents the prediction signal generated by the k-th intra prediction, IntraTMP j (x, y) represents the prediction signal generated by the j-th intra-frame TMP, w IPk represents the weighting parameter corresponding to the prediction signal generated by the k-th intra-frame prediction, w TMPj represents a weighting parameter corresponding to a prediction signal generated by the j-th intra-frame TMP, and i, j, n, and m are integers.
48. The method of claim 47, wherein n=1 or 2 or 3 and m=1, or Where n=1 or 2 or 3 and m=2, Where n=1 and m=1 or 2 or 3, Where n=2 and m=1 or 2 or 3, Where n=0 and m=2 or 3 or 4 or 5, or Wherein n=2 or 3 or 4 or 5 and m=0.
49. The method of claim 1, wherein a set of weighting parameters used to fuse the intra-frame TMP prediction signal and the intra-frame prediction signal is predefined, or wherein the set of weighting parameters is indicated, or Wherein the set of weighting parameters is derived.
50. The method of claim 49, wherein an intra prediction signal and an intra TMP prediction signal are used in the fusion of the intra TMP mode with the codec tool.
51. The method of claim 50, wherein P(x, y) = w IP *IP(x,y)+w TMP *IntraTMP(x,y), and Wherein P(x, y) represents the prediction of the video unit, IP(x, y) represents the intra-frame prediction signal, IntraTMP(x, y) represents the intra-frame TMP prediction signal, and w IP represents the weighting parameter corresponding to the intra-frame prediction signal n, w TMP represents the weighting parameter corresponding to the intra-frame TMP prediction signal, and w IP +w TMP =1.
52. The method of claim 50, wherein P(x, y) = (w IP *IP(x,y)+w TMP *IntraTMP(x,y)+offset)>>shift, and where P(x, y) represents the prediction of the video unit, IP(x, y) represents the intra prediction signal, IntraTMP(x, y) represents the intra TMP prediction signal, w IP represents the weighting parameter corresponding to the intra prediction signal n, w TMP represents the weighting parameter corresponding to the intra TMP prediction signal, w IP + w TMP = (1 << shift), and shift represents a parameter.
53. The method of claim 52, wherein offset = 0, or Where offset = 1 < < (shift - 1), or Where shift = 5 or 6 or 7 or 8.
54. The method of claim 52, wherein w IP =1, shift=1, or where w IP =1 or 2 or 3, shift=2, or where w IP =1 or 2 or 3 or 4 or 5 or 6 or 7, shift=3, or where w IP =1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9 or 10 or 11 or 12 or 13 or 14 or 15, shift=4.
55. The method of claim 49, wherein the set of weighting parameters is transmitted via a signal.
56. The method of claim 55, wherein the set of weighting parameters is constructed and an index indicating the set of weighting parameters is indicated.
57. The method of claim 49, wherein the set of weighting parameters is derived using codec information.
58. The method of claim 57, wherein the codec information comprises a codec mode of a neighboring video unit.
59. The method of claim 58, wherein the set of weighting parameters depends on whether one or more neighboring video units are coded using intra prediction or IBC mode.
60. The method of claim 57, wherein the codec information includes an intra prediction mode used to obtain an intra prediction signal.
61. The method according to claim 57, wherein the codec information comprises at least one of the following: the block size of the video unit, the block dimensions of the video unit, The block size of the adjacent video unit, or Block dimensions of adjacent video units.
62. The method of claim 57, wherein the set of weighting parameters is derived using a template matching method.
63. The method of claim 57, wherein the set of weighting parameters is derived using the codec information generated during a search for intra TMP candidates.
64. The method of claim 63, wherein the codec information includes a search cost.
65. The method of claim 57, wherein a template matching based method is used to derive the set of weighting parameters.
66. The method of claim 57, wherein the set of weighting parameters is derived using the LDL method or Gaussian elimination.
67. The method of claim 66, wherein one or more samples in the current template indicated by the intra TMP candidate are used, and / or wherein one or more samples in the reference template of the current template indicated by the intra TMP candidate are used, and / or The prediction signal of the current template is used.
68. The method of claim 49, wherein the set of weighting parameters depends on video content of the video unit.
69. The method of claim 68, wherein weighting parameters are different for natural sequences and screen content sequences.
70. The method of claim 49, wherein the set of weighting parameters is generated in the same or similar manner as weighting parameters generated by a geometric partitioning mode (GPM) or a spatial GPM (SGPM).
71. The method of claim 70, wherein at least one index or syntax element is indicated to indicate which set of weighting parameters is to be used.
72. The method of claim 70, wherein the set of weighting parameters is derived at a decoder without signaling which set of weighting parameters is to be used.
73. The method of claim 1, wherein codec information used for the fusion of intra TPM mode with the codec tool is used to encode and decode a subsequent video unit of the video unit.
74. The method of claim 73, wherein the block vector used to generate the intra TMP prediction signal is considered to be a block vector of at least one of: intra TMP or normal IBC.
75. The method of claim 74, wherein the block vector is added to an IBC history based motion vector prediction (HMVP) table.
76. The method of claim 74, wherein the block vector is used to construct an IBC Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video unit, and / or The blocks are used to construct an IBC Merge candidate list for the subsequent video unit.
77. The method of claim 73, wherein the IPM used to generate the intra prediction signal is regarded as an IPM of normal intra prediction.
78. The method of claim 77, wherein the IPM is used to construct a most probable mode (MPM) list for subsequent video units.
79. The method of claim 77, wherein the IPM is used for chroma prediction.
80. The method of claim 77, wherein the IPM is propagated for non-intra-coded video units.
81. The method of claim 1, wherein codec information used for the fusion of the intra TPM mode with the codec tool is not used to encode a subsequent video unit of the video unit.
82. The method of claim 81, wherein a predefined IPM is used for the subsequent video unit instead of the IPM used for the current video unit.
83. The method of claim 1, wherein a signal of an intra TMP prediction signal and a second predicted signal are fused by directly combining the intra TMP prediction signal and the second predicted signal based on position.
84. The method of claim 83, wherein the intra TMP prediction is applied and for at least one position the second prediction is applied.
85. The method of claim 1, wherein multiple fusion methods are applied to intra-frame TMP.
86. The method of claim 85, wherein which fusion method is used is predefined, or where which fusion method is used is indicated, or Which fusion method is used is derived.
87. The method of claim 1, wherein whether to apply the fusion of the intra-frame TPM mode and the codec tool for the video unit and / or the manner in which the fusion of the intra-frame TPM mode and the codec tool for the video unit is applied depends on codec information.
88. The method according to claim 87, wherein the codec information includes at least one of the following: Whether to allow intra-frame TMP or the intra-frame prediction method, Block dimensions, Block size, Block depth, Strip type, Image type, Split tree type, Time domain layer identification, block location, or Color component.
89. The method of claim 88, wherein the fusion of intra TPM mode with the codec tool is applied to an I slice or an I picture.
90. The method of claim 88, wherein the fusion of intra TPM mode with the codec tool is applied to all slice types or all picture types.
91. The method of claim 1, wherein the manner in which template matching is performed for a block depends on whether the block is to be fused by the intra TMP and a second prediction.
92. The method of claim 91, wherein the second prediction is taken into account during calculation of template costs.
93. The method of claim 1, wherein whether to apply the fusion of intra-TPM mode and the codec tool and / or the manner in which the fusion of intra-TPM mode and the codec tool is applied depends on at least one of: color format or color component.
94. The method of claim 93, wherein the fusion of intra TPM mode with the codec tool is applied to all color components.
95. The method of claim 93, wherein if the fusion of the intra TPM mode with the codec tool is applied to chroma components, the derivation of intra prediction differs from the derivation for luma components.
96. The method of claim 95, wherein the intra prediction is obtained using one of: CCLM, multi-model CCLM (MMLM), CCCM, chroma-DIMD, chroma-TIMD, a fusion of CCLM and angular mode, a fusion of MMLM and angular mode, or a fusion of CCCM and angular mode.
97. A method according to claim 93, wherein whether the fusion of intra-frame TPM mode and the codec tool is applied to the first component and / or the manner in which the fusion of intra-frame TPM mode and the codec tool is applied to the first component depends on whether the fusion of intra-frame TPM mode and the codec tool is applied to the second component.
98. The method of claim 97, wherein the first component comprises a chrominance component and the second component comprises a luma component.
99. The method of claim 97, wherein the fusion of intra TPM mode with the codec tool is applied to the first component in the same manner as applied to the second component.
100. The method of claim 97, wherein the fusion of intra TPM mode with the codec tool is applied differently to the first component than to the second component.
101. The method of claim 100, wherein the weighting parameters of the fusion are different.
102. The method of claim 93, wherein the fusion of intra TPM mode with the codec tool is applied to luma components but not to chroma components.
103. The method of claim 102, wherein the luma component comprises Y in a YCbCr color space or green (G) in a red, green, and blue (RGB) color space.
104. The method of claim 102, wherein the chrominance component comprises at least one of: Cb or Cr in a YCbCr color space, or The chrominance component includes at least one of the following: R or B in the RGB color space.
105. The method of any one of claims 1-104, wherein the indication of the fusion of intra-TPM mode with the codec tool is dynamically derived.
106. The method of any one of claims 1-104, wherein the indication of the fusion of intra-TPM mode with the codec tool is signaled based on a condition.
107. The method of claim 106, wherein the condition comprises at least one of the following: Whether to allow intra-frame TMP or the intra-frame prediction method, Block dimensions, Block size, Block depth, Strip type, Image type, Split tree type, Time domain layer identification, block location, or Color component.
108. The method of claim 107, wherein the fusion of intra TPM mode with the codec tool is indicated for an I slice or an I picture.
109. The method of claim 107, wherein the fusion of intra TPM mode with the codec tool is indicated for all slice types or all picture types.
110. The method of claim 107, wherein the fusion of intra TPM mode with the codec tool is not signaled for a block located at the upper left corner of a slice or an upper left corner of a picture.
111. The method of claim 106, wherein if the indication of the fusion of intra TPM mode with the codec tool is not signaled, the indication is inferred to be a default value.
112. The method of claim 106, wherein if the indication of the fusion of intra TPM mode with the codec tool is not signaled, then the indication is presumed to be false.
113. The method of claim 106, wherein if the indication of the fusion of intra TPM mode with the codec tool is not signaled, then the indication is presumed to be true.
114. The method of any one of claims 1-106, wherein whether the current block is coded using the fusion of intra TPM mode with the codec tool is signaled using one or more syntax elements.
115. The method of claim 114, wherein the one or more syntax elements are binarized as one of: a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code.
116. The method of claim 114, wherein the one or more syntax elements are context-coded, or The one or more syntax elements are bypassed and decoded.
117. The method of claim 116, wherein the context depends on coded information.
118. The method of claim 117, wherein the encoded information comprises at least one of: Block dimensions, Block size, Strip type, Image type, Information about neighboring blocks, Information about other codecs being used for the current block, or Time domain layer information.
119. The method of claim 114, wherein if the video unit is intra TPM coded, the fusion of intra TPM mode with the codec is signaled.
120. The method of claim 114, wherein the one or more syntax elements are indicated at one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
121. The method of claim 114, wherein the one or more syntax elements are encoded in a predictive manner.
122. The method of claim 114, wherein the one or more syntax elements are predicted from one or more syntax elements of a neighboring block.
123. The method of any one of claims 1-122, wherein the video unit comprises at least one of: Color component, Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Tree Block (CTB), Codec Unit (CU), Codec Tree Unit (CTU), CTU line, CTU group, strips, piece, sub-images, piece, a subregion within a block, or An area that includes more than one sample or pixel.
124. The method of any one of claims 1-123, wherein an indication of whether and / or how to derive the prediction or reconstruction of the video unit based on the fusion of intra TPM mode with the codec tool is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
125. The method of any one of claims 1-123, wherein the indication of whether and / or how to derive the prediction or reconstruction of the video unit based on the fusion of intra TPM mode with the codec tool is indicated in one of the following: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
126. The method of any one of claims 1-123, further comprising: Determining whether and / or how to derive the prediction or reconstruction of the video unit based on the fusion of the intra TPM mode and the codec tool is based on at least one of the following: Message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, TU, PU block, video codec unit, The location of one of the following: CU, PU, TU, block, video codec unit, the block dimensions of the current block and / or the neighboring blocks of the current block, the block shapes of the current block and / or the neighboring blocks of the current block, the encoding and decoding mode of the video unit, Indication of color format, Codec tree structure, Strip type, Chipset type, Image type, Color component, Time domain layer identification, A grade, level, or tier of standard.
127. The method of any one of claims 1-126, wherein the converting comprises encoding the video unit into the bitstream.
128. The method of any one of claims 1-126, wherein the converting comprises decoding the video unit from the bitstream.
129. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-128.
130. A non-transitory computer-readable storage medium storing instructions for causing a processor to execute the method according to any one of claims 1-128.
131. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: Determine the integration of intra template matching prediction (intra TMP) mode with codec tools; Derivation of prediction or reconstruction of video units of the video based on the fusion of the intra TMP mode and the codec tool; as well as The bitstream is generated based on the prediction or reconstruction of the video unit.
132. A method for storing a bitstream of a video, comprising: Determine the integration of intra template matching prediction (intra TMP) mode with codec tools; Derivation of prediction or reconstruction of video units of the video based on the fusion of the intra TMP mode and the codec tool; generating the bitstream based on the prediction or reconstruction of the video unit; as well as The bitstream is stored in a non-transitory computer-readable recording medium.