Method and device for video processing and medium

By determining the content type of the video unit and performing corresponding conversion, the requirements for improving the encoding and codec quality in the prior art are solved, and the adaptive video content detection and encoding and codec tools are implemented to improve the encoding and codec quality.

CN120530618APending Publication Date: 2025-08-22DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480007649.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-13
Filing Date
2024-01-12
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing video encoding and codec technology has room for improvement in encoding and codec quality, especially in adaptive video content detection and codec tool control.

Method used

By determining the content type of the video unit, converting based on the sample point value associated with the video unit, adaptive video content detection is realized, and the encoding and decoding tool is controlled to improve the encoding and decoding quality.

Benefits of technology

Adaptive video content detection is realized, and the encoding and decoding tools are controlled based on the detected video content, thereby improving the encoding and decoding quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120530618A_ABST
    Figure CN120530618A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a current video unit of a video and a bitstream of the video, determining a content type of the current video unit based on a value of a sample associated with the current video unit; and performing the conversion based on the content type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to video content detection. Background Art

[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding quality of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided, comprising: determining, for conversion between a current video unit of a video and a bitstream of the video, a content type of the current video unit based on values ​​of samples associated with the current video unit; and performing conversion based on the content type.

[0005] According to the method of the first aspect of the present disclosure, the content type of a video unit is determined based on the values ​​of samples associated with the video unit. Compared to traditional solutions, the proposed method can advantageously implement adaptive video content detection and thus support the control of codec tools based on the detected video content. In this way, codec quality can be improved.

[0006] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, the instructions causing a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method includes: determining a content type of a current video unit of the video based on values ​​of samples associated with the current video unit; and generating a bitstream based on the content type.

[0009] In a fifth aspect, a method for storing a bitstream of a video is provided, comprising: determining a content type of a current video unit based on values ​​of samples associated with the current video unit; generating a bitstream based on the content type; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0013] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;

[0014] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0015] Figure 4 Intra prediction mode is shown;

[0016] Figure 5A The reference samples used for wide-angle intra prediction are shown;

[0017] Figure 5B The reference samples used for wide-angle intra prediction are shown;

[0018] Figure 6 The discontinuity problem is shown when the orientation exceeds 45°;

[0019] Figure 7A A schematic diagram showing the definition of samples used by PDPC applied to diagonal top right diagonal mode and adjacent angle intra mode;

[0020] Figure 7B A schematic diagram showing the definition of samples used by PDPC applied to the diagonal lower left diagonal mode and the adjacent angle intra mode;

[0021] Figure 7CA schematic diagram showing the definition of samples used by PDPC applied to diagonal adjacent top right diagonal mode and adjacent angle intra mode;

[0022] Figure 7D A schematic diagram showing the definition of samples used by PDPC applied to diagonal adjacent left-bottom diagonal mode and adjacent angle intra mode;

[0023] Figure 8 An example of four reference rows adjacent to a prediction block is shown;

[0024] Figure 9A A schematic diagram showing the process of sub-segmentation depending on the block size;

[0025] Figure 9B A schematic diagram showing the process of sub-segmentation depending on the block size;

[0026] Figure 10 shows the matrix-weighted intra prediction process;

[0027] Figure 11 The spatial GPM candidates are shown;

[0028] Figure 12 A GPM template is shown;

[0029] Figure 13 GPM mixing is shown;

[0030] Figure 14 Shows the location of the spatial merge candidate;

[0031] Figure 15 shows the candidate pairs considered for redundancy check of spatial Merge candidates;

[0032] Figure 16 FIG. 4 is a schematic diagram showing motion vector scaling for a temporal Merge candidate;

[0033] Figure 17 The candidate positions for the time domain Merge candidate, C0 and C1, are shown;

[0034] Figure 18 The MMVD search point is shown;

[0035] Figure 19 shows the extended CU area used in BDOF;

[0036] Figure 20 A schematic diagram for a symmetric MVD mode is shown;

[0037] Figure 21 Decoding side motion vector refinement is shown;

[0038] Figure 22 Shown are the top and left neighboring blocks used in the derivation of CIIP weights;

[0039] Figure 23 An example of GPM partitioning grouped at the same angle is shown;

[0040] Figure 24 shows the unidirectional prediction MV selection for geometric partitioning mode;

[0041] Figure 25 shows an example generation of blending weights w0 using geometric partitioning mode;

[0042] Figure 26 The current CTU processing order and its available reference samples in the current CTU and the left CTU are shown;

[0043] Figure 27 The residual encoding and decoding passes for a transform skip block are shown;

[0044] Figure 28 An example of a block encoded and decoded in palette mode is shown;

[0045] Figure 29 shows sub-block based index map scanning for a palette, left for horizontal scanning and right for vertical scanning;

[0046] Figure 30 Shown is a decoding flow chart using ACT;

[0047] Figure 31 The intra-frame template matching search area used is shown;

[0048] Figure 32 Five locations in the reconstructed luminance samples are shown;

[0049] Figure 33 The prediction process of DBV mode is shown;

[0050] Figure 34 The low-frequency non-separable transform (LFNST) process is shown;

[0051] Figure 35 The SBT position, type and transformation type are shown;

[0052] Figure 36 The ROI for LFNST16 is shown;

[0053] Figure 37 The ROI for LFNST8 is shown;

[0054] Figure 38 Discontinuity measurements are shown;

[0055] Figure 39 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and

[0056] Figure 40 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0057] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION

[0058] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0059] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0060] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.

[0061] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0062] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment

[0063] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0064] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0065] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0066] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.

[0067] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0068] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.

[0069] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0070] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0071] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0072] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2are shown separately in the example.

[0073] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0074] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0075] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0076] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0077] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0078] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0079] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0080] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0081] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0082] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0083] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0084] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0085] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0086] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0087] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0088] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0089] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0090] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0091] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.

[0092] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0093] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0094] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0095] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0096] Motion compensation unit 302 may calculate interpolated values ​​for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.

[0097] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.

[0098] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0099] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0100] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to video codec technology. Specifically, it relates to codec tool on / off control and related technologies in image / video codecs. It can be applied to existing video codec standards such as HEVC and VVC. It is also applicable to future video codec standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly, and at the April 2018 JVET meeting, the new video codec standard was officially named the Versatile Video Codec (VVC), and the first version of the VVC Test Model (VTM) was released. The VVC working draft and the VTM test model are subsequently updated after each meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. 2.1 Existing Codec Tools 2.1.1 Intra-frame prediction 2.1.1.1 Intra-mode codec with 67 intra-prediction modes To capture arbitrary edge directions present in natural videos, the number of directional intra modes in VVC is extended from 33 used in HEVC to 65. New directional modes not in HEVC are Figure 4 These are depicted as red dashed arrows in , and the planar and DC modes remain the same. These more dense directional intra prediction modes apply to all block sizes and both luma intra prediction and chroma intra prediction. In VVC, for non-square blocks, several normal-angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes. In HEVC, each intra-coded block has a square shape, and the length of each of its sides is a power of 2. Therefore, no split operation is required to generate intra prediction values ​​using DC mode. In VVC, blocks can have a rectangular shape, which generally requires the use of a split operation for each block. To avoid split operations for DC prediction, only the longer sides are used to calculate the average value of non-square blocks. 2.1.1.2 Intra-mode encoding and decoding Figure 4 Intra prediction modes are shown. In order to keep the complexity of the most probable mode (MPM) list generation low, an intra mode codec with 6 MPMs is used by considering two available adjacent intra modes. The following three aspects are considered when constructing the MPM list: – Default intra mode; – Neighborhood intra mode; – Derived intra mode. A unified 6MPM list is used for intra blocks, regardless of whether MRL and ISP codecs are applied. The MPM list is constructed based on the intra modes of the left and above neighboring blocks. Assuming the mode on the left is denoted as Left and the mode of the above block is denoted as Above, the unified MPM list is constructed as follows: – When a neighboring block is not available, its intra mode is set to planar by default. – If both the Left mode and the Above mode are non-angle modes: – MPM list → {plane, DC, V, H, V-4, V+4}. – If one of the Left and Above modes is an angular mode and the other is a non-angular mode: – Set Mode Max to the larger of Left and Above. – MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2}. – If Left and Above are both angles and they are different: – Set Mode Max to the larger of Left and Above. – If the difference between the pattern Left and the pattern Above is in the range 2 to 62 (inclusive): – MPM list → {Plane, Left, Above, DC, Max-1, Max+1}. -otherwise: – MPM list → {Flat, Left, Above, DC, Max-2, Max+2}. – If Left and Above are both angles and they are the same: – MPM list → {plane, left, left-1, left+1, DC, left-2}. In addition, the first binary bit of the mpm index codeword is CABAC context coded. A total of three contexts are used, corresponding to whether the current intra block is MRL-enabled, ISP-enabled, or a normal intra block. During the 6MPM list generation process, deduplication is used to remove repeated patterns so that only unique patterns can be included in the MPM list.For entropy coding of the 61 non-MPM modes, truncated binary code (TBC) is used. 2.1.1.3 Wide-angle Intra Prediction for Non-square Blocks Conventional angular intra prediction directions are defined as running from 45 degrees to -135 degrees clockwise. In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original mode index, which is remapped to the wide-angle mode index after parsing. The total number of intra prediction modes remains unchanged at 67, and the intra mode encoding and decoding method remains unchanged. To support these predictions, Figure 5A and Figure 5B As shown, a top reference with a length of 2W+1 and a left reference with a length of 2H+1 are defined. The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra-frame prediction modes are shown in Table 2-1. Table 2-1 – Intra prediction modes replaced by Wide mode Aspect ratio Replaced intra prediction mode W / H==16 Modes 12, 13, 14, and 15 W / H==8 Modes 12 and 13 W / H==4 Mode 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H==2 Modes 2, 3, 4, 5, 6, 7 W / H==1 none W / H==1 / 2 Modes 61, 62, 63, 64, 65, 66 W / H==1 / 4 Modes 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 W / H==1 / 8 Modes 55 and 56 W / H==1 / 16 Modes 53, 54, 55, 56 like Figure 6 As shown in Figure 2, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap Δp. α negative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted using these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the non-fractional mode in the conventional prediction mode with the wide-angle mode. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​of the entries of the mapping table to more accurately convert the prediction angles for chroma blocks. 2.1.1.4 Mode-Dependent Intra-Frame Smoothing (MDIS) A four-tap intra interpolation filter is utilized to improve the accuracy of directional intra prediction. In HEVC, a two-tap linear interpolation filter has been used to generate intra prediction blocks in directional prediction mode (i.e., excluding planar and DC prediction values). In VVC, a simplified 6-bit 4-tap Gaussian interpolation filter is used only for directional intra mode. The non-directional intra prediction process is not modified. Selection of the 4-tap filter is performed according to the MDIS condition for directional intra prediction modes that provide non-fractional displacement (i.e., all directional modes except the following modes: 2, HOR_IDX, DIA_IDX, VER_IDX, 66). Depending on the intra prediction mode, the following reference sample processing is performed: – Directional intra prediction modes are classified into one of the following groups: – vertical mode or horizontal mode (HOR_IDX, VER_IDX), – diagonal mode (2, DIA_IDX, VDIA_IDX) that represents angles that are multiples of 45 degrees, – Remaining direction mode; – If the directional intra prediction mode is classified as belonging to group A, no filter is applied to the reference samples to generate the prediction samples; – Otherwise, if the mode falls into group B, the [1,2,1] reference sample filter may (depending on the MDIS condition) be applied to the reference samples to further copy these filtered values ​​into the intra prediction values ​​according to the selected direction, but no interpolation filter is applied; Otherwise, if the mode is classified as belonging to group C, only the intra reference sample interpolation filter is applied to the reference samples to generate prediction samples that fall between the reference samples at fractional or integer positions according to the selected direction (no reference sample filtering is performed). 2.1.1.5 Position-dependent intra prediction combinations In VVC, the results of intra prediction for DC, planar, and several angular modes are further modified by the Position Dependent Intra Prediction Combining (PDPC) method. PDPC is an intra prediction method that uses a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, bottom-left angular mode and its eight adjacent angular modes, and top-right angular mode and its eight adjacent angular modes. The prediction sample pred(x', y') is predicted using the intra prediction mode (DC, planar, angular) and a linear combination of the reference samples according to the following equations 3-8: pred(x',y')=(wL×R -1,y’+wT×R x’,-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x',y')+32)>>6 (2-1) where R x,-1 、R -1,y Respectively represent the reference sample points at the top and left boundaries of the current sample point (x, y), and R -1,-1 Indicates the reference sample located at the upper left corner of the current block. If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filters are required, as in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process for DC mode and planar mode is the same, avoiding the clipping operation. For angular mode, the PDPC scaling factor is adjusted so that no range check is required, and the angle condition for enabling PDPC is removed (using scale>=0). In addition, in all angular mode cases, the PDPC weights are based on 32. The PDPC weights depend on the prediction mode, as shown in Table 2-2. PDPC is applied to blocks with a width and height both greater than or equal to 4. Figures 7A-7D The reference samples (R x,-1 、R -1,y and R -1,-1 ) definition. The prediction sample point pred(x', y') is located at (x', y') in the prediction block. For example, the reference sample point R x,-1 The coordinate x of is given by: x=x'+y'+1, and the reference point R -1,y The coordinate y of is similarly given by: y=x'+y'+1 (for diagonal mode). For other angle modes, the reference point R x,-1 and R -1,y Can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position is used. Table 2-2 - Example of PDPC weights according to prediction mode Prediction Model wT wL wTL Diagonal upper right 16>>((y’<<1)>>shift) 16>>((x’<<1)>>shift) 0 Diagonal lower left 16>>((y’<<1)>>shift) 16>>((x’<<1)>>shift) 0 Adjacent diagonal upper right 32>>((y’<<1)>>shift) 0 0 Adjacent diagonal lower left 0 32>>((x’<<1)>>shift) 0 2.1.1.6 Multiple Reference Line (MRL) Intra Prediction Multiple Reference Line (MRL) intra prediction uses more reference lines for intra prediction. Figure 8 In

[15] , an example of 4 reference lines is depicted, where the samples of segments A and F are not obtained from reconstructed neighboring samples, but are filled with the nearest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used. The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row idx greater than 0, only additional reference row modes are included in the MPM list, and only the mpm index is signaled, while the remaining modes are not signaled. The reference row index is signaled before the intra prediction mode, and if a non-zero reference row index is signaled, planar mode is excluded from the intra prediction mode. MRL is disabled for the first row of blocks inside a CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in DC intra prediction mode for non-zero reference row index is aligned with its derivation for reference row index 0. MRL requires storage of 3 adjacent luma reference rows along with the CTU to generate the prediction. The Cross Component Linear Model (CCLM) tool also requires 3 adjacent luma reference rows for its downsampling filter. The definition of MLR using the same 3 rows is aligned with CCLM which reduces the storage requirements for the decoder. 2.1.1.7 Intra-frame Sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into 4 sub-partitions. It has been noted that M×128 (with M≤64) and 128×N (with N≤64) ISP blocks may generate potential problems with 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into four M×32TBs (only horizontal division is possible), each TB is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, both chroma components will have a size larger than a 32×32 block. Similarly, a 128×NCU using ISP can produce a similar situation. Therefore, these two situations are problems for a 64×64 decoder pipeline. Therefore, the CU size that can use ISP is limited to a maximum of 64×64. Figure 9A and Figure 9B Examples of two possibilities are shown. Figure 9A Examples of sub-partitioning for 4x8 and 8x4 CUs are shown, and Figure 9B Examples of sub-partitions for CUs other than 4x8, 8x4, and 4x4 are shown. All sub-partitions satisfy the condition of having at least 16 samples. In ISP, 1xN / 2xN sub-block predictions are not allowed to rely on the reconstructed values ​​of previously decoded 1xN / 2xN sub-blocks of the codec block, resulting in a minimum prediction width of four samples for each sub-block. For example, an 8xN (N>4) codec block encoded using ISP with vertical partitioning is partitioned into two prediction regions of size 4xN each and four transforms of size 2xN. Furthermore, a 4xN codec block encoded using ISP with vertical partitioning is predicted using a full 4xN block; four transforms are used, each of size 1xN. Although 1xN and 2xN transform sizes are allowed, transforms for these blocks in a 4xN region can be performed in parallel. For example, when a 4xN prediction region contains four 1xN transforms, no transform is performed horizontally; the transform in the vertical direction can be performed as a single 4xN transform in the vertical direction. Similarly, when a 4xN prediction region contains two 2xN transform blocks, the transform operations for the two 2xN blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, processing these smaller blocks does not increase latency compared to processing intra blocks for 4x4 regular codecs. Table 2-3 – Entropy codec coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization and inverse transformation. Therefore, the reconstructed sample values ​​of each sub-partition can be used to generate a prediction for the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the sub-partition containing the upper left sample of the CU, and then continues downwards (horizontal partitioning) or to the right (vertical partitioning). As a result, the reference samples used to generate the sub-partition prediction signal are only located to the left and above the row. All sub-partitions share the same intra mode. The following is an overview of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: The size of the entropy coding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 2-3. Note that the new size only affects blocks generated by ISP where one of the dimensions is less than 4 samples. In all other cases, the coefficient group remains 4×4 in size. –CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions and the first n-1 subpartitions have yielded zero CBF, the CBF of the nth subpartition is assumed to be 1. – MPM usage: In blocks encoded and decoded by ISP mode, the MPM flag will be inferred to be one, and the MPM list is modified to exclude DC mode and give priority to horizontal intra mode for ISP horizontal partitioning and vertical intra mode for vertical partitioning. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. –PDPC: When the CU uses ISP codec mode, the PDPC filter will not be applied to the resulting sub-partition. –MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and the flag will not be sent to the decoder. Therefore, the encoder will not perform RD tests for the different available transforms for each resulting sub-split. Instead, the transform selection for ISP mode will be fixed and selected based on the utilized intra mode, processing order, and block size. Therefore, no signaling is required. For example, let t H and t V are the horizontal and vertical transforms selected for the w×h sub-partition, respectively, where w is the width and h is the height. The transforms are selected according to the following rules: If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w=2 or w>32, then t H =DCT-II. – If h = 2 or h > 32, then t V =DCT-II. – Otherwise, the transformation is selected as in Table 2-4. Table 2-4 – Transform selection depending on intra mode In ISP mode, all 67 intra modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the conditions for intra interpolation filter selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.1.1.8 Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a newly added intra prediction technology in VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed adjacent boundary samples on the left side of the block and a row of W reconstructed adjacent boundary samples above the block as input. If the reconstructed samples are not available, the reconstructed samples are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation, as shown in Figure 10 shown. ●Average of neighboring points Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging the adjacent boundary samples according to a predefined rule depending on the block size. top and bdry left Shrinking to smaller boundaries and Then, the two shrinking boundaries and Spliced ​​to the reduced boundary vector bdry red , so for a block of shape 4×4, bdry red Size is 4, and for all other block shapes, bdry red The size is 8. If mode refers to MIP mode, the splicing is defined as follows. Matrix multiplication Using the averaged samples as input, a matrix-vector multiplication is performed and then an offset is added. The result is a scaled-down prediction signal on a downsampled set of samples in the original block. From the scaled-down input vector bdry red Generate a reduced prediction signal pred red ,,pred red , is the width W red And the height is H red Here, W red and W red is defined as: Reduced prediction signal pred red It is calculated by taking the matrix-vector product and adding the offset: pred red =A·bdry red +b. Here, if W=H=4, then A is a red ·H red rows and 4 columns, and in all other cases a matrix with 8 columns. b is a matrix of size W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, and S2. An index idx = idx(W,H) is defined as follows: Here, each coefficient of matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices and 16 offset vectors Each matrix has 16 rows and 4 columns, and each offset vector has a size of 16. The matrices and offset vectors of this set are used for blocks of size 4×4. Set S1 consists of 8 matrices and 8 offset vectors Each matrix has 16 rows and 8 columns, and each offset vector has a size of 16. Set S2 consists of 6 matrices and 6 offset vectors Each matrix has 64 rows and 8 columns, and each offset vector has a size of 64. Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the downsampled set by linear interpolation, which is a single-step linear interpolation in each direction. Regardless of the block shape or block size, the interpolation is performed first in the horizontal direction and then in the vertical direction. ●MIP mode signaling and coordination with other codec tools For each codec unit (CU) in intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, a transposed flag (isTransposed) that determines whether the mode is transposed, and a MIP mode Id (modeId) that determines which matrix to use for a given MIP mode are derived as follows: isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-6) The MIP codec mode is coordinated with other codec tools by taking into account the following aspects: – For MIPs on large blocks, LFNST is enabled. Here, the planar LFNST transform is used. - Reference sample derivation for MIP is performed exactly the same as for regular intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled reference samples. – Clipping is performed before upsampling, rather than after upsampling. – Regardless of the maximum transform size, MIPs are limited to 64x64. - For sizeld=0, the number of MIP modes is 32, for sizeld=1, the number of MIP modes is 16, and for sizeld=2, the number of MIP modes is 12. 2.1.1.9 Airspace GPM (SGPM) In the spatial domain GPM, a candidate list including partitioning and two intra prediction modes is constructed. An MPM with no more than 11 intra prediction modes is used to form a combination, and the length of the candidate list is set to be equal to 16. The selected candidate index is transmitted through the signal. use Figure 11 The templates shown reorder the list. The GPM blending process is not used in the templates, and the SAD between the prediction and reconstruction of the templates is used for sorting. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in inter-frame GPM. Consider the following projects: ●Airspace GPM segmentation mode: 26 predefined modes. An adaptive inference algorithm based on the ratio of horizontal gradient to vertical gradient. ●Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation mode, an IPM list is derived for each part using the intra-inter GPM list. The IPM list size is 3. In the list, the TIMD-derived pattern is replaced by 2 derived patterns with horizontal and vertical directions (using the top or left template), or the TIMD-derived pattern is excluded. MPM List: A unified MPM list (maximum 11 elements) is used for all segmentation modes. ● Template size (left and top): 1 or 4. ●Extended block size: Spatial GPM is extended to be further applied to 4x8, 8x4, 4x16 and 16x4 blocks, which can be described as 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. Adaptive Hybrid: Adaptive mixing is tested for spatial GPM, where the mixing depth τ is derived as follows: ■If min(width, height) == 4, then 1 / 2τ is selected. ■ Otherwise, if min(width, height) == 8, then τ is selected. ■ Otherwise, if min(width, height) == 16, then 2τ is selected. ■ Otherwise, if min(width, height) == 32, then 4τ is selected. ■Otherwise, 8τ is selected. 2.1.2 Inter-frame prediction For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, as well as additional information required by the new codec features of VVC that will be used for inter-prediction sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector deltas or reference picture indices. A Merge mode is specified whereby the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and the additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list and other required information are explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC also includes many new and refined inter-frame prediction coding tools listed below: – Extended Merge predictions. –Merge mode with MVD (MMVD). – Symmetrical MVD (SMVD) signaling. – Affine motion compensated prediction. – Sub-block based temporal motion vector prediction (SbTMVP). – Adaptive Motion Vector Resolution (AMVR). – Motion field storage: 1 / 16th luminance sample MV storage and 8x8 motion field compression. – Bidirectional prediction with CU-level weighting (BCW). – Bidirectional Optical Flow (BDOF). – Decoder-side motion vector refinement (DMVR). – Geometric Partitioning Mode (GPM). – Joint Intra-frame and Inter-frame Prediction (CIIP). The following text provides details about these inter prediction methods specified in VVC. 2.1.2.1 Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Spatial MVP from spatially adjacent CUs. 2) Temporal MVP from the co-located CU. 3) History-based MVP from FIFO table. 4) Average MVP in pairs. 5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU codec in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded using context, and bypass coding is used for the remaining binary bits. This section provides the derivation process for each category of Merge candidates. As done in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a specific size. 2.1.2.1.1 Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. At most four Merge candidates are selected from the candidates at the positions shown in the figure above. The derivation order is B 0, 、A 0, 、B 1, , A1 and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1 and A1 are not available (for example, because it belongs to another slice or piece) or is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the list, so that the coding efficiency is improved. In order to reduce computational complexity, not all possible candidate pairs are considered in the redundancy check mentioned. Instead, only the pairs linked by arrows in the figure below are considered, and the candidate is added to the list only when the corresponding candidates used for the redundancy check do not have the same motion information. 2.1.2.1.2 Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. As shown by the dotted line in the figure below, the scaled motion vector of the temporal Merge candidate is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set equal to zero. As shown in the figure, the position of the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. 2.1.2.1.3 Merge Candidate Derivation Based on History History-based MVP (HMVP) Merge candidates are added to the Merge list after spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (cleared). As long as there is a non-subblock inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that no more than 6 historically based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find whether there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. Redundancy check is applied to HMVP candidates for spatial or temporal Merge candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: 1. The number of HMPV candidates used for Merge list generation is set to (N<=4) × M: (8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process starting from HMVP is terminated. 2.1.2.1.4 Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list. The predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge index of the merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference pictures. If only one motion vector is available, it is used directly. If no motion vector is available, the list remains invalid. When the merge list is not full after adding pairwise average merge candidates, zero MVPs are inserted at the end until the maximum number of merge candidates is reached. 2.1.2.2Merge Estimation Area The Merge Estimation Region (MER) allows independent derivation of the Merge candidate lists of CUs in the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. In addition, the update process of the history-based motion vector prediction candidate list is only updated when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel), where (xCb, yCb) is the top left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled as log2_parallel_merge_level_minus2 in the sequence parameter set. 2.1.2.3 Merge Mode with MVD (MMVD) In addition to the Merge mode in which the implicitly derived motion information is directly used for prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is also introduced in VVC. The MMVD flag is transmitted by signal immediately after the Skip flag and Merge flag are sent to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it is further refined by signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. A merge candidate flag is signaled to specify which candidate is used. The distance index specifies the motion magnitude and indicates a predefined offset from the starting point. As shown in the figure above, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is shown in Table 2-5. Table 2-5 – Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 2-6. It should be noted that the meaning of the MVD symbol can change according to the information of the starting MV. When the starting MV is a unidirectionally predicted MV or a bidirectionally predicted MV in which the two lists point to the same side of the current picture (that is, the POCs of the two references are both greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 2-6 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectionally predicted MV with two MVs pointing to different sides of the current picture (that is, the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the symbols in Table 2-6 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 2-6 – Signs of MV offsets specified by direction index Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.4 Bidirectional Prediction with CU-Level Weighting (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-7) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectionally predicted CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights are conditionally checked only for 1-pixel and 4-pixel motion vector precision if the current picture is a low-latency picture. When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – Unequal weights are only conditionally checked when the two reference pictures in bidirectional prediction are the same. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using one context codec bit, followed by a bypass codec bit. The first context codec bit indicates whether equal weights are used; if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with cross-fading. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For MergeCUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, affine motion information is constructed based on motion information of no more than 3 blocks. The BCW index of a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded or decoded using CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.1.2.5 Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and multiplier size. BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: The CU is coded using “true” bi-prediction mode, ie, one of the two reference pictures precedes the current picture in display order, and the other follows the current picture in display order. – The distances from the two reference pictures to the current picture (i.e., the POC difference) are the same. – Both reference images are short-term reference images. –CU is not encoded or decoded using Affine mode or ATMVP Merge mode. – The CU has more than 64 luma samples. – CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. –Do not enable WP for the current CU. –CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 sub-block, the motion refinement (v x ,v y ). Then, motion refinement is used to adjust the bidirectional prediction sample values ​​in the 4x4 sub-block. The following steps are applied in the BDOF process. First, the horizontal and vertical gradients of the two predicted signals are calculated by directly calculating the difference between two adjacent samples. and k=0,1, that is, Among them I (k) (i, j) is the sample value of the prediction signal at coordinate (i, j) in list k, k=0,1, and shift1 is calculated based on the luma bit depth bitDepth as shift1=max(6, bitDepth-6). Then, the autocorrelations and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as: in where Ω is a 6x6 window around the 4x4 sub-block, and n a and n b The values ​​of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. Then, using the following equation, motion refinement (v x ,v y ) is derived using the cross-correlation and autocorrelation terms: in th′ BIO =2 max(5,BD-7) . is the rounding function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I(0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift (2-13) These values ​​are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, some predicted samples I in the list k outside the current CU boundary (k) (i,j)(k=0,1) needs to be generated. Figure 19 As shown, BDOF in VVC uses an extended row / column around the CU boundary. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly obtaining the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values ​​are only used in gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. In order to avoid the additional complexity of the SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.1.2.6 Symmetric MVD Encoding and Decoding In VVC, in addition to the normal uni-prediction and bi-prediction mode MVD signaling, a symmetric MVD mode for bi-prediction MVD signaling is also applied. In symmetric MVD mode, the motion information including the reference picture indices of both list 0 and list 1 and the MVD of list 1 is not transmitted through signaling but derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference index for list 0 and list 1 is set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown in the following formula. In the encoder, symmetric MVD motion estimation starts with initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for MVD motion search. 2.1.2.7 Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of MV in Merge mode, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In bidirectional prediction operation, the refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and list L1. Figure 21 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs that are coded or decoded using the following modes and features: – CU-level Merge mode with bi-predictive MV. - Relative to the current picture, one reference picture is in the past and the other reference picture is in the future. – The distances from the two reference pictures to the current picture (i.e., the POC difference) are the same. – Both reference images are short-term reference images. – The CU has more than 64 luma samples. – CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. – Do not enable WP for the current block. – CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction for future picture encoding and decoding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.1.2.7.1 Search Scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR (represented by the candidate MV pair (MV0, MV1)) obeys the following two equations: MV0′=MV0+MV_offset (2-15) MV1′=MV1-MV_offset (2-16) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainty, it is proposed to bias towards the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation, rather than an additional search using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates with the center having the minimum SAD in either the first or second iteration of the search. In the sub-pixel offset estimation based on the parameter error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit the two-dimensional parabolic error surface equation of the following form: E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (2-17) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values ​​of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-18) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-19) Since all cost values ​​are positive and the minimum is E(0,0), then x min and y min The value of is automatically constrained between -8 and 8. This corresponds to a half-pixel shift with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to get the sub-pixel accurate refinement delta MV. 2.1.2.7.2 Bilinear interpolation and sample filling In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the initial fractional pixel MV with integer sample offsets, so for the DMVR search process, those fractional samples need to be interpolated. To reduce computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples within the 2-sample search range compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. In order to not access more reference samples than the normal MC process, samples that are not required by the interpolation process based on the original MV but are required by the interpolation process based on the refined MV are padded from those available samples. 2.1.2.7.3 Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16. 2.1.2.8 Joint Intra-Frame and Inter-Frame Prediction (CIIP) In VVC, when a CU is encoded and decoded using Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the inter-intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter The intra prediction signal P is derived using the same inter prediction process applied to the conventional Merge mode; intra It is derived after the regular intra prediction process with planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values ​​are calculated according to the codec mode of the top neighbor block and the left neighbor block as follows: – If the top neighboring block is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighboring block is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) is equal to 2, then wt is set to 3; – Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, wt is set to 2; – Otherwise, set wt to 1. The CIIP forecast is formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (2-20) 2.1.2.9 Multiple Hypothesis Prediction (MHP) Except for Inter AMVP mode, normal Merge mode and MMVD mode, no more than two additional prediction values ​​are signaled. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. 2.1.2.10 Overlapped Sub-Block Motion Compensation (OBMC) When OBMC is applied, top and left boundary pixels of a CU are refined using motion information of neighboring blocks with weighted prediction. The conditions under which OBMC should not be applied are as follows: ●When OBMC is disabled at the SPS level. ●When the current block has intra mode or IBC mode. ●When the current block applies LIC. ●When the current luminance block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. To enable it for sub-block based codecs: ●Affine AMVP mode; Affine Merge mode and sub-block based temporal motion vector prediction (SbTMVP); ●Sub-block based bilateral matching. 2.1.2.11 Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its prediction block as a function of the local illumination variation between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, which form a linear equation, i.e., α*p[x]+β, to compensate for illumination variation, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When surround motion compensation is enabled, the MV should be clipped to account for the surround offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for AMVP mode to indicate the use of LIC. The local illumination compensation proposed in JVET-O0066 is used for unidirectionally predicted inter CUs with the following modifications. ● Neighboring samples within the frame can be used to derive LIC parameters; ● Disable LIC for blocks with fewer than 32 luma samples; For non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than based on the partial template block samples corresponding to the first top-left 16x16 unit; • The samples of the reference block template are generated by using MC with the block MV without rounding it to integer pixel precision. 2.1.2.12 Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode for inter-frame prediction is supported. The geometric partitioning mode is signaled as a Merge mode using a CU level flag. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. The geometric partitioning mode is for each possible CU size w×h=2 m ×2 n (where m,n∈{3…6} excludes 8x64 and 64x8) a total of 64 splits are supported. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Figure 23 ). The position of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each part has a motion vector and a reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, only two motion-compensated predictions are required for each CU. If geometric partitioning mode is used for the current CU, a geometric partitioning index and two Merge indices (one for each partition) indicating the partitioning mode (angle and offset) of the geometric partitioning are further transmitted by signal. The number of maximum GPM candidate sizes is explicitly transmitted by signal in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, a hybrid process with adaptive weights is used to adjust the sample values ​​along the geometric partitioning edge. This is a prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.2.12.1 One-way prediction candidate list construction The unidirectional prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are Figure 24 In the case where the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partition mode. 2.1.2.12.2 Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. The blending weight for each position of the CU is derived based on the distance between the individual position and the partition edge. The distance from position (x,y) to the segmentation edge is derived as: where i,j are the indices of the angle and offset for the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j The sign of depends on the angle index i. The weight of each part of the geometric segmentation is derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-25) w1(x,y)=1-w0(x,y) (2-27) partIdx depends on the angle index i. An example of the weight w0 is shown below. 2.1.2.12.3 Motion Field Storage for Geometric Partitioning Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and Mv which is a combination of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-28) where motionIdx is equal to d(4x+2,4y+2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, then Mv, a combination of Mv0 and Mv2, is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.1.2.12.4 GPM with Inter and Intra Prediction (GPM Inter-Intra) With GPM Inter-Intra, in addition to the Merge candidates for each non-rectangular partitioned area in the CU to which GPM is applied, a predefined intra prediction mode for the geometric partition line can also be selected. In the proposed method, for each GPM separable area with a flag from the encoder, it is determined whether it is an intra prediction mode or an inter prediction mode. When it is an inter prediction mode, the unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, when it is an intra prediction mode, the unidirectional prediction signal is generated from the neighboring pixels for the intra prediction mode specified by the index from the encoder. The variation of possible intra prediction modes is limited by the geometric shape. Finally, the two unidirectional prediction signals are mixed in the same way as ordinary GPM. 2.1.3 Screen Content Encoding and Decoding Tools 2.1.3.1 Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. As we all know, IBC significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra prediction mode or the inter prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4x4 sub-block. For larger current block sizes, a hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as either IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC coded blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if encoded with IBC). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.1.3.1.1 IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows reconstructed portions of a predefined area, which includes the area of ​​the current CTU and some areas of the left CTU. Figure 26 The reference area of ​​the IBC mode is shown, where each block represents a 64x64 luma sample unit. Depending on the location of the current codec CU position within the current CTU, the following applies: – If the current block falls into the upper left 64x64 block of the current CTU, in addition to the reconstructed samples in the current CTU, the CPR mode can also be used to reference the reference samples in the lower right 64x64 block of the left CTU. The current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU. – If the current block falls into the upper right 64x64 block of the current CTU, in addition to the samples that have been reconstructed in the current CTU, if the luma position (0, 64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls into the lower left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU. Otherwise, the current block can also use the CPR mode to refer to the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls into the lower right 64x64 block of the current CTU, the current block can only refer to the reconstructed samples in the current CTU using the CPR mode. This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementations. 2.1.3.1.2 Interaction between IBC and other codecs The interaction between IBC mode and other inter-frame codec tools in VVC (such as paired merge candidates, history-based motion vector predictor (HMVP), intra / inter joint prediction mode (CIIP), merge mode with motion vector difference (MMVD) and geometric partition mode (GPM)) is as follows: – IBC can be used with pairwise merge candidates and HMVP. A new pairwise IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference. – IBC cannot be used in conjunction with the following interframe tools: affine motion, CIIP, MMVD, and GPM. – When DUAL_TREE partitioning is used, IBC is not allowed for chroma codec blocks. Unlike the HEVC screen content codec extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation process of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects are applied: – IBC shares the same process as regular MV Merge, including with paired merge candidates and history-based motion prediction values, but TMVP and zero vectors are not allowed because they are not valid for IBC mode. – Separate HMVP buffers (5 candidates each) are used for regular MV and IBC. – Block vector constraints are implemented as bitstream consistency constraints. The encoder needs to ensure that no invalid vectors exist in the bitstream, and if the merge candidate is invalid (out of range or 0), then the merge should not be used. This bitstream consistency constraint is expressed in terms of virtual buffers, as described below. – For deblocking, IBC is handled as inter mode. If the current block is coded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to only indicate whether the MV is inter pixels or 4 integer pixels. - The number of IBC Merge candidates may be signaled in the slice header separately from the number of normal Merge candidates, sub-block Merge candidates, and geometric Merge candidates. The concept of a virtual buffer is used to describe the allowable reference area and valid block vectors for the IBC prediction mode. Denoting the CTU size as ctbSize, the virtual buffer ibcBuf has a width of wIbcBuf = 128x128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows. – At the beginning of decoding each CTU line, flush the entire ibcBuf with an invalid value of -1. – At the start of decoding the VPDU (xVPDU, yVPDU) relative to the upper left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU% wIbcBuf, ..., xVPDU% wIbcBuf+W v -1;y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1. – After decoding the CU containing (x, y) relative to the top left corner of the picture, set ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]. For a block covering coordinates (x, y), the block is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, the block is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1. 2.1.3.2 Block Differential Pulse Coded Modulation (BDPCM) VVC supports Block Differential Pulse Coded Modulation (BDPCM) for screen content encoding and decoding. At the sequence level, the BDPCM enable flag is signaled in the SPS; this flag is only signaled when transform skip mode (described in the next section) is enabled in the SPS. When BDPCM is enabled, if the CU size in terms of luma samples is less than or equal to MaxTsSize by MaxTsSize, and if the CU is intra-coded, a flag is transmitted at the CU level, where MaxTsSize is the maximum block size allowed for transform skip mode. The flag indicates whether regular intra-coding or BDPCM is used. If BDPCM is used, a BDPCM prediction direction flag is transmitted to indicate whether the prediction is horizontal or vertical. The block is then predicted using a regular horizontal intra-frame prediction process or a vertical intra-frame prediction process with unfiltered reference samples. The residuals are quantized, and the difference between each quantized residual and its predicted value (i.e., the previously coded residual of the neighboring position, either horizontally or vertically (depending on the BDPCM prediction direction)) is coded. For a block of size M (height) × N (width), let r i,j ,0≤i≤M-1,0≤j≤N-1 is the prediction residual. Let Q(r i,j ), 0≤i≤M-1,0≤j≤N-1 represents the residual r i,j quantized version of . BDPCM is applied to the quantized residual values, and the result is a quantized version of The modified M×N array in is predicted from its neighboring quantized residual values. For vertical BDPCM prediction mode, for 0≤j≤(N-1), the following is used to derive For the horizontal BDPCM prediction mode, for 0≤i≤(M-1), the following is used to derive At the decoder side, the above process is reversed to calculate Q(r i,j ), 0≤i≤M-1,0≤j≤N-1, as follows: If vertical BDPCM is used (2-31) If horizontal BDPCM is used (2-32) Dequantized residual Q -1 (Q(r i,j )) is added to the intra block prediction value to produce the reconstructed sample value. Predicted quantized residual value The residual codec is sent to the decoder using the same residual codec process as in transform skip mode residual codec. For lossless codecs, if slice_ts_residual_coding_disabled_flag is set to 1, the quantized residual values ​​are sent to the decoder using regular transform residual codec. For MPM modes for future intra mode codecs, if the BDPCM prediction direction is horizontal or vertical, the horizontal prediction mode or vertical prediction mode is stored for the BDPCM-coded CU, respectively. For deblocking, if both blocks on either side of a block boundary are coded using BDPCM, then that particular block boundary is not deblocked. 2.1.3.3 Residual Coding and Decoding for Transform Skip Mode VVC allows transform skip mode to be used for luminance blocks of size not exceeding MaxTsSize by MaxTsSize, where the value of MaxTsSize is signaled in the PPS and can be up to 32. When a CU is encoded and decoded in transform skip mode, its prediction residual is quantized and encoded using the transform skip residual encoding and decoding process. This process is modified from the transform coefficient encoding and decoding process. In transform skip mode, the residual of the TU is also encoded and decoded in units of non-overlapping sub-blocks of size 4x4. For better coding and decoding efficiency, some modifications are made to customize the residual coding and decoding process to the characteristics of the residual signal. The following summarizes the differences between transform skip residual coding and conventional transform residual coding: – Forward scan order is applied to scan sub-blocks within the transform block and positions within the sub-block; – no signaling of the final (x, y) position; – When all previous flags are equal to 0, coded_sub_block_flag is coded for each sub-block except the last sub-block; –sig_coeff_flag context modeling uses a reduced template, and the context model of sig_coeff_flag depends on the top neighboring value and the left neighboring value; – The context model of the abs_level_gt1 flag also depends on the sig_coeff_flag value on the left and the sig_coeff_flag value on the top; –par_level_flag uses only one context model; – Additional flags greater than 3, 5, 7, 9 are signaled to indicate coefficient levels, one context per flag; – The binarization of the remainder values ​​is derived using a Rice parameter with a fixed order of 1; – The context model of the sign flag is determined based on the left neighbor and the above neighbor, and the sign flag is parsed after sig_coeff_flag to keep all context-coded bins together. For each subblock, if coded_subblock_flag is equal to 1 (i.e. there is at least one non-zero quantized residual in the subblock), the encoding and decoding of the quantized residual level is performed in three scanning passes (see Figure 27 ): - First scan pass: The significance flag (sig_coeff_flag), the sign flag (coeff_sign_flag), the flag that the absolute level is greater than 1 (abs_level_gtx_flag[0]), and the parity check (par_level_flag) are encoded and decoded. For a given scan position, if sig_coeff_flag is equal to 1, then coeff_sign_flag is encoded and decoded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[0] is equal to 1, then par_level_flag is additionally encoded and decoded to specify the parity check of the absolute level. - Scan passes greater than x: For each scan position where the absolute level is greater than 1, no more than four abs_level_gtx_flag[i] (i=1...4) are encoded to indicate whether the absolute level at the given position is greater than 3, 5, 7 or 9, respectively. – Remainder scan pass: The remainder of the absolute level abs_remainder is encoded and decoded in bypass mode. The remainder of the absolute level is binarized using a fixed Rice parameter value of 1. The bins in scan pass #1 and scan pass #2 (the first scan pass and scan passes greater than x) are context-coded until the maximum number of context-coded bins in the TU has been exhausted. The maximum number of context-coded bins in the residual block is limited to 1.75*block_width*block_height, or equivalently, an average of 1.75 context-coded bins per sample position. The bins in the last scan pass (the remainder scan pass) are bypass-coded. The variable RemCcbs is first set to the maximum number of context-coded bins for the block and is decremented by 1 each time a context-coded bin is coded. When RemCcbs is greater than or equal to 4, the syntax elements in the first coding pass (including sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag) are coded using context-coded bins. If RemCcbs becomes smaller than 4 when encoding and decoding the first pass, the remaining coefficients that have not been encoded in the first pass are encoded in the remainder scanning pass (pass #3). After completing the first pass of encoding and decoding, if RemCcbs is greater than or equal to 4, the syntax elements in the second encoding and decoding pass (including abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag) are encoded using the bins of context encoding. If RemCcbs becomes less than 4 during the second pass of encoding and decoding, the remaining coefficients that have not been encoded in the second pass are encoded in the remainder scan pass (pass #3). Figure 27 The transform skip residual coding process is shown. The asterisks mark the locations where the context coded bits are exhausted, at which point all remaining bits are coded using the bypass codec. In addition, for blocks that are not coded in BDPCM mode, a level mapping mechanism is applied to transform skip residual coding until the maximum number of bins for context coding is reached. Level mapping uses the top neighboring coefficient level and the left neighboring coefficient level to predict the current coefficient level to reduce signaling cost. For a given residual position, denote absCoeff as the absolute coefficient level before mapping and absCoeffMod as the coefficient level after mapping. Let X0 denote the absolute coefficient level of the left neighboring position and let X1 denote the absolute coefficient level of the upper neighboring position. Level mapping is performed as follows: The absCoeffMod value is then encoded as described above.After all context encoded bins are exhausted, level mapping is disabled for all remaining scan positions in the current block. 2.1.3.4 Palette Mode In VVC, palette mode is used for encoding and decoding screen content in all chroma formats supported in the 4:4:4 profile (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome). When palette mode is enabled, if the CU size is less than or equal to 64x64 and the number of samples in the CU is greater than 16, a flag is transmitted at the CU level to indicate whether palette mode is used. Considering that applying palette mode to small CUs introduces insignificant coding gain and brings additional complexity to small blocks, palette mode is disabled for CUs with less than or equal to 16 samples. A palette-encoded codec unit (CU) is treated as a prediction mode different from intra prediction, inter prediction, and intra block copy (IBC) mode. If palette mode is used, the sample values ​​in the CU are represented by a set of representative color values. This set of representative color values ​​is called a palette. For positions where the sample values ​​are close to the palette colors, the palette index is transmitted through the signal. Samples outside the palette can also be specified by signaling a jump symbol. For samples encoded and decoded using a jump symbol within the CU, the component values ​​of these samples are directly signaled using (possibly) quantized component values. This is in Figure 28 The quantized escape symbols are binarized using a fifth-order Exp-Golomb binarization process (EG5). In order to encode and decode the palette, the palette prediction value is maintained. In the non-wavefront case, the palette prediction value is initialized to 0 at the beginning of each slice. For the WPP case, the palette prediction value at the beginning of each CTU row is initialized to the prediction value derived from the first CTU in the previous CTU row, so that the initialization scheme between the palette prediction value and the CABAC synchronization is unified. For each entry in the palette prediction value, the reuse flag is transmitted by signal to indicate whether it is part of the current palette in the CU. The reuse flag is sent using a run length codec of zero. Afterwards, the number of new palette entries and the component values ​​of the new palette entries are transmitted by signal. After encoding a palette-encoded CU, the palette prediction value will be updated using the current palette, and entries from previous palette prediction values ​​that have not been reused in the current palette will be added to the end of the new palette prediction value until the maximum allowed size is reached. A jump flag is transmitted by signal for each CU to indicate whether a jump symbol exists in the current CU. If a tab symbol exists, the palette table is incremented by one and the last index is assigned as the tab symbol. In a manner similar to the coefficient groups (CGs) used in transform coefficient coding, a CU coded using palette mode is divided into multiple row-based coefficient groups, each row-based coefficient group consisting of m (i.e., m=16) samples, where for each CG index run, the palette index value and the quantized color for the jump mode are sequentially encoded / parsed. As in HEVC, horizontal or vertical traversal scanning can be applied to scan the samples, such as Figure 29 shown. The coding order for palette run-length coding in each segment is as follows: For each sample position, one context-coded binary bit run_copy_flag=0 is transmitted by signal to indicate whether the pixel has the same mode as the previous sample position, that is, whether the previously scanned sample and the current sample are both of run type COPY_ABOVE, or whether the previously scanned sample and the current sample are both of run type INDEX and have the same index value. Otherwise, run_copy_flag=1 is transmitted by signal. If the current sample and the previous sample have different modes, one context-coded binary bit copy_above_palette_indices_flag is transmitted by signal to indicate the run type of the current sample, that is, INDEX or COPY_ABOVE. Here, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type because INDEX mode is used by default. In the same way, if the previously parsed run type is COPY_ABOVE, the decoder does not have to parse the run type. After the samples are palette run-length encoded in one codec pass, the index values ​​(for INDEX mode) and the quantized jump colors are grouped and encoded in another codec pass using CABAC bypass codec. This separation of the context-encoded bits and the bypass-encoded bits can improve the throughput within each row CG. For slices with dual luma / chroma trees, the palettes are applied separately to luma (Y component) and chroma (Cb component and Cr component), where the luma palette entries contain only Y values ​​and the chroma palette entries contain Cb and Cr values. For slices with a single tree, the palette will be applied jointly to the Y, Cb, and Cr components, i.e., each entry in the palette contains a Y value, a Cb value, and a Cr value, unless the CU is encoded and decoded using a local dual tree, in which case the encoding and decoding of luma and chroma are handled separately. In this case, if the corresponding luma block or chroma block is encoded and decoded using palette mode, their palettes are applied in a similar manner to the dual tree case (this is related to non-4:4:4 codecs and will be further explained in 2.1.3.4.1). For slices coded with dual-tree, the maximum palette predictor size is 63, and the maximum palette table size for the codec of the current CU is 31. For slices coded with dual-tree, the maximum predictor size and the maximum palette table size are halved, i.e., for each of the luma palette and chroma palette, the maximum predictor size is 31, and the maximum table size is 15. For deblocking, palette-coded blocks on both sides of a block boundary are not deblocked. 2.1.3.4.1 Palette Mode for Non-4:4:4 Content In a similar manner to the palette modes in HEVC SCC, the palette modes in VVC support all chroma formats. For non-4:4:4 content, the following customizations are applied: 1. When the pop-out value for a given sample position is signaled, if the sample position has only luma components and no chroma components due to chroma downsampling, only the luma pop-out value is signaled. This is the same as in HEVC SCC. 2. For local dual-tree blocks, the palette mode is applied to the block in the same way as the palette mode applied to single-tree blocks, with two exceptions: a. The process of palette prediction value update is slightly modified as follows. Since the local dual-tree block contains only luma (or chroma) components, the prediction value update process uses the signaled value of the luma (or chroma) component and fills in the "missing" chroma (or luma) component by setting the chroma (or luma) component to the default value (1<<(component bit depth-1)). b. The maximum palette prediction value size remains at 63 (because slices are coded using a single tree), but the maximum palette table size for luma / chroma blocks remains at 15 (because blocks are coded using separate palettes). 3. For palette mode in monochrome format, the number of color components in a palette-encoded block is set to 1 instead of 3. 2.1.3.4.2 Encoder Algorithm for Palette Mode On the encoder side, the following steps are used to generate the palette table for the current CU: 1. First, in order to derive the initial entries in the palette table of the current CU, simplified K-means clustering is applied. The palette table of the current CU is initialized to an empty table. For each sample position in the CU, the SAD between the sample and each palette table entry is calculated, and the minimum SAD among all palette table entries is obtained. If the minimum SAD is less than the predefined error limit errorLimit, the current sample is clustered with the palette table entry with the minimum SAD. Otherwise, a new palette table entry is created. The threshold errorLimit is QP-dependent and is retrieved from a lookup table containing 57 elements covering the entire QP range. After all samples of the current CU are processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entries after the 31st entry are discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroid of each cluster from step 1 or using one of the palette colors in the palette prediction value. The option with the lower rate-distortion cost is selected as the final color of the palette table. If a cluster has only a single sample and the corresponding palette entry is not in the palette prediction value, the corresponding sample is converted to a jump symbol in the next step. 3. The palette table thus generated contains some new entries from the centroids of the clusters in step 1 and some entries from the palette predictions. Therefore, the table is reordered again so that all new entries (i.e. centroids) are placed at the beginning of the table, followed by entries from the palette predictions. Given the palette table of the current CU, the encoder selects the palette index for each sample position in the CU. For each sample position, the encoder checks all index values ​​corresponding to the palette table entries and the RD cost of the index representing the escape symbol, and selects the index with the smallest RD cost using the following equation: RD cost = distortion × (isChroma? 0.8:1) + lambda × bypass codec bits (2-33) After determining the index mapping for the current CU, each entry in the palette table is checked to see if the entry in the palette table is used by at least one sample position in the CU. Any unused palette entry will be removed. After the index mapping for the current CU is determined, grid RD optimization is applied to find the best value of run_copy_flag and run type for each sample position by comparing the RD cost of the following three options: the same as the previously scanned position, run type COPY_ABOVE, or run type INDEX. When calculating the SAD value, the sample value is scaled down to 8 bits unless the CU is coded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. In addition, in the case of lossless codecs, only the rate is used in the above-mentioned rate-distortion optimization step (because lossless codecs do not produce distortion). 2.1.3.5 Adaptive Color Transformation In the HEVC SCC extension, Adaptive Color Transform (ACT) is applied to reduce redundancy between the three color components in the 444 chroma format. ACT has also been adopted into the VVC standard to improve the codec efficiency of the 444 chroma format. As in HEVC SCC, ACT performs an in-loop color space conversion in the prediction residual domain by adaptively converting the residual from the input color space to the YCgCo space. Figure 30 The decoding flow chart for applying ACT is shown. The two color spaces are adaptively selected by signaling 1 ACT flag at the CU level. When the flag is equal to 1, the residual of the CU is encoded and decoded in the YCgCo space; otherwise, the residual of the CU is encoded and decoded in the original color space. Additionally, similar to the HEVC ACT design, for inter-frame CUs and IBC CUs, ACT is enabled only when there is at least one non-zero coefficient in the CU. For intra-frame CUs, ACT is enabled only when the chroma component selects the same intra-frame prediction mode as the luminance component (i.e., DM mode). 2.1.3.5.1ACT Mode In the HEVC SCC extension, ACT supports both lossless and lossy codecs based on the lossless flag (i.e., cu_transquant_bypass_flag). However, no flag is signaled in the bitstream to indicate whether lossy or lossless codec is applied. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown below. Because the YCgCo-R transform is not normalized, to compensate for the dynamic range changes in the residual signal before and after color transformation, QP adjustments (-5, 1, 3) are applied to the transformed residuals of the Y, Cg, and Co components, respectively. The adjusted quantization parameters only affect the quantization and inverse quantization of the residual within the CU. For other codec processes (such as deblocking), the original QP is still applied. Additionally, because forward and inverse color transforms require access to the residuals of all three components, ACT mode is always disabled for single tree partitioning and ISP mode (where the prediction block sizes for different color components are different). When ACT is applied, transform skip (TS) and block differential pulse codec modulation (BDPCM) extended to codec chroma residuals are also enabled. 2.1.3.5.2ACT Fast Encoding Algorithm In order to avoid brutal RD searches in both the original and converted color spaces, the following fast encoding algorithm is applied to the VTM reference software to reduce encoder complexity when ACT is enabled. – The order of enabling / disabling RD checking for ACT depends on the original color space of the input video. For RGB video, the RD cost of ACT mode is checked first; for YCbCr video, the RD cost of non-ACT mode is checked first. The RD cost of the second color space is checked only if there is at least one non-zero coefficient in the first color space. – When a CU is retrieved through a different split path, the same ACT enable / disable decision is reused. Specifically, when the CU is first encoded and decoded, the selected color space used to encode and decode the residual of a CU is stored. Then, when the same CU is retrieved through another split path, the stored color space decision is directly reused instead of checking the RD cost of the two spaces. The RD cost of the parent CU is used to determine whether to check the RD cost of the second color space for the current CU. For example, if the RD cost of the first color space is less than the RD cost of the second color space for the parent CU, the second color space is not checked for the current CU. To reduce the number of codec modes tested, the selected codec mode is shared between the two color spaces. Specifically, for intra mode, pre-selected intra mode candidates based on SATD-based intra mode selection are shared between the two color spaces. For inter mode and IBC mode, block vector search or motion estimation is performed only once. Block vectors and motion vectors are shared between the two color spaces. 2.1.3.6 Intra-frame Template Matching (IntraTMP) Intra Template Matching (IntraTM) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. The encoder searches for the template most similar to the current template in the reconstructed portion of the current frame within a predefined search range and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. By combining the L-shaped causal neighbors of the current block with Figure 31 The prediction signal is generated by matching another block in a predefined search area in the , which consists of: R1: current CTU, R2: Upper left CTU, R3: Upper CTU, R4: left CTU. SAD is used as the cost function. In each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW, SearchRange_h=a*BlkH. Where "a" is a constant that controls the gain / complexity tradeoff. In practice, "a" is equal to 5. The intra template matching tool is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.1.3.6.1 Using Block Vectors Derived from IntraTMP for IBC The block vector (BV) derived from intra template matching prediction (IntraTMP) is used for intra block copy (IBC). The stored IntraTMP BV and IBC BV of neighboring blocks are used as spatial BV candidates in IBC candidate list construction. 2.1.3.6.2 Direct Block Vector (DBV) Mode for Chroma Prediction For chroma components, when the chroma dual tree is activated in an intra slice, if one of the luma blocks (five positions) is coded using MODE_IBC, its block vector bvL is used and scaled to derive the chroma block vector bvC. The scaling factor depends on the chroma format sampling structure. Then, by using the position (xCb, yCb) of the current chroma block and its bvC, the corresponding offset position (xCb+bvC[0], yCb+bvC[1]) is determined, and block copy prediction is performed. As shown in Table 2-7, a CU level flag is signaled to indicate whether the proposed DBV mode is applied. Table 2-7 Binarization process for intra_chroma_pred_mode in the proposed method intra_chroma_pred_mode binary bit string Chroma Intra mode 0 11100 list[0] 1 11101 list[1] 2 11110 list[2] 3 11111 list[3] 4 110 DIMD chromaticity 5 10 DM 6 0 DBV 2.1.4 Transformation and Quantization 2.1.4.1 Large Block Size Transformation with High-Frequency Zeroing In VVC, large block size transforms with a size not exceeding 64×64 are enabled, which are mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are set to zero so that only the low-frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M is equal to 64, only the left 32 columns of transform coefficients are retained. Similarly, when N is equal to 64, only the first 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without setting any value to zero. In addition, transform displacement is removed in transform skip mode. VTM also supports a configurable maximum transform size in SPS, so that the encoder can flexibly select a transform size not exceeding 32 lengths or 64 lengths according to the needs of a specific implementation. 2.1.4.2 Multiple Transformation Selection (MTS) for Kernel Transformations In addition to the DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-8 shows the selected DST / DCT basis functions. Table 2-8 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain orthogonality of the transform matrix, the transform matrix is ​​quantized more accurately than the transform matrix in HEVC. To keep the intermediate values ​​of the transform coefficients within 16 bits, all coefficients have 10 bits after horizontal transform and after vertical transform. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is applied only to luma. MTS signaling is skipped when one of the following conditions applies: - The position of the last significant coefficient of the luma TB is less than 1 (ie, DC only). – The last significant coefficient of the luminance TB is located within the MTS zero region. If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to one, two other flags are additionally signaled to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 2-9. A unified transform selection for ISP and implicit MTS is used by removing the intra mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are turned on, only DST7 is used for both the horizontal transform kernel and the vertical transform kernel. When it comes to transform matrix precision, an 8-bit main transform kernel is used. Therefore, all transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform kernels including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8 use an 8-bit main transform kernel. Table 2-9 - Transformation and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, high-frequency transform coefficients are set to zero for DST-7 blocks and DCT-8 blocks with size (width or height, or both) equal to 32. Only coefficients in the 16x16 low-frequency region are retained. As in HEVC, the residual of a block can be coded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. When MTS is enabled for an inter-coded block, implicit MTS can also be enabled. 2.1.4.3 Low-Frequency Non-Separable Transform (LFNST) In VVC, LFNST is applied between the forward main transform and quantization (at the encoder) and between dequantization and inverse main transform (at the decoder side). In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied depending on the block size. For example, 4×4 LFNST is applied to small blocks (i.e., min(width, height) <8), and 8×8 LFNST is applied to larger blocks (i.e., min(width, height)>4). Using the input as an example, the description of the application of the non-separable transform used in LFNST is as follows. To apply 4x4 LFNST, the 4x4 input block X First, it is represented as a vector The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. The 16x1 coefficient vector Subsequently, it is reorganized into 4x4 blocks using the scan order (horizontal, vertical, or diagonal) for this block. Coefficients with smaller indices will be placed at positions with smaller scan indices in the 4x4 coefficient block. 2.1.4.3.1 Simplified non-separable transform The LFNST (Low Frequency Non-Separable Transform) applies the non-separable transform based on a direct matrix multiplication approach, enabling it to be implemented in a single pass without multiple iterations. However, the dimension of the non-separable transform matrix needs to be reduced to minimize the computational complexity and the memory space for storing the transform coefficients. Therefore, the Simplified Non-Separable Transform (or RST) method is used in LFNST. The main idea of the Simplified Non-Separable Transform is to map an N-dimensional vector (for 8x8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix is not an NxN matrix but becomes an R×N matrix as follows: Where the R rows of the transform are the R basis of the N-dimensional space. The inverse transform matrix of RT is the transpose of the forward transform of RT. For 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix (which is the size of a conventional 8x8 non-separable transform matrix) is reduced to a 16x48 direct matrix. Therefore, a 48×16 inverse RST matrix is ​​used on the decoder side to generate the core (main) transform coefficients in the 8×8 upper left region. When a 16x48 matrix is ​​applied instead of a 16x64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4x4 blocks in the upper left 8x8 block in addition to the lower right 4x4 block. With the reduced dimension, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, and the performance degradation is reasonable. To reduce complexity, LFNST is restricted to being applicable only when all coefficients outside the first coefficient subgroup are insignificant. Therefore, when LFNST is applied, all main transform coefficients must be zero. This allows LFNST index signaling to be conditional on the last significant position, thus avoiding the additional coefficient scan required in current LFNST designs to check only for significant coefficients at specific positions. The worst-case processing of LFNST (in terms of per-pixel multiplications) restricts the non-separable transforms for 4x4 blocks and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In these cases, when LFNST is applied, the last significant scan position must be less than 8, and less than 16 for other sizes. For blocks of shape 4xN and Nx4 with N > 8, the proposed restriction means that LFNST is now applied only once, and only to the top-left 4x4 region. Since all primary coefficients are zero when LFNST is applied, the number of operations required for the primary transform is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when the LFNST transform is tested. Rate-distortion-optimized quantization must be performed for at most the first 16 coefficients (in scan order), with the remaining coefficients forced to zero. 2.1.4.3.2 LFNST Transform Selection In LFNST, a total of 4 transform sets are used, and 2 inseparable transform matrices (kernels) are used for each transform set. As shown in Table 2-10, the mapping from intra prediction mode to transform set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an LFNST index explicitly transmitted by signal. The index is transmitted once in the bitstream for each intra CU after the transform coefficient. Table 2-10 - Transformation selection table IntraPredMode Transform Set Index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0 2.1.4.3.3 LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to being applicable only when all coefficients outside the first coefficient subgroup are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-encoded, but does not depend on the intra prediction mode, and only the first binary bit is context-encoded. In addition, LFNST is applied to intra CUs in both intra and inter slices, and to both luma and chroma. If dual-tree is enabled, the LFNST indices for luma and chroma are signaled separately. For inter slices (dual-tree disabled), a single LFNST index is signaled and used for both luma and chroma. Considering that large CUs larger than 64x64 are implicitly partitioned (TU slicing) due to the existing maximum transform size limit (64x64), LFNST index search can quadruple the data cache for a certain number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST is limited to 64x64. Note that LFNST is only enabled with DCT2. LFNST index signaling is placed before MTS index signaling. The use of the scaling matrix for perceptual quantization is not obvious, and the scaling matrix specified for the primary matrix can be used for the LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied. 2.1.4.4 Sub-block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only a sub-portion of the residual block is encoded and decoded for the CU. When an inter-predicted CU has cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-portion of the residual block is encoded and decoded. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, part of the residual block is encoded and decoded using the inferred adaptive transform, and the other part of the residual block is cleared to zero. When SBT is used for inter-coded CU, SBT type and SBT location information are signaled in the bitstream. Figure 35As shown, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is like a binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is like an asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 in luminance samples, 1:3 / 3:1 partitioning along that dimension is not allowed. There are up to 8 SBT modes for a CU. Position-dependent transform kernel selection is applied to the luma transform block in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are selected in Figure 35 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DCT-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slice, cbf, and horizontal and vertical kernel transform types for the residual block. SBT is not applied to CUs coded using the joint intra-inter mode. 2.1.4.5 Maximum Transform Size and Transform Coefficient Clearing The CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the largest intra-frame codec block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, no normal clearing operation is applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are normalized to zero. 2.1.4.6 Enhanced MTS for Intra-frame Coding and Decoding In the current VVC design, for MTS, only the DST7 transform core and the DCT8 transform core are utilized, which are used for intra-frame coding and decoding and inter-frame coding and decoding. Additional main transforms including DCT5, DST4, DST1 and identity transform (IDT) are used. In addition, the MTS set also depends on the TU size and intra-frame mode information. 16 different TU sizes are considered, and for each TU size, 5 different categories are considered based on the intra-frame mode information. For each category, 4 different transform pairs are considered, the same as for VVC. Note that although a total of 80 different categories are considered, some of these different categories often share exactly the same transform set. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For angle modes, the joint symmetry of TU shape and intra prediction is taken into account. Therefore, mode i (i>34) with TU shape AxB will be mapped to the same category corresponding to mode j=(68-i) with TU shape BxA. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For wide-angle mode, the nearest conventional angle mode is used for determination of the transform set. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. The MTS index [0,3] is signaled using a 2-bit fixed length codec. 2.1.4.7 Secondary Transformation: Extension of LFNST with Large Kernel The LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is extended to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: ○ For predModeIntra<2, lfnstTrSetIdx is equal to 2. ○For predModeIntra in [0,34], lfnstTrSetIdx=predModeIntra. ○ For predModeIntra in [35,66], lfnstTrSetIdx=68−predModeIntra. • Three different kernels LFNST4, LFNST8, and LFNST16 are defined to indicate the LFNST kernel set, which are applied to 4xN / Nx4 (N≥4), 8xN / Nx8 (N≥8), and MxN (M, N≥16), respectively. The kernel dimensions are specified by: (LFSNT4,LFNST8*,LFNST16*)=(16x16,32x64,32x96). Forward LFNST is applied to the low-frequency region on the upper left, called the region of interest (ROI). When LFNST is applied, the main transform coefficients present in the region other than the ROI are cleared to zero, which is unchanged compared to the VVC standard. The ROI of LFNST16 is as follows Figure 36As shown in Figure 1. This ROI consists of six 4x4 sub-blocks, consecutive in scan order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be Rx96. In this contribution, R is chosen to be 32, and accordingly, 32 coefficients (two 4x4 sub-blocks) are generated from the forward LFNST16, which are placed to follow the coefficient scan order. The ROI of LFNST8 is as follows Figure 37 The forward LFNST8 matrix can be Rx64, and R is chosen to be 32. The generated coefficients are positioned in the same way as LFNST16. The mapping of intra prediction modes to these sets is shown in Table 2-11. Table 2-11. Mapping of intra prediction modes and LFNST set indices 2.1.4.8 Non-separable primary transform (NSPT) for intra-frame coding and decoding For block sizes 4x4, 4x8, 8x4, and 8x8, DCT-II+LFNST is replaced by NSPT. NSPT follows the design of LFNST, i.e., 3 candidates and 35 sets based on intra-mode selection. The kernel sizes are as follows: NSPT4x4: 16x16, NSPT4x8 / NSPT8x4: 32x20, NSPT8x8: 64x32. Therefore, 12 coefficients and 32 coefficients are cleared for NSPT4x8 / NSPT8x4 and NSPT8x8, respectively. 2.1.4.9 Symbol Prediction The basic idea of ​​the coefficient sign prediction methods (JVET-D0031 and JVET-J0021) is to compute the reconstructed residuals for both negative and positive sign combinations for applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal symbols, the cost function is defined as a measure of discontinuity across block boundaries, as Figure 38 The cost function is measured for all hypotheses, and the hypothesis with the smallest cost is chosen as the predicted value of the coefficient sign. The cost function is defined as the sum of the absolute second-order derivatives in the residual domain of the upper row and left column as follows: Where R is the reconstructed neighboring block, P is the prediction of the current block, and r is the residual hypothesis. -1 +2R0-P1) can be calculated only once for each block, and only the residual hypothesis is subtracted. 2.2 About Block-Level Adaptive OBMC-Neighboring Prediction Mode in Video Coding The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In this section, regarding "blocks encoded and decoded using mode N", the "mode N" here can be a prediction mode (for example, MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding and decoding technology (for example, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.). It should be noted that the following terms are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 2.2.1 In one example, whether OBMC is applied to the current block may depend on the prediction modes of spatial / temporal neighboring blocks that are adjacent / non-adjacent to the current block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on whether there are neighboring blocks that utilize IBC codec. 4) For example, it can be based on whether there are neighboring blocks that utilize PLT codec. 5) For example, it can be based on whether there are neighboring blocks that utilize intraTMP codec. 6) For example, it can be based on whether there are neighboring blocks that are coded with BDPCM. 7) For example, it can be based on whether there are neighboring blocks that utilize transform skip coding. 8) For example, the neighboring block may be adjacent to the current block. 9) For example, the neighboring block may not be adjacent to the current block. 10) For example, the neighboring block may be a spatial neighboring block within the current picture. 11) For example, the neighboring block may be a time domain block in a reference picture. 12) For example, the neighboring block may be a sub-block (eg, 4x4 or 8x8) smaller than the current block. 13) For example, the neighboring block may be a video unit that is larger than or equal to the current block. 14) For example, the neighboring blocks may be sample locations. 15) For example, a series of adjacent neighboring blocks / sub-blocks to the left and / or above the current block may be checked one by one (eg, following a predefined position and a predefined checking order). a. For example, if there is a neighbor that utilizes a specific mode codec, the process is terminated and OBMC is considered not to be applied to the current block. b. For example, if there is a neighbor that is coded using inter-frame mode, it is further checked whether the reference block of the neighbor is coded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded neighbor and the position of such inter-frame coded neighbor), and if the reference block is coded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in a reference picture. c. For example, if there is a neighbor that is coded using intraTMP, it can be further checked whether the reference block of the neighbor is coded using a specific mode (for example, the reference block is identified by adding the block vector associated with the neighbor that is coded using such intraTMP and the position of the neighbor that is coded using such intraTMP), and if the reference block is coded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in the current picture. d. For example, the specific mode may be IBC and / or PLT. e. For example, the specific mode may be intraTMP. f. For example, the specific mode may be BDPCM. g. For example, the specific mode may be transform skip. 16) For example, a series of non-adjacent neighboring blocks / sub-blocks in an already coded area of ​​the current picture may be checked one by one (eg, following a predefined position and a predefined checking order). h. For example, if there is a neighbor that utilizes a specific mode codec, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, if there is a neighbor that is coded using inter-frame mode, it is further checked whether the reference block of the neighbor is coded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded neighbor and the position of such inter-frame coded neighbor), and if the reference block is coded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in a reference picture. j. For example, if there is a neighbor that is coded using intraTMP, it is further checked whether the reference block of the neighbor is coded using a specific mode (for example, the reference block is identified by adding the block vector associated with the neighbor that is coded using such intraTMP and the position of the neighbor that is coded using such intraTMP), and if the reference block is coded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in the current picture. k. For example, the specific mode may be IBC and / or PLT. 1. For example, the specific mode may be intraTMP. m. For example, the specific mode may be BDPCM. n. For example, the specific mode may be transform skip. 17) For example, a series of temporal blocks / sub-blocks in a reference picture may be checked one by one (eg, following a predefined position and order). o. For example, if there is a time domain block that is coded using a specific mode, the process is terminated and OBMC is considered not to be applied to the current block. p. For example, if there is a time domain block coded using inter-frame mode, it is further checked whether the reference block is coded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded time domain block and the position of such inter-frame coded time domain block), and if the reference block is coded using the specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. q. For example, if there is a time domain block encoded using intraTMP, it is further checked whether the reference block of the time domain block is encoded using a specific mode (for example, the reference block is identified by adding a block vector associated with such intraTMP encoded time domain block and the position of such intraTMP encoded time domain block), and if the reference block is encoded using the specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. r. For example, the specific mode may be IBC and / or PLT. s. For example, a specific mode may be intraTMP. t. For example, the specific mode may be BDPCM. u. For example, the specific mode may be transform skip. 2.2.2 In one example, whether OBMC is applied to the current block may depend on the prediction mode of the reference block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, the reference block may be a block / sub-block identified based on adding a displacement (eg, predefined, or based on a motion vector, or based on a block vector) to the position of the first block. a. For example, the first block may be the current block. b. For example, the first block may be a neighboring block adjacent to the current block. c. For example, the first block may be a neighboring block that is not adjacent to the current block. d. For example, the first block may be a reference block for the current block. e. For example, the first block may be a reference block of a neighboring block. f. For example, the reference block may be identified based on the position of the current block in inter-mode coding and its motion information associated with the current block (eg, motion vector and reference index). i. For example, in this case, the reference block is in a reference picture. g. For example, the reference block may be identified based on the position of the current block encoded in the IntraTMP mode and its motion information (eg, block vector) associated with the current block. i. For example, in this case, the reference block is in the current picture. h. For example, a reference block may be identified based on the location of an inter-mode coded neighboring block and its motion information (eg, motion vector and reference index) associated with the inter-mode coded neighbor. i. For example, in this case, the reference block is in a reference picture. i. For example, a reference block may be identified based on the location of an IntraTMP mode coded neighboring block and its motion information (eg, block vector) associated with the IntraTMP mode coded neighbor. i. For example, in this case, the reference block is in the current picture. j. For example, the reference block may be identified based on the position of the inter-mode coded reference block and its motion information (eg, motion vector and reference index) associated with the inter-mode coded reference block. i. For example, in this case, the reference block is in another reference picture (not the reference picture where the reference block of inter-mode coding is located). k. For example, the reference block may be identified based on the position of the reference block coded in the IntraTMP mode and its motion information (eg, block vector) associated with the reference block coded in the IntraTMP mode. i. For example, in this case, the reference block is in the same reference picture as the reference block encoded in the IntraTMP mode. 4) For example, in the case where a neighboring block is coded using the inter mode, the reference block may be checked. a. If the neighboring block is inter-coded, the reference block is then identified by adding the motion vector associated with such inter-coded neighbor and the position of such inter-coded neighbor, and if the reference block is coded using a specific mode, OBMC is considered not to be applied to the current block. b. For example, the specific mode may be IBC and / or PLT. c. For example, the specific mode may be intraTMP. d. For example, the specific mode may be BDPCM. e. For example, the specific mode may be transform skip. 5) For example, in the case where the neighboring blocks are coded using the IntraTMP mode, the reference blocks can be checked. a. If the neighboring block is intraTMP coded, the reference block is then identified by adding the block vector associated with such intraTMP coded neighbor and the position of such intraTMP coded neighbor, and if the reference block is coded using a specific mode, OBMC is considered not to be applied to the current block. b. For example, the specific mode may be IBC and / or PLT. c. For example, the specific mode may be BDPCM. d. For example, the specific mode may be transform skip. 2.2.3 In one example, whether OBMC is applied to the current block may depend on the prediction mode of the reference block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, a reference block of a reference block may be identified by adding a motion vector associated with an inter-mode coded reference block and the position of such an inter-coded reference block. 4) For example, a reference block of a reference block may be identified by adding a block vector associated with an IntraTMP mode coded reference block and the position of such an IntraTMP coded reference block. 5) For example, historical / propagated prediction patterns can be stored in a cache. a. For example, if the block itself is coded using a specific mode or the block once had a reference block coded using a specific mode, the specific mode may be stored in a cache associated with the block information, indicating that it has information about the history / propagation of the specific mode. 6) For example, the specific mode may be IBC and / or PLT. 7) For example, the specific mode may be intraTMP. 8) For example, the specific mode may be BDPCM. 9) For example, the specific mode may be transform skip. 2.2.4 In one example, whether a block in the current picture is coded into a specific mode may be stored in a buffer. 1) For example, the specific mode may be IBC. 2) For example, the specific mode may be PLT. 3) For example, the specific mode may be intraTMP. 4) For example, the specific mode may be BDPCM. 5) For example, the specific mode may be transform skip. 6) For example, whether a block is encoded or decoded using IBC or PLT can be stored using shared parameters / buffers. a. For example, a single parameter / buffer can be used for storage. 7) For example, whether a block is coded using IBC or PLT or intraTMP or BDPCM may be stored using a separate parameter / buffer. b. For example, multiple parameters / buffers can be used for storage. 8) For example, it can be stored at a granularity of MxM (eg, M=4 or M=8) sub-blocks. 2.2.5 In one example, whether a block and / or its reference blocks are coded into a specific mode may be stored in a cache. 1) For example, the specific mode may be IBC. 2) For example, the specific mode may be PLT. 3) For example, the specific mode may be intraTMP. 4) For example, the specific mode may be BDPCM. 5) For example, the specific mode may be transform skip. 6) For example, the block or its reference blocks are encoded using IBC or PLT, a parameter equal to true (eg, indicating that it is a historical / propagated screen content block) may be stored in the cache. a. Alternatively, on the other hand, a parameter equal to false (eg, indicating that it is not a historical / propagated screen content block) may be stored in the cache. 7) For example, whether a block is coded using IBC or PLT, and whether the reference blocks of such block are coded using IBC or PLT, may be stored as separate parameters and stored in separate buffers. 8) For example, it can be stored at a granularity of MxM (eg, M=4 or M=8) sub-blocks. 2.2.6 In one example, whether OBMC is enabled may be coupled to whether a particular tool is enabled. 1) For example, the specific tool may be an IBC. 2) For example, the specific tool may be a PLT. 3) For example, the specific tool may be intraTMP. 4) For example, the specific tool may be BDPCM. 5) For example, a specific tool may be a transform skip. 6) In one example, whether a specific tool is applied may be controlled by a first syntax element (SE), such as in a VPS / SPS / PPS / slice header / CTU / CU / etc. 7) In one example, whether to apply OBMC may be controlled by a second syntax element (SE), such as in VPS / SPS / PPS / slice header / CTU / CU / etc. 8) In one example, it may be constrained that if a first SE indicates that a particular tool is enabled, then a second SE must indicate that OBMC is disabled. 9) In one example, it may be set at the encoder that if the first SE indicates that a particular tool is enabled, the second SE must indicate that OBMC is disabled. 2.2.7 Whether to apply and / or how to apply the method disclosed above can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 2.2.8 Whether and / or how to apply the methods disclosed above may be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. 2.2.9 Whether and / or how to apply the method disclosed above may depend on the information being coded, such as block size, color format, single / dual tree partitioning, color component, slice type / picture type. 2.3 Block-level Adaptive OBMC-Sample Value Correlation in Video Coding The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In this section, regarding "blocks encoded and decoded using mode N", the "mode N" here can be a prediction mode (for example, MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding and decoding technology (for example, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.). It should be noted that the following terms are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 2.3.1 In one example, whether OBMC is applied to the current block may depend on sample values ​​of samples within the current block (and / or adjacent to the current block). 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on the prediction samples of the current block (before OBMC). 4) For example, it can be based on prediction samples that are adjacent to the current block. 5) For example, it can be based on the reconstructed samples that are adjacent to the current block. 6) For example, it can be based on the gradient / direction / angle (or gradient histogram / direction histogram / angle histogram) of samples inside the current block (and / or adjacent to the current block). a. For example, the current prediction sample point before OBMC can be used. b. For example, adjacent reconstruction points can be used. c. For example, a gradient histogram / direction histogram / angle histogram may be calculated based on counting gradients along a specific direction / angle. i. For example, a specific direction / angle may be predefined. ii. For example, the specific direction / angle may be based on the direction of an intra prediction angle mode in a video codec. iii. For example, for a specific direction / angle, the gradient magnitude may be calculated based on counting the gradient (magnitude of the gradient) of at least one sample point in the current block. 1. For example, the gradient (magnitude of the gradient) at a specific position (eg, the center) in the current block may be counted. 2. For example, the gradients (magnitudes of the gradients) at a series of specific positions in the current block may be counted. 3. For example, the gradients (magnitudes of gradients) of all samples in the current block may be counted. 4. For example, the gradients (magnitudes of gradients) of all samples in the current block except the first row, last row, first column, and last column of samples may be counted. iv. For example, for a specific direction / angle, the gradient magnitude may be calculated based on counting the gradient (magnitude of the gradient) of at least one sample point adjacent to the current block. 1. For example, gradients (magnitudes of gradients) at specific locations adjacent to the current block may be counted. 2. For example, the gradients (magnitudes of gradients) of all samples adjacent to (left and / or top of) the current block may be counted. d. For example, a gradient histogram / direction histogram / angle histogram may be calculated based on dividing the entire range of directions / angles into a series of intervals / bins. e. For example, a gradient histogram / direction histogram / angle histogram may be calculated based on counting the gradient (magnitude of the gradient) in each bin / bin / direction / angle. 7) For example, it can be based on the color / brightness / intensity (or color histogram / brightness histogram / intensity histogram) of samples inside the current block (and / or adjacent to the current block). a. For example, the current prediction sample point before OBMC can be used. b. For example, adjacent reconstruction points can be used. c. For example, the color histogram / luminance histogram / intensity histogram can be calculated based on counting sample values ​​in the Y component domain and / or U component domain and / or V component domain (or, R component domain and / or G component domain and / or B component domain). i. For example, the sample values ​​at a series of specific positions in the current block can be counted. ii. For example, the sample values ​​of all samples in the current block may be counted. iii. For example, sample values ​​at specific locations adjacent to the current block may be counted. iv. For example, the sample values ​​of all samples adjacent to (to the left and / or top of) the current block may be counted. d. For example, a color histogram / brightness histogram / intensity histogram may be calculated based on dividing the entire range of color values / brightness values / intensity values ​​into a series of intervals / bins. e. For example, the color histogram / brightness histogram / intensity histogram can be calculated based on counting the number of samples in each interval / bin. 8) For example, it can be based on the number of main gradients / number of main directions / number of main angles / number of main colors / number of main brightness / number of main intensities of samples inside the current block (and / or adjacent to the current block). a. For example, it can be calculated based on the prediction samples inside the current block (before OBMC). b. For example, it can be calculated based on the reconstructed samples adjacent to the current block. c. For example, the main gradient / main direction / main angle / main color / main brightness / main intensity can be derived based on the gradient histogram / direction histogram / angle histogram / color histogram / brightness histogram / intensity histogram. d. For example, the main gradient / main direction / main angle / main color / main brightness / main intensity can be derived based on how many intervals / bins in the histogram show values ​​greater than a threshold (e.g., gradient magnitude, color value, brightness value). e. For example, the main gradient / main direction / main angle / main color / main brightness / main intensity can be derived based on how many intervals / bins in the histogram provide values ​​(e.g., gradient magnitude, color value, brightness value) that are much larger than the values ​​of other intervals / bins. i. For example, the values ​​of the intervals / bins in the histogram can be sorted first - assuming that the values ​​after sorting (e.g., from large to small) are X0, X1, X2, ..., X n-2 、X n-1 Indicates which n intervals / bins are included in the histogram—if X i > = a*(X i+1 ), then the interval / bin from 0 to i can be regarded as the main gradient, where a represents the scaling factor (for example, a can be equal to a constant between 2 and 20). ii. For example, if the number of main gradients / the number of main directions / the number of main angles / the number of main colors / the number of main brightness / the number of main intensities is less than a specific number (e.g., 1, or 2, or 3, or 4, or 5, or 6, or 7, or 8, or 9), OBMC may not be applied to the block. 2.3.2 In one example, whether OBMC is applied to the current block may depend on the template cost. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on a first non-hybrid template cost and a second hybrid template cost. a. For example, the first template cost may be calculated based on the SAD between the current template and a reference template (where the reference template is identified by adding the current motion vector to the position of the current template). b. For example, the second template cost may be calculated based on the SAD between the current template and a hybrid reference template (where the hybrid reference template may be generated by blending the template identified by the current motion vector and templates identified by neighboring motion vectors). c. For example, if the non-hybrid template cost is smaller, OBMC may not be applied to the block. i. Alternatively, OBMC can be applied to the block if the hybrid template is less expensive. 2.3.3 In one example, whether OBMC is applied to the current block may depend on the motion vector accuracy of the current block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on whether the motion vector of the block is an integer (rather than fractional) precision motion vector. 4) For example, it may be based on whether the motion vector differences of the blocks are integer (rather than fractional) precision motion vector differences. 2.3.4 In one example, when the encoding and decoding information of the second block is used in the encoding and decoding process of the current block, the position of the second block may be restricted based on a specific rule. 1) For example, the encoding and decoding process of the current block may refer to at least one of the following: a. Mode decision. b. Motion candidate derivation. c. Movement list generation. d. Block vector candidate derivation. e. Block vector list generation. f. Intra-mode candidate derivation. g. Generation of intra-frame luminance MPM list. h. Intra-frame chroma block vector candidate derivation under dual tree. i. Intra-frame chroma mode candidate derivation under dual-tree. j. Block-level OBMC on / off decision. 2) For example, the second block may be a reference block of the current block. 3) For example, the second block may be a reference block of a reference block of the current block. 4) For example, the second block may be a co-located luminance block or a non-co-located luminance block of the current chrominance block. 5) For example, the second block may be required not to exceed the valid range. a. For example, the effective search range may be predefined. b. For example, the effective search range can be based on CTU size / information. c. For example, the effective search range may be based on the VPDU size / information. d. For example, the effective search range can be based on the tile size / information. e. For example, the effective search range can be based on sub-picture size / information. 6) For example, when the second block represents a reference block derived by a motion vector (e.g., from inter mode), the requirement for the position of the reference block may be based on the position of the CTU / CTU row / slice / sub-picture where the current block is located. a. For example, the reference block may be required to be located no further than the co-located CTU (i.e., the CTU in the reference picture and co-located with the current CTU) and X (e.g., X=3) sample columns to the right of the co-located CTU. b. For example, the reference block may be required to be located no further than the co-located CTU and the CTU to the right adjacent to the co-located CTU. c. For example, it may be required that the location of the reference block does not exceed the same CTU row. d. For example, the reference block may be required to be located no more than Y (eg, Y=3) sample rows above the co-located CTU. e. For example, it may be required that the location of the reference block does not exceed the location of the same sub-picture. f. For example, it may be required that the location of the reference block does not exceed that of the co-located slice. 7) For example, when the second block represents a reference block B of a reference block A of a current block, the requirement for the position of the reference block B may be based on the position of the CTU / CTU row / slice / sub-picture where the current block is located. a. For example, the reference block B may be required to be located no more than X (eg, X=3) sample columns to the right of the co-located CTU and the adjacent co-located CTU. b. For example, the reference block may be required to be located no further than the co-located CTU and the CTU to the right adjacent to the co-located CTU. c. For example, it may be required that the location of the reference block does not exceed the same CTU row. d. For example, the reference block may be required to be located no more than Y (eg, Y=3) sample rows above the co-located CTU. e. For example, the reference block may be required to be located no further than the co-located CTU and the CTU to the left adjacent to the co-located CTU. f. For example, the reference block may be required to be located no further than the co-located CTU, the CTU adjacent to the left of the co-located CTU, and the CTU adjacent to the right of the co-located CTU. g. For example, it may be required that the location of the reference block does not exceed that of the co-located sub-picture. h. For example, it may be required that the location of the reference block does not exceed the co-located slice. 8) For example, when the second block is the luminance block of the current chrominance block, it may be required that the position of the luminance block does not exceed the co-located luminance CU. a. Alternatively, the position of the luma block may be required not to exceed the current luma CTU. b. Alternatively, the position of the luma block may be required to be no more than the current luma CTU and a CTU to the right adjacent to the current luma CTU. c. Alternatively, the position of the luma block may be required not to exceed the current luma CTU row. 9) For example, when the second block exceeds the valid range (or, is outside the required position range), the second block may be deemed unavailable. a. For example, in this case, predefined codec information may be used instead. b. For example, in this case, the codec information of the second block is not used. 2.3.5 For example, how many prediction samples and / or which prediction samples are used for mode decision may be based on codec information and / or predefined rules. 1) For example, the mode decision may refer to at least one of the following: a. Determination of OBMC on / off based on gradient / DIMD. b. Determination of the transformation kernel based on gradient / DIMD. c. Gradient / DIMD based intra mode derivation for main transform. d. Gradient / DIMD based intra mode derivation for secondary transform. e. Gradient / DIMD based intra mode derivation for separable transform. f. Gradient / DIMD based intra mode derivation for inseparable transforms. g. LFNST kernel derivation for a specific mode (e.g., MIP mode). h. NSPT kernel derivation for a specific mode (eg, MIP mode). 2) For example, not all prediction samples in the current block are used for mode decision. 3) For example, some prediction samples within the current block are used for mode decision. 4) For example, the prediction samples used for mode decision can be downsampled. 5) For example, the prediction samples within the current block may be downsampled by a downsampling factor. a. For example, the downsampling factor in the width direction may be equal to 1, or 2, or 4, or 8. b. For example, the downsampling factor in the height direction may be equal to 1, or 2, or 4, or 8. c. For example, the values ​​of the downsampling factors in the width direction and / or the height direction may be derived based on the block width and / or the block height. i. For example, larger downsampling factors can be used for larger blocks. ii. For example, a smaller downsampling factor can be used for smaller blocks. 6) For example, whether the prediction samples are downsampled may be determined based on block information. a. For example, it can be based on the number of samples in the block. b. For example, it can be based on block width. c. For example, it can be based on block height. d. For example, the downsampling method may be based on at least one threshold. 7) For example, the first row and / or last row and / or first column and / or last column of samples in the prediction block may not be used for mode decision. a. For example, the prediction block can be downsampled. b. For example, the prediction block may not be downsampled. 8) For example, partial samples / downsampled samples can be used for mode decision. 9) For example, gradients and / or gradient histograms / color histograms / luminance histograms / intensity histograms may be calculated based on partial samples / downsampled samples. 2.3.6 Whether to apply and / or how to apply the method disclosed above can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 2.3.7 Whether and / or how to apply the methods disclosed above may be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. 2.3.8 Whether and / or how to apply the method disclosed above may depend on the information being coded, such as block size, color format, single / dual tree partitioning, color component, slice type / picture type. 3. Question In ECM-7.0, codec tools are applied to both natural and screen content tools. For some codec tools (such as IBC, PLT, etc.), the on / off of the tool can be controlled by sequence level syntax elements, however, there is no adaptive block level video content detection. 4. Detailed solution The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In this section, regarding "blocks encoded and decoded using mode N", the "mode N" here can be a prediction mode (for example, MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding and decoding technology (for example, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.). It should be noted that the following terms are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 4.1 For example, the content type of a video unit can be determined based on sample values ​​within or adjacent to the video unit. 1) For example, a video unit may be a block / sub-block / CU / PU / TU / slice / slice / sub-picture. 2) For example, it can be based on the predicted samples of the current video unit. a. For example, the prediction samples can be before OBMC mixing / fusion / weighting. b. For example, the prediction samples can be before MHP mixing / fusion / weighting. c. For example, prediction samples can be prior to BCW mixing / fusion / weighting. d. For example, prediction samples can be prior to CIIP mixing / fusion / weighting. e. For example, prediction samples can be before GPM / SGPM mixing / fusion / weighting. f. For example, prediction samples can be preceded by bidirectional mixing / fusion / weighting. g. For example, the prediction samples may precede the prediction sample refinement process (such as BDOF, PROF, LIC, OBMC, etc.). h. For example, the prediction samples may be preceded by a sample filtering process (such as PDPC, CIIP-PDPC, gradient-PDPC, reference sample filtering / smoothing, prediction sample filtering / smoothing, etc.). 3) For example, it can be based on the prediction samples that are adjacent to the current video unit. 4) For example, it can be based on reconstructed samples adjacent to the current video unit. a. For example, the reconstructed samples may be preceded by a sample filtering process (such as bilateral filtering, deblocking, neural network-based filtering, LMCS, SAO, CCSAO, ALF, CCALF, motion-compensated temporal filtering, etc.). 5) For example, the samples used for determination may be derived based on a downsampling process. a. For example, the downsampling factor can be determined based on the width / height of the current block. b. For example, the downsampling factor can be determined based on the width / height of neighboring blocks. c. For example, the downsampling factor may be predefined as a fixed value. 6) For example, it can be based on the gradient / direction / angle (or gradient histogram / direction histogram / angle histogram) of a particular sample point. a. For example, a gradient histogram / direction histogram / angle histogram can be calculated based on counting gradients along a specific direction / angle. i. For example, a specific direction / angle may be predefined. ii. For example, the specific direction / angle may be based on the direction of an intra prediction angle mode in a video codec. iii. For example, for a specific direction / angle, the gradient magnitude may be calculated based on counting the gradient (magnitude of the gradient) of at least one sample point in a designated area. b. For example, the gradient histogram / direction histogram / angle histogram can be calculated based on dividing the entire range of directions / angles into a series of intervals / bins. c. For example, a gradient histogram / direction histogram / angle histogram may be calculated based on counting the gradient (magnitude of the gradient) in each bin / bin / direction / angle. 7) For example, it can be based on the color / brightness / intensity (or color histogram / brightness histogram / intensity histogram) of a particular sample. a. For example, the color histogram / luminance histogram / intensity histogram can be calculated based on counting sample values ​​in the Y component domain and / or U component domain and / or V component domain (or, R component domain and / or G component domain and / or B component domain). i. For example, sample values ​​at a series of specific positions in the current video unit can be counted. ii. For example, the sample values ​​of all samples in the current video unit may be counted. iii. For example, sample values ​​at specific locations adjacent to the current video unit may be counted. iv. For example, the sample values ​​of all samples adjacent to (to the left and / or top of) the current video unit may be counted. b. For example, a color histogram / brightness histogram / intensity histogram may be calculated based on dividing the entire range of color values / brightness values / intensity values ​​into a series of intervals / bins. c. For example, the color histogram / brightness histogram / intensity histogram can be calculated based on counting the number of samples in each interval / bin. 8) For example, it can be based on the number of main gradients / number of main directions / number of main angles / number of main colors / number of main luminances / number of main s of samples inside (and / or adjacent to) the current video unit. a. For example, it can be derived based on gradient histogram / direction histogram / angle histogram / color histogram / brightness histogram / intensity histogram. b. For example, it can be derived based on how many bins / bins in the histogram (e.g., gradient magnitude, color value, brightness value) show values ​​larger than a threshold. c. For example, it can be derived based on how many bins / bins in the histogram provide values ​​(e.g., gradient magnitude, color value, brightness value) that are much larger than the values ​​of other bins / bins. i. For example, the values ​​of the intervals / bins in the histogram can be sorted first - assuming that the values ​​after sorting (e.g., from large to small) are X0, X1, X2, ..., X n-2 、X n-1Indicates which n intervals / bins are included in the histogram—if X i > = a*(X i+1 ), then the interval / bin from 0 to i can be regarded as the main gradient, where a represents the scaling factor (for example, a can be equal to a constant between 2 and 20). ii. For example, if the number of main gradients / the number of main directions / the number of main angles / the number of main colors / the number of main brightness / the number of main intensities is less than a specific number (e.g., 1, or 2, or 3, or 4, or 5, or 6, or 7, or 8, or 9), a specific codec tool may not be applied to the video unit. 4.2 For example, the content type of a video unit may be determined based on the prediction mode of neighboring video units. 1) For example, the neighboring block may be a spatial / temporal neighboring block that is adjacent / non-adjacent to the current block. a. For example, the neighboring block may be adjacent to the current block. b. For example, the neighboring block may not be adjacent to the current block. c. For example, the neighboring block may be a spatially neighboring block within the current picture. d. For example, the neighboring block can be a time domain block in a reference picture. e. For example, the neighboring block may be a sub-block (eg, 4x4 or 8x8) smaller than the current block. f. For example, the neighboring block may be a video unit that is larger than or equal to the current block. 2) For example, the neighboring block may be a reference block. a. For example, the reference block may be a block / sub-block identified based on adding a displacement (eg, predefined, or based on a motion vector, or based on a block vector) to the position of the first block. a. For example, the first block may be the current block. b. For example, the first block may be a neighbor of the current block. c. For example, the first block may be a reference block for the current block. 3) For example, the neighbor may be a reference block of a reference block. a. For example, a reference block of a reference block may be identified by adding a motion vector associated with an inter-mode coded reference block and the position of such an inter-coded reference block. b. For example, a reference block of a reference block may be identified by adding a block vector associated with a reference block coded in IntraTMP mode and the position of such an IntraTMP coded reference block. c. For example, a reference block of a reference block may be identified by adding a block vector associated with an IBC mode coded reference block and the position of such IBC coded reference block. 4.3 For example, the prediction mode decision can be implicitly determined by the results of video content detection. 1) For example, if the content type detected by the video content indicates that the current video unit belongs to a specific type (eg, screen content), a specific tool may not be allowed / used / applied to the video unit. a. For example, tool on / off flags for specific prediction modes at the video unit level may not be signaled. b. For example, a tool on / off flag for a particular prediction mode at the video unit level may be determined implicitly without signaling (eg, the flag is inferred to be a value indicating that the particular prediction mode is not used for the video unit). c. For example, the specific prediction mode can be at least one of the following: i. Intra-frame luma fusion and / or its variants. ii. Chroma Fusion and / or its variants. iii. IntraTMP fusions and / or variants thereof. iv. DIMD hybrid and / or its variants. v. TIMD hybrid and / or its variants. vi. Conventional inter-frame AMVP / Merge and / or its variants. vii. Affine AMVP / Merge and / or its variants. viii. Affine AMVP / Merge and / or its variants of a specific granularity (eg, 1x1 / pixel / sample based, 4x4 sub-block based). ix. OBMC and / or its variants. x.LIC (and / or its variants). xi.SGPM hybrid and / or its variants. xii. GPM hybrid and / or its variants. xiii. CIIP and / or its variants. xiv. MHP and / or variants thereof. xv.BDOF and / or its variants. xvi. BDOF of a specific granularity (e.g., sample-based, 4x4 / 8x8 / 16x16 sub-block-based). xvii.PROF and / or its variants. xviii.DMVR and / or its variants. xix. DMVR of a specific granularity (e.g., sample-based, 4x4 / 8x8 / 16x16 sub-block-based). xx.AMVP-Merge and / or its variants. xxi.MMVD and / or its variants. xxii.TM and / or its variations. d. For example, a specific prediction model can be a combination of at least two of the above tools: i. For example, OBMC for affine AMVP. ii. For example, OBMC for affine merge. iii. For example, OBMC for inter-frame AMVP. iv. For example, OBMC for inter-frame merging. v. For example, OBMC for MHP when the underlying assumption is affine AMVP / Merge. 2) For example, if the content type obtained by video content detection indicates that the current video unit belongs to a specific type (eg, natural content, camera-captured content), specific tools may not be used for the video unit. a. For example, block level on / off flags may not be transmitted via signals. b. For example, a block-level on / off flag may be derived implicitly (eg, equal to a value indicating that a particular tool is not used for a video unit). c. For example, a specific tool may be at least one of the following: i. IBC and / or its variants. ii. PLT and / or its variants. iii. BDPCM and / or its variants. iv. IntraTMP and / or its variants. v. Affine AMVP / Merge based on 1x1 / pixel / sample. 4.4 In one example, once the inter-frame propagation mode is used for encoding and decoding of the current block, the location of the temporal reference blocks used to derive the inter-frame propagation mode can be restricted to the co-located CTU row. 1) Alternatively, it can be restricted to the co-located CTU and the CTU to the right of the co-located CTU. 2) Alternatively, it can be limited to the co-located CTU and the CTU to the left of the co-located CTU. 3) Alternatively, it can be restricted to no more than M (eg, M=3 or M=4 or M=8) rows of samples above the co-located CTU. 4) Furthermore, alternatively, the inter-frame propagation pattern may be stored in the buffer only when the temporal reference block is within the restricted range / region. 4.5 In one example, once the time domain mode is used for encoding and decoding of the current block, the position of the time domain reference block used to derive the time domain mode can be restricted to no more than M (e.g., M=3 or M=4 or M=8) rows of samples above the co-located CTU. 1) Alternatively, it can be restricted to the co-located CTU and the CTU to the right of the co-located CTU. 2) Alternatively, it can be limited to the co-located CTU and the CTU to the left of the co-located CTU. 3) Alternatively, it can be restricted to co-located CTU rows. 4.6 Whether to apply and / or how to apply the method disclosed above can be transmitted through a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 4.7 Whether and / or how to apply the methods disclosed above may be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other kinds of regions containing more than one sample or pixel. 4.8 Whether and / or how to apply the method disclosed above may depend on the information being coded, such as block size, color format, single / dual tree partitioning, color component, slice type / picture type.

[0101] The following will describe more details of embodiments of the present disclosure related to video content detection. The embodiments of the present disclosure should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these embodiments can be applied alone or combined in any way.

[0102] As used herein, the term "video unit" may refer to a block, a sub-block, a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a codec unit (CU), a prediction unit (PU), a transform unit (TU), a prediction block (PB), a transform block (TB), a slice, a slice, a sub-picture, a video processing unit including a plurality of samples / pixels, etc. A video unit may be rectangular or non-rectangular.

[0103] Figure 39 3900 is a flowchart of a method for video processing according to some embodiments of the present disclosure. The method 3900 may be implemented during the conversion between the current video unit of the video and the bit stream of the video. Figure 39 As shown, method 3900 begins at 3902, where a content type of a current video unit is determined based on values ​​of samples associated with the current video unit.

[0104] For example, the samples associated with the current video unit may be within the current video unit. Additionally or alternatively, the samples associated with the current video unit may be adjacent to the current video unit. For example, the samples associated with the current video unit may include samples within the current video unit and samples surrounding one or more boundaries of the current video unit.

[0105] In some embodiments, the content type of the current video unit may be determined based on the gradient of a sample point, the direction of a sample point, the angle of a sample point, a gradient histogram of a sample point, a direction histogram of a sample point, or an angle histogram of a sample point, etc. This will be described in detail below.

[0106] At 3904, conversion is performed based on the content type. By way of example and not limitation, if the content type of the current video unit is screen content, overlapped sub-block based motion compensation (OBMC) may be disabled for the current video unit. In one example, conversion may include encoding the current video block into a bitstream. Alternatively or additionally, conversion may include decoding the current video block from the bitstream. It should be understood that the above illustrations and / or examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this respect.

[0107] In view of the above, the content type of a video unit is determined based on the values ​​of the samples associated with the video unit. Compared to traditional solutions, the proposed method can advantageously implement adaptive video content detection and thus support the control of codec tools based on the detected video content. In this way, codec quality can be improved.

[0108] In some embodiments, the samples associated with the current video unit may include prediction samples of the current video unit. For example, the prediction samples may include one of the following: prediction samples before overlapping sub-block based motion compensation (OBMC) mixing, prediction samples before OBMC fusion, prediction samples before OBMC weighting, prediction samples before multi-hypothesis prediction (MHP) mixing, prediction samples before MHP fusion, prediction samples before MHP weighting, prediction samples before bidirectional prediction (BCW) mixing with CU level weighting, prediction samples before BCW fusion, prediction samples before BCW weighting, prediction samples before intra-frame and inter-frame joint prediction. The prediction samples before CIIP (CIIP) mixing, the prediction samples before CIIP fusion, the prediction samples before CIIP weighting, the prediction samples before geometric partitioning mode (GPM) mixing, the prediction samples before GPM fusion, the prediction samples before GPM weighting, the prediction samples before spatial geometric partitioning mode (SGPM) mixing, the prediction samples before SGPM fusion, the prediction samples before SGPM weighting, the prediction samples before bidirectional mixing, the prediction samples before bidirectional fusion, and the prediction samples before bidirectional weighting.

[0109] Alternatively, the prediction samples may include prediction samples before a prediction sample refinement process. For example, the prediction sample refinement process may include bidirectional optical flow (BDOF), prediction refinement using optical flow (PROF), local illumination compensation (LIC), OBMC, etc. In other embodiments, the prediction samples may include prediction samples before a sample filtering process. For example, the sample filtering process may include position-dependent intra prediction combination (PDPC), CCIP-PDPC, gradient-PDPC, reference sample filtering, reference sample smoothing, prediction sample filtering, prediction sample smoothing, etc.

[0110] In some embodiments, the samples associated with the current video unit may include predicted samples adjacent to the current video unit. Alternatively, the samples associated with the current video unit may include reconstructed samples adjacent to the current video unit. For example, the reconstructed samples may include reconstructed samples prior to a sample filtering process such as bilateral filtering, deblocking, neural network-based filtering, luma mapping with chroma scaling (LMCS), sample adaptive offset (SAO), cross-component SAO (CCSAO), adaptive loop filter (ALF), cross-component ALF (CCALF), or motion-compensated temporal filtering.

[0111] In some embodiments, the samples associated with the current video unit may be obtained based on a downsampling process. In one example, the downsampling process may be applied to the predicted samples. In another example, the downsampling process may be applied to the reconstructed samples.

[0112] In some embodiments, the downsampling factor used in the downsampling process can be predetermined. For example, the downsampling factor can be predetermined as a fixed value. Alternatively, the downsampling factor can be determined based on one of the following: the width of the current video unit, the height of the current video unit, the width of the video unit adjacent to the current video unit, and / or the height of the video unit adjacent to the current video unit.

[0113] In some embodiments, a gradient histogram, direction histogram, or angle histogram used to determine the content type may be determined based on the result of counting gradients along one or more directions or angles. By way of example and not limitation, the gradient direction of a set of sample points may be determined based on a gradient histogram. Additionally or alternatively, an angle for intra-frame prediction may be determined based on a gradient histogram.

[0114] In some embodiments, one or more directions or angles may be predetermined. Alternatively, one or more directions or angles may be determined based on one or more directions of an intra prediction angle mode, such as Figure 4 shown.

[0115] In some embodiments, for a direction or angle, the gradient magnitude may be determined based on the result of counting the gradient magnitudes of at least one sample point in a region. Alternatively, for a direction or angle, the gradient magnitude may be determined based on the result of counting the gradient magnitudes of at least one sample point in a region. The region may be predetermined.

[0116] In some embodiments, the gradient histogram, direction histogram, or angle histogram can be determined based on dividing the entire range of directions or angles into a series of intervals or bins. For example, the gradient histogram, direction histogram, or angle histogram can be determined based on the result of counting the gradient or the magnitude of the gradient in each interval, or each bin, or each direction, or each angle.

[0117] In some embodiments, the content type of the current video unit can be determined based on the color of the sample, the brightness of the sample, the intensity of the sample, the color histogram of the sample, the brightness histogram of the sample, the intensity histogram of the sample, etc. For example, the color histogram, the brightness histogram, or the intensity histogram can be determined based on the result of counting the values ​​of the samples in the Y component domain, the U component domain, and / or the V component domain. Alternatively, the color histogram, the brightness histogram, or the intensity histogram can be determined based on the result of counting the values ​​of the samples in the red (R) component domain, the green (G) component domain, and / or the blue (B) component domain.

[0118] In one example embodiment, the counted values ​​of samples may include the values ​​of samples at a range of locations in the current video unit. In another example embodiment, the counted values ​​of samples may include the values ​​of all samples in the current video unit. In another example embodiment, the counted values ​​of samples may include the values ​​of samples at locations adjacent to the current video unit. In yet another example embodiment, the counted values ​​of samples may include the values ​​of all samples adjacent to the current video unit. For example, the sample values ​​of samples to the left of the current video unit may be counted. Additionally or alternatively, the sample values ​​of samples above the current video unit may be counted.

[0119] In some embodiments, a color histogram can be determined based on dividing the entire range of colors into a series of intervals or bins. In some other embodiments, a luminance histogram can be determined based on dividing the entire range of luminance into a series of intervals or bins. In still other embodiments, an intensity histogram can be determined based on dividing the entire range of intensity into a series of intervals or bins. For example, a color histogram, a luminance histogram, or an intensity histogram can be determined based on counting the number of samples in each interval or bin.

[0120] In some embodiments, the content type of the current video unit may be determined based on the number of dominant gradients of the samples, the number of dominant directions of the samples, the number of dominant angles of the samples, the number of dominant colors of the samples, the number of dominant luminances of the samples, and / or the number of dominant intensities of the samples.

[0121] In some embodiments, the number of primary gradients for a point can be determined based on a gradient histogram for the point. For example, the number of primary gradients for a point can be determined based on the number of intervals or bins in the gradient histogram that show values ​​greater than a threshold. Alternatively, the number of primary gradients for a point can be determined based on the number of intervals or bins in the gradient histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0122] In some embodiments, the number of dominant directions of a sample point can be determined based on a directional histogram of the sample point. For example, the number of dominant directions of a sample point can be determined based on the number of intervals or bins in the directional histogram that show values ​​greater than a threshold. Alternatively, the number of dominant directions of a sample point can be determined based on the number of intervals or bins in the directional histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0123] In some embodiments, the number of primary angles of a sample point can be determined based on an angle histogram of the sample point. For example, the number of primary angles of a sample point can be determined based on the number of intervals or bins in the angle histogram that show values ​​greater than a threshold. Alternatively, the number of primary angles of a sample point can be determined based on the number of intervals or bins in the angle histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0124] In some embodiments, the number of dominant colors of a sample can be determined based on a color histogram of the sample. For example, the number of dominant colors of a sample can be determined based on the number of intervals or bins in the color histogram that show values ​​greater than a threshold. Alternatively, the number of dominant colors of a sample can be determined based on the number of intervals or bins in the color histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0125] In some embodiments, the number of dominant luminances of the sample points may be determined based on a luminance histogram of the sample points. For example, the number of dominant luminances of the sample points may be determined based on the number of intervals or bins in the luminance histogram that show values ​​greater than a threshold. Alternatively, the number of dominant luminances of the sample points may be determined based on the number of intervals or bins in the luminance histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0126] In some embodiments, the number of principal intensities of a sample point can be determined based on an intensity histogram of the sample point. For example, the number of principal intensities of a sample point can be determined based on the number of intervals or bins in the intensity histogram that show values ​​greater than a threshold. Alternatively, the number of principal intensities of a sample point can be determined based on the number of intervals or bins in the intensity histogram that provide values ​​greater than the values ​​of other intervals or bins.

[0127] In some embodiments, the values ​​of the intervals or bins in the gradient histogram can be sorted. For example, the sorted values ​​are X0, X1, X2, ..., X n-2 、X n-1 If X i > = k*X i+1 , then the intervals or bins from 0 to i are determined to be primary gradients, where k represents the scaling factor, n represents the number of intervals or bins in the gradient histogram, and i is in the range from 0 to n-1. By way of example and not limitation, the scaling factor can be equal to a constant between 2 and 20.

[0128] In some embodiments, codec tools (such as OBMC, sample affine mode, etc.) may not be applied to the current video unit if one of the following is less than a threshold number: the number of dominant gradients of the samples, the number of dominant directions of the samples, the number of dominant angles of the samples, the number of dominant colors of the samples, the number of dominant luminances of the samples, or the number of dominant intensities of the samples. For example, the threshold number may be 1, 2, 3, 4, 5, 6, 7, 8, 9, etc.

[0129] In view of the foregoing, the solutions according to some embodiments of the present disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding quality.

[0130] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. In the method, based on the values ​​of samples associated with a current video unit of the video, the content type of the current video unit is determined. Furthermore, the bitstream is generated based on the content type.

[0131] According to further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, the content type of a current video unit of the video is determined based on the values ​​of samples associated with the current video unit. Furthermore, a bitstream is generated based on the content type, and the bitstream is stored in a non-transitory computer-readable recording medium.

[0132] Embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0133] Item 1. A method for video processing, comprising: for converting between a current video unit of a video and a bitstream of the video, determining a content type of the current video unit based on values ​​of samples associated with the current video unit; and performing the conversion based on the content type.

[0134] Item 2. The method of Item 2, wherein the samples associated with the current video unit are within or adjacent to the current video unit.

[0135] Item 3. A method according to any one of Items 1 to 2, wherein the current video unit comprises one of the following: a block, a sub-block, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a slice, a slice, or a sub-picture.

[0136] Item 4. The method of any one of Items 1 to 3, wherein the samples associated with the current video unit comprise predicted samples for the current video unit.

[0137] Item 5. The method according to Item 4, wherein the prediction sample includes one of the following: a prediction sample before overlapping sub-block motion compensation (OBMC) mixing, a prediction sample before OBMC fusion, a prediction sample before OBMC weighting, a prediction sample before multi-hypothesis prediction (MHP) mixing, a prediction sample before MHP fusion, a prediction sample before MHP weighting, a prediction sample before bidirectional prediction (BCW) mixing with CU level weighting, a prediction sample before BCW fusion, a prediction sample before BCW weighting, a prediction sample before intra-frame and inter-frame joint prediction (CIIP) mixing The prediction samples before CIIP fusion, CIIP weighting, GPM mixing, GPM fusion, GPM weighting, SGPM mixing, SGPM fusion, SGPM weighting, bidirectional mixing, bidirectional fusion, bidirectional weighting, and prediction samples before the refinement process or the sample filtering process.

[0138] Item 6. The method of Item 5, wherein the prediction sample refinement process comprises one of the following: bidirectional optical flow (BDOF), prediction refinement using optical flow (PROF), local illumination compensation (LIC), or OBMC.

[0139] Item 7. A method according to any one of Items 5 to 6, wherein the sample filtering process includes one of the following: position-dependent intra prediction combination (PDPC), CCIP-PDPC, gradient-PDPC, reference sample filtering, reference sample smoothing, prediction sample filtering, or prediction sample smoothing.

[0140] Item 8. The method of any one of Items 1 to 3, wherein the samples associated with the current video unit comprise prediction samples neighboring the current video unit.

[0141] Item 9. The method of any one of Items 1 to 3, wherein the samples associated with the current video unit comprise reconstructed samples adjacent to the current video unit.

[0142] Item 10. The method of Item 9, wherein the reconstructed samples include reconstructed samples prior to a sample filtering process.

[0143] Item 11. A method according to Item 10, wherein the sample filtering process includes one of the following: bilateral filtering, deblocking, neural network-based filtering, luma mapping with chroma scaling (LMCS), sample adaptive offset (SAO), cross-component SAO (CCSAO), adaptive loop filter (ALF), cross-component ALF (CCALF), or motion compensated temporal filtering.

[0144] Item 12. The method of any one of Items 1 to 11, wherein the samples associated with the current video unit are obtained based on a downsampling process.

[0145] Item 13. A method according to Item 12, wherein the downsampling factor used in the downsampling process is predetermined, or the downsampling factor is determined based on one of the following: the width of the current video unit, the height of the current video unit, the width of a video unit adjacent to the current video unit, or the height of the video unit adjacent to the current video unit.

[0146] Item 14. A method according to any one of Items 1 to 13, wherein the content type of the current video unit is determined based on one of: the gradient of the sample point, the direction of the sample point, the angle of the sample point, the gradient histogram of the sample point, the direction histogram of the sample point, or the angle histogram of the sample point.

[0147] Item 15. The method of Item 14, wherein the gradient histogram, the direction histogram, or the angle histogram is determined based on counting gradients along one or more directions or angles.

[0148] Item 16. The method of Item 15, wherein the one or more directions or angles are predetermined, or the one or more directions or angles are determined based on one or more directions of an intra prediction angle mode.

[0149] Item 17. The method according to any one of Items 15 to 16, wherein for a direction or angle, the gradient magnitude is determined based on a result of counting the gradient or the magnitude of the gradient at at least one sample point in the region.

[0150] Item 18. The method of Item 14, wherein the gradient histogram, the direction histogram, or the angle histogram is determined based on dividing the entire range of directions or angles into a series of intervals or bins.

[0151] Item 19. A method according to Item 14, wherein the gradient histogram, the direction histogram or the angle histogram is determined based on the result of counting the gradient or the magnitude of the gradient in each interval, or each bin, or each direction, or each angle.

[0152] Item 20. A method according to any one of Items 1 to 13, wherein the content type of the current video unit is determined based on one of: the color of the sample, the brightness of the sample, the intensity of the sample, the color histogram of the sample, the brightness histogram of the sample, or the intensity histogram of the sample.

[0153] Item 21. A method according to Item 20, wherein the color histogram, the brightness histogram or the intensity histogram is determined based on the result of counting the values ​​of the samples in at least one of the following items: the Y component domain, the U component domain or the V component domain, or the color histogram, the brightness histogram or the intensity histogram is determined based on the result of counting the values ​​of the samples in at least one of the following items: the R component domain, the G component domain or the B component domain.

[0154] Item 22. A method according to Item 21, wherein the counted values ​​of the samples include one of the following: values ​​of samples at a series of positions in the current video unit, values ​​of all samples in the current video unit, values ​​of samples at positions adjacent to the current video unit, or values ​​of all samples adjacent to the current video unit.

[0155] Item 23. A method according to Item 20, wherein the color histogram is determined based on dividing the entire range of colors into a series of intervals or bins, or wherein the brightness histogram is determined based on dividing the entire range of brightness into a series of intervals or bins, or wherein the intensity histogram is determined based on dividing the entire range of intensity into a series of intervals or bins.

[0156] Item 24. The method of Item 20, wherein the color histogram, the brightness histogram, or the intensity histogram is determined based on a result of counting the number of samples in each interval or bin.

[0157] Item 25. A method according to any one of Items 1 to 13, wherein the content type of the current video unit is determined based on at least one of: the number of primary gradients of the samples, the number of primary directions of the samples, the number of primary angles of the samples, the number of primary colors of the samples, the number of primary luminances of the samples, or the number of primary intensities of the samples.

[0158] Item 26. A method according to Item 25, wherein the number of the main gradients of the sample is determined based on the gradient histogram of the sample, or the number of the main directions of the sample is determined based on the direction histogram of the sample, or the number of the main angles of the sample is determined based on the angle histogram of the sample, or the number of the main colors of the sample is determined based on the color histogram of the sample, or the number of the main brightness of the sample is determined based on the brightness histogram of the sample, or the number of the main intensities of the sample is determined based on the intensity histogram of the sample.

[0159] Item 27. A method according to Item 26, wherein the number of the main gradients of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the gradient histogram, or the number of the main directions of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the direction histogram, or the number of the main angles of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the angle histogram, or the number of the main colors of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the color histogram, or the number of the main brightness of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the brightness histogram, or the number of the main intensities of the sample points is determined based on the number of intervals or bins showing values ​​greater than a threshold in the intensity histogram.

[0160] Item 28. A method according to Item 26, wherein the number of the main gradients of the sample points is determined based on the number of the following intervals or bins in the gradient histogram, which intervals or bins provide values ​​that are larger than the values ​​of other intervals or bins, or the number of the main directions of the sample points is determined based on the number of the following intervals or bins in the direction histogram, which intervals or bins provide values ​​that are larger than the values ​​of other intervals or bins, or the number of the main angles of the sample points is determined based on the number of the following intervals or bins in the angle histogram, which intervals or bins provide values ​​that are larger than the values ​​of other intervals or bins, or the number of the main colors of the samples is determined based on the number of intervals or bins in the color histogram, which intervals or bins provide values ​​larger than the values ​​of other intervals or bins, or the number of the main brightness of the samples is determined based on the number of intervals or bins in the brightness histogram, which intervals or bins provide values ​​larger than the values ​​of other intervals or bins, or the number of the main intensities of the samples is determined based on the number of intervals or bins in the intensity histogram, which intervals or bins provide values ​​larger than the values ​​of other intervals or bins.

[0161] Item 29. The method of Item 28, wherein the values ​​of the bins or bins in the gradient histogram are sorted.

[0162] Item 30. The method of Item 29, wherein if X i > = k*X i+1 , then the interval or bin from 0 to i is determined as the main gradient, where k represents the scaling factor, and the values ​​after the sorting are X0, X1, X2, ..., X n-2 、X n-1 , n represents the number of intervals or bins in the gradient histogram, and i ranges from 0 to n-1.

[0163] Item 31. The method of Item 30, wherein the scaling factor is equal to a constant between 2 and 20.

[0164] Item 32. A method according to any one of Items 25 to 31, wherein the codec tool is not applied to the current video unit if one of the following is less than a threshold number: the number of the main gradients of the samples, the number of the main directions of the samples, the number of the main angles of the samples, the number of the main colors of the samples, the number of the main luminances of the samples, or the number of the main intensities of the samples.

[0165] Item 33. The method of Item 32, wherein the threshold number is one of: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

[0166] Item 34. A method according to any one of Items 1 to 33, wherein the converting comprises encoding the current video unit into the bitstream.

[0167] Item 35. The method of any one of Items 1 to 33, wherein the converting comprises decoding the current video unit from the bitstream.

[0168] Item 36. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 35.

[0169] Item 37. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 35.

[0170] Item 38. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a content type of a current video unit of the video based on a value of a sample associated with the current video unit; and generating the bitstream based on the content type.

[0171] Item 39. A method for storing a bitstream of a video, comprising: determining a content type of a current video unit of the video based on a value of a sample associated with the current video unit; generating the bitstream based on the content type; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0172] Figure 40 A block diagram of a computing device 4000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0173] It should be understood that Figure 40 The computing device 4000 shown in FIG. 4 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0174] like Figure 40As shown, computing device 4000 comprises a general computing device 4000. Computing device 4000 may include at least one or more processors or processing units 4010, memory 4020, storage unit 4030, one or more communication units 4040, one or more input devices 4050, and one or more output devices 4060.

[0175] In some embodiments, the computing device 4000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4000 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0176] The processing unit 4010 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 4020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 4000. The processing unit 4010 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0177] The computing device 4000 typically includes various computer storage media. Such media can be any media accessible by the computing device 4000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4020 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 4030 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 4000.

[0178] The computing device 4000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 40 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0179] The communication unit 4040 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 4000 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 4000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0180] Input device 4050 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 4060 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 4040, computing device 4000 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 4000, or, if desired, any device that enables computing device 4000 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).

[0181] In some embodiments, some or all components of the computing device 4000 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers in a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0182] In an embodiment of the present disclosure, the computing device 4000 may be used to implement video encoding / decoding. The memory 4020 may include one or more video encoding / decoding modules 4025 having one or more program instructions. These modules are accessible and executable by the processing unit 4010 to perform the functions of the various embodiments described herein.

[0183] In an example embodiment performing video encoding, an input device 4050 may receive video data as input to be encoded 4070. The video data may be processed, for example, by a video codec module 4025 to generate an encoded bitstream. The encoded bitstream may be provided as output 4080 via an output device 4060.

[0184] In an example embodiment performing video decoding, an input device 4050 may receive an encoded bitstream as input 4070. The encoded bitstream may be processed, for example, by a video codec module 4025 to generate decoded video data. The decoded video data may be provided as output 4080 via an output device 4060.

[0185] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For conversion between a current video unit of a video and a bitstream of the video, determining a content type of the current video unit based on values ​​of samples associated with the current video unit; as well as The conversion is performed based on the content type.

2. The method of claim 2, wherein the samples associated with the current video unit are within or adjacent to the current video unit.

3. The method according to any one of claims 1 to 2, wherein the current video unit comprises one of the following: piece, sub-block, Codec Unit (CU), Prediction Unit (PU), Transformation Unit (TU), piece, strips, or Sub-picture.

4. The method of any one of claims 1 to 3, wherein the samples associated with the current video unit comprise prediction samples of the current video unit.

5. The method according to claim 4, wherein the predicted sample point comprises one of the following: Prediction samples before overlapping sub-block based motion compensation (OBMC) blending, Prediction sample points before OBMC fusion, The predicted sample points before OBMC weighting, The prediction samples before the multi-hypothesis prediction (MHP) mixing, Prediction sample points before MHP fusion, The predicted sample points before MHP weighting, The prediction samples before using CU-level weighted bidirectional prediction (BCW) mixing, The predicted sample points before BCW fusion, The predicted sample points before BCW weighting, The prediction samples before the CIIP mixing, Prediction sample points before CIIP fusion, Forecast sample points before CIIP weighting, Prediction samples before GPM mixing, Prediction samples before GPM fusion, The predicted sample points before GPM weighting, The predicted samples before the spatial geometric partitioning mode (SGPM) mixing, Prediction samples before SGPM fusion, The predicted sample points before SGPM weighting, The prediction samples before bidirectional mixing, The predicted sample points before bidirectional fusion, The predicted sample points before bidirectional weighting, The prediction samples before the prediction sample refinement process, or The predicted samples before the sample filtering process.

6. The method according to claim 5, wherein the prediction sample point refinement process comprises one of the following: Bidirectional Optical Flow (BDOF), Using optical flow prediction refinement (PROF), Local Illumination Compensation (LIC), or OBMC.

7. The method according to any one of claims 5 to 6, wherein the sample filtering process comprises one of the following: Position-dependent intra prediction combination (PDPC), CCIP-PDPC, Gradient-PDPC, Reference sample filtering, Reference sample smoothing, Prediction sample filtering, or Prediction sample smoothing.

8. The method of any one of claims 1 to 3, wherein the samples associated with the current video unit comprise prediction samples neighboring the current video unit.

9. The method of any one of claims 1 to 3, wherein the samples associated with the current video unit include reconstructed samples adjacent to the current video unit.

10. The method of claim 9, wherein the reconstructed samples include reconstructed samples before a sample filtering process.

11. The method according to claim 10, wherein the sample filtering process comprises one of the following: Bilateral filtering, Go to the block, Neural network-based filtering, Luminance Mapping with Chroma Scaling (LMCS), Sample Adaptive Overshoot (SAO), Cross-Component SAO (CCSAO), Adaptive loop filter (ALF), Cross-Component ALF (CCALF), or Temporal filtering based on motion compensation.

12. The method according to any one of claims 1 to 11, wherein the samples associated with the current video unit are obtained based on a downsampling process.

13. The method according to claim 12, wherein the downsampling factor used in the downsampling process is predetermined, or The downsampling factor is determined based on: The width of the current video unit, The height of the current video unit, the width of the video unit adjacent to the current video unit, or The heights of the video units adjacent to the current video unit.

14. The method according to any one of claims 1 to 13, wherein the content type of the current video unit is determined based on one of: The gradient of the sample point, The direction of the sample point, The angle of the sample point, The gradient histogram of the sample point, The direction histogram of the sample points, or Angle histogram of the sample points. 15 . The method according to claim 14 , wherein the gradient histogram, the direction histogram, or the angle histogram is determined based on a result of counting gradients along one or more directions or angles.

16. The method of claim 15, wherein the one or more directions or angles are predetermined, or The one or more directions or angles are determined based on one or more directions of the intra prediction angle mode. 17 . The method according to claim 15 , wherein for a direction or an angle, a gradient magnitude is determined based on a result of counting the gradient or the magnitude of the gradient of at least one sample point in the region.

18. The method of claim 14, wherein the gradient histogram, the direction histogram, or the angle histogram is determined based on dividing the entire range of directions or angles into a series of intervals or bins.

19. The method according to claim 14, wherein the gradient histogram, the direction histogram or the angle histogram is determined based on a result of counting the gradient or the magnitude of the gradient in each interval, or each bin, or each direction, or each angle.

20. The method of any one of claims 1 to 13, wherein the content type of the current video unit is determined based on one of: The color of the sample point, The brightness of the sample point, The intensity of the sample point, The color histogram of the sample points, the brightness histogram of the sample points, or The intensity histogram of the sample points.

21. The method according to claim 20, wherein the color histogram, the brightness histogram or the intensity histogram is determined based on a result of counting the values ​​of the samples in at least one of the following: a Y component domain, a U component domain or a V component domain, or The color histogram, the brightness histogram, or the intensity histogram is determined based on a result of counting the values ​​of the samples in at least one of the following: an R component domain, a G component domain, or a B component domain.

22. The method of claim 21 , wherein the counted value of the sample points comprises one of: the values ​​of samples at a series of positions in the current video unit, The values ​​of all samples in the current video unit, The value of the sample at a position adjacent to the current video unit, or The values ​​of all samples adjacent to the current video unit.

23. The method of claim 20, wherein the color histogram is determined based on dividing the entire range of colors into a series of intervals or bins, or wherein the brightness histogram is determined based on dividing the entire range of brightness into a series of intervals or bins, or The intensity histogram is determined based on dividing the entire range of intensities into a series of intervals or bins.

24. The method according to claim 20, wherein the color histogram, the brightness histogram, or the intensity histogram is determined based on a result of counting the number of samples in each interval or bin.

25. The method of any one of claims 1 to 13, wherein the content type of the current video unit is determined based on at least one of: The number of main gradients at the sample point, the number of main directions of the sample points, the number of main angles of the sample points, The number of main colors of the sample point, The number of main luminances of the sample points, or The number of main intensities of the sample point.

26. The method of claim 25, wherein the number of the main gradients of the sample point is determined based on a gradient histogram of the sample point, or The number of the main directions of the sample points is determined based on a direction histogram of the sample points, or The number of the main angles of the sample point is determined based on an angle histogram of the sample point, or The number of the dominant colors of the sample is determined based on a color histogram of the sample, or The number of the main luminances of the samples is determined based on a luminance histogram of the samples, or The number of the main intensities of the sample points is determined based on an intensity histogram of the sample points.

27. The method of claim 26, wherein the number of the main gradients of the sample point is determined based on the number of bins or segments showing values ​​greater than a threshold value in the gradient histogram, or The number of the main directions of the sample points is determined based on the number of bins or sections showing values ​​greater than a threshold value in the direction histogram, or The number of the main angles of the sample point is determined based on the number of bins or sections showing values ​​greater than a threshold value in the angle histogram, or The number of the dominant colors of the sample is determined based on the number of bins or segments in the color histogram showing values ​​greater than a threshold, or The number of the main luminances of the sample points is determined based on the number of bins or sections showing values ​​greater than a threshold value in the luminance histogram, or The number of the main intensities of the sample points is determined based on the number of bins or segments in the intensity histogram showing values ​​greater than a threshold value.

28. The method of claim 26, wherein the number of the main gradients of the sample point is determined based on the number of intervals or bins in the gradient histogram that provide values ​​that are larger than the values ​​of other intervals or bins, or The number of main directions of the sample points is determined based on the number of intervals or bins in the direction histogram that provide values ​​that are greater than the values ​​of other intervals or bins, or The number of the main angles of the sample point is determined based on the number of intervals or bins in the angle histogram that provide values ​​that are larger than the values ​​of other intervals or bins, or The number of the dominant colors of the sample is determined based on the number of intervals or bins in the color histogram that provide values ​​that are greater than the values ​​of other intervals or bins, or The number of the dominant luminances of the sample points is determined based on the number of bins or sections in the luminance histogram that provide values ​​that are greater than the values ​​of other bins or sections, or The number of the main intensities of the sample points is determined based on the number of bins or sections in the intensity histogram that provide values ​​that are greater than the values ​​of other bins or sections.

29. The method of claim 28, wherein the values ​​of the bins or segments in the gradient histogram are sorted.

30. The method of claim 29, wherein if X i > = k*X i+1 , then the interval or bin from 0 to i is determined as the main gradient, where k represents the scaling factor, and the values ​​after the sorting are X0, X1, X2, ..., X n-2 、X n-1 , n represents the number of intervals or bins in the gradient histogram, and i ranges from 0 to n-1. The method of claim 30 , wherein the scaling factor is equal to a constant between 2 and 20.

32. The method of any one of claims 25 to 31, wherein codec tools are not applied to the current video unit if one of the following is less than a threshold number: the number of said main gradients of said sample points, the number of the main directions of the sample points, the number of the main angles of the sample points, the number of the primary colors of the samples, The number of the main luminances of the samples, or The number of the main intensities of the sample points.

33. The method of claim 32, wherein the threshold number is one of: 1, 2, 3, 4, 5, 6, 7, 8, or 9.

34. The method of any one of claims 1 to 33, wherein the converting comprises encoding the current video unit into the bitstream.

35. The method of any one of claims 1 to 33, wherein the converting comprises decoding the current video unit from the bitstream.

36. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 35.

37. A non-transitory computer-readable storage medium storing instructions, the instructions causing a processor to execute the method according to any one of claims 1 to 35.

38. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining a content type of a current video unit of the video based on values ​​of samples associated with the current video unit; as well as The bitstream is generated based on the content type.

39. A method for storing a bitstream of a video, comprising: determining a content type of a current video unit of the video based on values ​​of samples associated with the current video unit; generating the bitstream based on the content type; as well as The bitstream is stored in a non-transitory computer-readable recording medium.