Method and device for video processing and medium

By using adaptive OBMC technology in video encoding and decoding, analyzing the prediction mode and codec tools of the video unit, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and a higher encoding and decoding gain is achieved.

CN120359754APending Publication Date: 2025-07-22DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085362.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-12
Filing Date
2023-12-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially when processing conversion between video units, it is difficult for existing methods to make full use of the prediction modes and codec tools of adjacent blocks to improve the codec gain.

Method used

Motion compensation (OBMC) technology based on overlapping sub-blocks is used to analyze the prediction mode of the nearest neighbor block, the prediction mode of the reference block and the target codec tool to determine whether it is applied to the current block, and realize the block-level adaptive OBMC to improve the encoding and decoding efficiency.

Benefits of technology

Through adaptive OBMC technology, the encoding and codec gain of video encoding and codec is improved and the encoding and codec efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359754A_ABST
    Figure CN120359754A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises the following steps: aiming at conversion between a video unit of a video and a bit stream of the video unit, determining whether overlapping sub-block-based motion compensation (OBMC) is applied to a current block of the video unit based on at least one of: a prediction mode of a neighbor block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; and performing the conversion based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to motion compensation (OBMC) based on block-level adaptive overlapping sub-blocks that rely on neighboring prediction modes in video coding and decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding and decoding efficiency of video coding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a video unit of a video and a bitstream of the video unit, determining whether overlapping block motion compensation (OBMC) is applied to a current block of the video unit based on at least one of the following: the prediction mode of a neighboring block, the prediction mode of a reference block, the prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; and performing the conversion based on the determination. In this way, block-level adaptive OBMC that takes into account the prediction mode of neighboring blocks can bring higher coding and decoding gains and improve coding and decoding efficiency.

[0005] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video, and the bitstream of the video is generated by a method executed by a device for video processing. The method includes: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; and generating a bitstream based on the determination.

[0008] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0009] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 A block diagram showing an exemplary video codec system according to some embodiments of the present disclosure is shown;

[0012] Figure 2 A block diagram showing a first exemplary video encoder according to some embodiments of the present disclosure is shown;

[0013] Figure 3 A block diagram showing an exemplary video decoder according to some embodiments of the present disclosure is shown;

[0014] Figure 4 An intra prediction mode is shown;

[0015] Figure 5 Reference sample points for wide-angle intra prediction are shown;

[0016] Figure 6 The problem of discontinuity in the case where the direction exceeds 45° is shown;

[0017] Figure 7A Schematic diagram showing the definition of samples used by PDPC of the diagonal upper right pattern applied to the diagonal and adjacent angular intra modes;

[0018] Figure 7B Schematic diagram showing the definition of samples used by PDPC of the diagonal lower left pattern applied to the diagonal and adjacent angular intra modes;

[0019] Figure 7C Schematic diagram showing the definition of samples used by PDPC of the adjacent diagonal upper right pattern applied to the diagonal and adjacent angular intra modes;

[0020] Figure 7D Schematic diagram showing the definition of samples used by PDPC of the adjacent diagonal lower left pattern applied to the diagonal and adjacent angular intra modes;

[0021] Figure 8 Example of four reference lines adjacent to the prediction block is shown;

[0022] Figure 9A Schematic diagram showing the process of sub - division depending on the block size;

[0023] Figure 9B Schematic diagram showing the process of sub - division depending on the block size;

[0024] Figure 10 Matrix weighted intra prediction process is shown;

[0025] Figure 11 Spatial GPM candidates are shown;

[0026] Figure 12 GPM template is shown;

[0027] Figure 13 GPM mixing is shown;

[0028] Figure 14 Position of spatial Merge candidates is shown;

[0029] Figure 15 Candidate pairs considered for redundancy check for spatial Merge candidates are shown;

[0030] Figure 16 Schematic diagram showing the motion vector scaling for temporal Merge candidates;

[0031] Figure 17 Candidate positions for temporal Merge candidates (C0 and C1) are shown;

[0032] Figure 18Shows the MMVD search points;

[0033] Figure 19 Shows the extended CU region used in BDOF;

[0034] Figure 20 Shows the schematic diagram for the symmetric MVD mode;

[0035] Figure 21 Shows the motion vector refinement on the decoding side;

[0036] Figure 22 Shows the top and left neighboring blocks used in CIIP weight derivation;

[0037] Figure 23 Shows an example of GPM partitioning grouped at the same angle;

[0038] Figure 24 Shows the unidirectional prediction MV selection for the geometric segmentation mode;

[0039] Figure 25 Shows the exemplary generation of the hybrid weight w0 using the geometric segmentation mode;

[0040] Figure 26 Shows the current CTU processing order and its available reference sample points in the current CTU and the left CTU;

[0041] Figure 27 Shows the residual encoding and decoding process for the transform skip block;

[0042] Figure 28 Shows an example of the block encoded and decoded in the palette mode;

[0043] Figure 29 Shows the sub-block based index map scanning for the palette, with the left side for horizontal scanning and the right side for vertical scanning;

[0044] Figure 30 Shows the decoding flowchart using ACT;

[0045] Figure 31 Shows the in-frame template matching search area used;

[0046] Figure 32 Shows five positions in the reconstructed luminance samples;

[0047] Figure 33 Shows the prediction process of the DBV mode;

[0048] Figure 34 Shows the low-frequency non-separable transform (LFNST) process;

[0049] Figure 35 Shows the SBT position, type, and transformation type;

[0050] Figure 36 Shows the ROI for LFNST16;

[0051] Figure 37 Shows the ROI for LFNST8;

[0052] Figure 38 Shows the discontinuity measurement;

[0053] Figure 39 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure;

[0054] Figure 40 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.

[0055] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description

[0056] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.

[0057] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.

[0058] As used in the present disclosure, the phrases "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0059] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0060] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Exemplary Environment

[0061] Figure 1 is a block diagram showing an exemplary video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0062] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.

[0063] Video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 over the network 130A. The encoded video data may also be stored on the storage medium / server 130B for access by the destination device 120.

[0064] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.

[0065] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.

[0066] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.

[0067] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0068] In some embodiments, the video encoder 200 may include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.

[0069] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0070] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are shown separately in the Figure 2 examples.

[0071] The splitting unit 201 may split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0072] The mode selection unit 203 may select, for example, one coding mode among multiple coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a Combined Intra-Inter Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0073] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.

[0074] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on a current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to portions of a picture composed of macroblocks that are independent of macroblocks within the same picture.

[0075] In some examples, the motion estimation unit 204 may perform uni-directional prediction on a current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0076] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction on a current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find one reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, the plurality of reference indices indicating the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicating the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0077] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.

[0078] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0079] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0080] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0081] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0082] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0083] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0084] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0085] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0086] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transformed coefficient video block, respectively, to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0087] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce block effect artifacts in the video block.

[0088] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.

[0089] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 an example of the video decoder 124 in the system 100 shown.

[0090] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 3 an example, the video decoder 300 includes multiple functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0091] In Figure 3 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.

[0092] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.

[0093] The motion compensation unit 302 can generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.

[0094] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0095] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice can be the entire picture or can also be a region of the picture.

[0096] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0097] The reconstruction unit 306 can obtain the decoded block, for example, by adding a residual block to a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also produces the decoded video for presentation on a display device.

[0098] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to the multi-functional video codec or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video codec technologies. Specifically, the present disclosure relates to overlapping sub-block based motion compensation (OBMC) in image / video coding and related technologies. It can be applied to existing video codec standards such as HEVC, VVC, etc. It can also be applicable to future video codec standards or video codecs. 2. Introduction Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. Since H.262, video coding standards are based on a hybrid video coding structure in which temporal prediction and transform coding are exploited. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. JVET meetings are held quarterly, and the new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was also released at that time. The working draft of VVC and the test model VTM are updated after each meeting. The VVC project achieved technical completion (FDIS) at the meeting in July 2020. 2.1. Existing Coding Tools 2.1.1. Intra Prediction 2.1.1.1. Intra Mode Coding with 67 Intra Prediction Modes To capture any edge direction presented in natural videos, the number of directional intra modes in VVC is extended from 33 used in HEVC to 65. The new directional modes not in HEVC are Figure 4 depicted as red dashed arrows in, and the planar mode and DC mode remain the same. These denser directional intra prediction modes apply to all block sizes and are used for both luma intra prediction and chroma intra prediction. In VVC, for non-square blocks, multiple conventional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. In HEVC, each intra-coded block has a square shape and the length of each of its sides is a power of 2. Thus, no division operation is required to generate the intra prediction factor using the DC mode. In VVC, blocks can have a rectangular shape, which in general requires the use of division operations for each block. To avoid division operations for DC prediction, only the longer side is used to calculate the average value for non-square blocks. 2.1.1.2. Intra Mode Coding To maintain the low complexity of the most probable mode (MPM) list generation, an intra mode coding method with 6 MPMs is used by considering two available neighboring intra modes. The following three aspects are considered to construct the MPM list: – Default intra mode; – Intra-neighboring mode; – Derived intra mode. A unified 6-MPM list is used for intra blocks regardless of whether the MRL and ISP codec tools are applied. The MPM list is constructed based on the intra modes of the left and upper neighboring blocks. Assume that the mode of the left neighbor is denoted as Left, and the mode of the upper block is denoted as Above. Then the unified MPM list is constructed as follows: – When the neighboring blocks are not available, their intra modes are default set to Planar. – If both modes Left and Above are non-angular modes: – MPM list → {Planar, DC, V, H, V-4, V+4}. – If one of the modes Left and Above is an angular mode and the other is a non-angular mode: – Set the mode Max to the larger mode of Left and Above – MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2}. – If both Left and Above are angular modes and they are different: – Set the mode Max to the larger mode of Left and Above – If the difference between the modes Left and Above is in the range of 2 to 62 (including 2 and 62) – MPM list → {Planar, Left, Above, DC, Max-1, Max+1} – Otherwise – MPM list → {Planar, Left, Above, DC, Max-2, Max+2}. – If both Left and Above are angular modes and they are the same: – MPM list → {Planar, Left, Left-1, Left+1, DC, Left-2}. In addition, the first binary bit of the mpm index codeword is context encoded by CABAC. A total of three contexts are used, corresponding to whether the current intra block is MRL-enabled, ISP-enabled, or a normal intra block. During the 6MPM list generation process, deduplication is used to remove duplicate modes so that only unique modes can be included in the MPM list. For the entropy coding of the 61 non-MPM modes, Truncated Binary Code (TBC) is used. 2.1.1.3. Wide-angle Intra Prediction for Non-square Blocks The conventional angular intra prediction directions are defined from 45 degrees clockwise to -135 degrees. In VVC, for non-square blocks, multiple conventional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode encoding and decoding methods also remain unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined as Figure 5 shown. The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 1. Table 1 - Intra prediction modes replaced by wide-angle modes As Figure 6 shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction applications to reduce the increased gap Δp α 's negative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, i.e., [-14, -12, -10, -6, 72, 76, 78, 80]. When the block is predicted by these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, this modification also aligns the design of non-fractional modes in the conventional prediction mode with the wide-angle mode. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was initially ported from HEVC, and the number of entries in this table was extended from 35 to 67 to align with the extension of the intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luminance intra prediction modes in the range from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the entries in the mapping table to more accurately convert the prediction angles for chroma blocks. 2.1.1.4. Mode-Dependent Intra Smoothing (MDIS) A four-tap intra interpolation filter is utilized to improve the directional intra prediction accuracy. In HEVC, a two-tap linear interpolation filter has been used to generate intra prediction blocks in the directional prediction mode (i.e., excluding the planar and DC predictors). In VVC, a simplified 6-bit four-tap Gaussian interpolation filter is only used for the directional intra mode. The non-directional intra prediction process is not modified. The selection of the four-tap filter is performed according to the MDIS condition for the directional intra prediction mode that provides non-fractional displacements (i.e., all directional modes except the following: 2, HOR_IDX, DIA_IDX, VER_IDX, 66). According to the intra prediction mode, the following reference sample processing is performed: – The directional intra prediction mode is classified into one of the following groups: – Vertical or horizontal modes (HOR_IDX, VER_IDX), – Diagonal modes (2, DIA_IDX, VDIA_IDX) representing multiples of 45 degrees, – The remaining directional modes; – If the directional intra prediction mode is classified as belonging to Group A, no filter is applied to the reference samples to generate the prediction samples; – Otherwise, if the mode falls into Group B, a [1,2,1] reference sample filter can be applied (depending on the MDIS condition) to the reference samples to further copy these filtered values into the intra predictors according to the selected direction, but no interpolation filter is applied; – Otherwise, if the mode is classified as belonging to Group C, only the intra reference sample interpolation filter is applied to the reference samples to generate prediction samples that fall at fractional or integer positions between the reference samples according to the selected direction (no reference sample filtering is performed). 2.1.1.5. Position-Dependent Intra Prediction Combination In VVC, the intra prediction results of the DC, planar, and multiple angle modes are further modified by the position-dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, lower left angle mode and its eight adjacent angle modes, and upper right angle mode and its eight adjacent angle modes. The prediction sample pred(x',y') is predicted according to Equation 3-8 using a linear combination of the intra prediction mode (DC, planar, angle) and the reference samples: pred(x',y') = (wL × R -1,y’+wT×R x ’ ,-1 -wTL×R -1,-1 +(64 - wL - wT+wTL)×pred(x',y') + 32) >> 6(2 - 1) where R x,-1 、R -1,y respectively represent the reference samples located at the top and left boundaries of the current sample (x, y), and R -1,-1 represents the reference sample located at the upper left corner of the current block. If PDPC is applied to DC, planar, horizontal, and vertical intra - modes, no additional boundary filters are required, as are required in the case of the HEVC DC - mode boundary filter or the horizontal / vertical - mode edge filter. The PDPC processes for DC and planar modes are the same, and the clipping operation is avoided. For angular modes, the pdpc scaling factor is adjusted such that no range check is required, and the pdpc - related condition on the angle is removed (scale >= 0 is used). Additionally, in all angular - mode cases, the PDPC weights are based on 32. The PDPC weights depend on the prediction mode and are shown in Table 2. PDPC is applied to blocks where both the width and height are greater than or equal to 4. Figures 7A - 7D shows the definition of the reference samples (R x,-1 、R -1,y and R -1,-1 ) for PDPC applied to various prediction modes. The predicted sample pred(x',y') is located at (x',y') within the prediction block. For example, for the diagonal mode, the x - coordinate of the reference sample R x,-1 is given by: x = x'+y'+1, and the y - coordinate of the reference sample R -1,y is similarly given by: y = x'+y'+1. For other ring - shaped modes, the reference samples R x,-1 and R -1,y can be located at fractional - sample positions. In this case, the sample value at the nearest integer - sample position is used. Table 2 - Examples of PDPC Weights According to Prediction Mode 2.1.1.6. Multiple - Reference - Line (MRL) Intra - Prediction Multiple - Reference - Line (MRL) intra - prediction uses more reference lines for intra - prediction. In Figure 8In it, examples of 4 reference lines are depicted, where the samples of segment A and segment F are not obtained from the reconstructed neighboring samples, but are filled respectively using the closest samples from segment B and segment E. HEVC intra picture prediction uses the closest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used. The index (mrl_idx) of the selected reference line is signaled and used to generate the intra prediction factor. For reference line idx greater than 0, only the additional reference line mode is included in the MPM list, and only the mpm index is signaled without signaling the remaining modes. The reference line index is signaled before the intra prediction mode, and in the case where a non-zero reference line index is signaled, the planar mode is excluded from the intra prediction mode. MRL is disabled for the first row of blocks inside the CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional lines are used. For the MRL mode, the derivation of the DC value in the DC intra prediction mode for non-zero reference line indices is aligned with the derivation for reference line index 0. MRL requires storing 3 neighboring luma reference lines and CTUs to generate the prediction. The cross-component linear model (CCLM) tool also requires 3 neighboring luma reference lines for the downsampling filter of the cross-component linear model (CCLM) tool. The definition of MLR using the same 3 lines is aligned with CCLM to reduce the storage requirements for the decoder. 2.1.1.7. Intra Sub-Partitioning (ISP) Intra Sub-Partitioning (ISP) divides the luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions according to the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is greater than 4x8 (or 8x4), the corresponding block is divided by 4 sub-partitions. It has been noted that ISP blocks of M×128 (when M≤64) and 128×N (when N≤64) may generate potential problems related to the 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding chroma TBs. If the CU uses ISP, the luma TB will be divided into four M×32 TBs (only horizontal division is possible), and each TB is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, both chroma components will have a size greater than 32×32 blocks. Similarly, a 128×N CU using ISP can create a similar situation. Therefore, these two cases are problems for the 64×64 decoder pipeline. For this reason, the CU sizes for which ISP can be used are limited to a maximum of 64×64. Figure 9A and 9BAn example showing two possibilities is presented. All sub - partitions satisfy the condition of having at least 16 samples. In the ISP, the dependence of the 1xN / 2xN sub - block prediction on the reconstructed values of the previously decoded 1xN / 2xN sub - blocks of the coded - decoded block is not allowed, such that the minimum prediction width for the sub - block becomes four samples. For example, an 8xN (N > 4) coded - decoded block coded - decoded using an ISP with vertical partitioning is divided into two prediction regions each of size 4xN and four transforms of size 2xN. Additionally, a 4xN coded - decoded block coded - decoded using an ISP with vertical partitioning uses the full 4xN block for prediction; four 1xN transforms are used. Although 1xN and 2xN transform sizes are allowed, it is claimed that the transforms of these blocks in the 4xN region can be executed in parallel. For example, when the 4xN prediction region contains four 1xN transforms, there is no transform in the horizontal direction; the transform in the vertical direction can be executed as a single 4xN transform in the vertical direction. Similarly, when the 4xN prediction region contains two 2xN transform blocks, the transform operations of the two 2xN blocks in each direction (horizontal and vertical) can be performed in parallel. Thus, no additional delay is added when processing these smaller blocks compared to processing an intra - block of 4x4 conventional coding - decoding. Table 3 - Entropy Coding - Decoding Coefficient Group Sizes Block size Coefficient group size 1×N, N≥16 1×16 N×1, N≥16 16×1 2×N, N≥8 2×8 N×2, N≥8 8×2 All other possible M×N cases 4×4 For each sub - partition, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated through processes such as entropy decoding, inverse quantization, and inverse transformation. Thus, the reconstructed sample values of each sub - partition can be used to generate the prediction for the next sub - partition, and each sub - partition is processed iteratively. Additionally, the first sub - partition to be processed is the sub - partition containing the top - left sample of the CU, and then continues down (horizontal partitioning) or to the right (vertical partitioning). Thus, the reference samples used to generate the sub - partition prediction signal are only located to the left and above the row. All sub - partitions share the same intra - mode. The following is a summary of the interaction between the ISP and other coding - decoding tools. – Multiple Reference Lines (MRL): If the block has an MRL index other than 0, the ISP coding - decoding mode will be presumed to be 0, and thus the ISP mode information will not be sent to the decoder. – Entropy Coding - Decoding Coefficient Group Sizes: The sizes of the entropy - coded - decoded sub - blocks have been modified such that they have 16 samples in all possible cases, as shown in Table 3. Note that the new sizes only affect the blocks where one of the dimensions produced by the ISP is less than 4 samples. In all other cases, the coefficient group remains at 4×4 dimension. – CBF Coding - Decoding: Assume that at least one sub - partition has a non - zero CBF. Thus, if n is the number of sub - partitions and the first n - 1 sub - partitions produce zero CBF, the CBF of the nth sub - partition is presumed to be 1. – MPM Usage: The MPM flag is presumed to be 1 in blocks coded in ISP mode, and the MPM list is modified to exclude the DC mode and prioritize the intra - horizontal mode for ISP horizontal partitioning and the intra - vertical mode for ISP vertical partitioning. – Transform Size Limit: All ISP transforms with length greater than 16 samples use DCT - II. – PDPC: When the CU uses the ISP coding mode, the PDPC filter will not be applied to the resulting sub - partitions. – MTS Flag: If the CU uses the ISP coding mode, the MTS CU flag will be set to 0 and it will not be sent to the decoder. Thus, the encoder will not perform RD tests for different available transforms for each resulting sub - partition. Instead, the transform selection for the ISP mode is fixed and selected according to the intra - mode utilized, processing order, and block size. Thus, no signaling is required. For example, assume t H and t V are the horizontal and vertical transforms selected for a w×h sub - partition respectively, where w is the width and h is the height. Then the transform is selected according to the following rules: – If w = 1 or h = 1, there is no horizontal transform or vertical transform respectively. – If w = 2 or w>32, t H = DCT - II – If h = 2 or h>32, t V = DCT - II – Otherwise, the transform is selected as shown in Table 4. Table 4 - Transform Selection Depends on Intra - mode In ISP mode, all 67 intra - modes are allowed. If the corresponding width and height are at least 4 samples long, PDPC is also applied. Additionally, the conditions for intra - interpolation filter selection no longer exist, and in ISP mode, the Cubic (DCT - IF) filter is always applied for fractional - position interpolation. 2.1.1.8. Matrix - weighted Intra - prediction (MIP) The Matrix Weighted Intra Prediction (MIP) method is a newly added intra prediction technique in VVC. To predict the samples of a rectangular block with width W and height H, the Matrix Weighted Intra Prediction (MIP) takes as input a row of H reconstructed neighboring boundary samples to the left of the block and a row of W reconstructed neighboring boundary samples above the block. If the reconstructed samples are not available, they are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation, as Figure 10 shown. · Average neighboring samples Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, according to a predefined rule depending on the block size, by averaging the neighboring boundary samples, the input boundary bdry top and the input boundary bdry left are reduced to smaller boundaries and boundary Then, the two reduced boundaries and boundary are spliced into the reduced boundary vector bdry red , so that for a block of shape 4×4, its size is 4, and for all other shaped blocks, its size is 8. If the mode refers to the MIP mode, this splicing is defined as follows: · Matrix multiplication Taking the averaged samples as input, matrix-vector multiplication is performed, followed by adding an offset. The result is a reduced prediction signal for a subsampled set of samples in the original block. The reduced prediction signal pred red , is generated from the reduced input vector bdry red The reduced prediction signal pred red, is a signal for a downsampled block with width W red and height H red Here, W red and H red are defined as: By calculating the matrix-vector product and adding an offset, the reduced prediction signal pred red is calculated: pred red = A·bdry red + b. Here, if W = H = 4, A is a matrix with W red ·H red rows and 4 columns, and in all other cases 8 columns A is a matrix with W red·H red matrix of 8 rows and 8 columns. b is a vector of size W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, S 2. in. Define the index idx = idx(W, H) as follows: Here, each coefficient of the matrix A is represented with 8-bit precision. The set S0 consists of 16 matrices i ∈ {0, …, 15} and 16 offset vectors i ∈ {0, …, 16}, each matrix having 16 rows and 4 columns, and each offset vector having a size of 16. The matrices and offset vectors of this set are used for blocks of size 4×4. The set S1 consists of 8 matrices i ∈ {0, …, 7} and 8 offset vectors i ∈ {0, …, 7}, each matrix having 16 rows and 8 columns, and each offset vector having a size of 16. The set S2 consists of 6 matrices i ∈ {0, …, 5} and 6 offset vectors i ∈ {0, …, 5}, each matrix having 64 rows and 8 columns, and each offset vector having a size of 64. · Interpolation The predicted signal at the remaining positions is generated by linear interpolation from the predicted signals regarding the subsampled sets, which is single-step linear interpolation in each direction. Regardless of the block shape or block size, the interpolation is first performed in the horizontal direction and then in the vertical direction. · Signaling of the MIP mode and coordination with other coding tools For each coding unit (CU) in the intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transpose flag (isTransposed) determining whether the mode is transposed, and the MIP mode Id (modeId) determining which matrix is to be used for a given MIP mode are derived as follows: isTransposed = predModeIntra & 1 modeId = predModeIntra >> 1 (2 - 6) The MIP coding mode is coordinated with other coding tools by considering the following aspects: – Enable LFNST for MIP on large blocks. Here, the LFNST transform in the planar mode is used. – The reference sample derivation for MIP is performed in exactly the same way as for conventional intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used, instead of the downsampled reference samples. – Clipping is performed before upsampling, rather than after upsampling. – MIP is allowed up to 64x64 regardless of the maximum transform size. – The number of MIP modes is 32 for sizeId = 0, 16 for sizeId = 1, and 12 for sizeId = 2. 2.1.1.9. Spatial GPM (SGPM) In spatial GPM, a candidate list including split partitions and two intra prediction modes is established. The MPM of up to 11 intra prediction modes is used to form combinations, and the length of the candidate list is set to be equal to 16. The selected candidate index is signaled. List usage Figure 11 The shown template is reordered. The GPM mixing process is not used in the template, and the SAD between the prediction and reconstruction of the template is used for sorting. Figure 12 The GPM template is shown. Figure 13 The GPM split boundary is shown. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in inter-frame GPM. The following items are considered: ● Spatial GPM split mode: 26 predefined modes Adaptive derivation algorithm based on the ratio of horizontal gradient and vertical gradient ● Intra prediction mode selection: IPM lists with and without TIMD: For each split mode, the IPM list is derived for each part using intra-inter GPM list derivation. The IPM list size is 3. In the list, the TIMD-derived mode is replaced by 2 derived modes with horizontal and vertical orientations (using the top or left template), or the TIMD-derived mode is excluded. MPM list: A unified MPM list (up to 11 elements) is used for all split modes. ● Template size (left and above): 1 or 4 ● Extended block size: The spatial GPM is extended to be further applied to 4x8, 8x4, 4x16, and 16x4 blocks, which can be described as 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32. ● Adaptive blending: Adaptive blending is tested for the spatial GPM, where the blending depth τ is derived as follows: ■ If min(width, height) == 4, then select 1 / 2τ ■ Otherwise, if min(width, height) == 8, then select τ ■ Otherwise, if min(width, height) == 16, then select 2τ ■ Otherwise, if min(width, height) == 32, then select 4τ ■ Otherwise, select 8τ. 2.1.2. Inter - frame prediction For each inter - frame prediction CU, the motion parameters consist of a motion vector, a reference picture index, and a reference picture list index, as well as additional information required for the new decoding features of VVC for inter - frame prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is encoded / decoded using the skip mode, the CU is associated with a PU, and there are no significant residual coefficients, no encoded / decoded motion vector differences, or reference picture indices. The Merge mode is defined, where the motion parameters for the current CU are obtained from neighboring CUs. The Merge mode includes spatial candidates and temporal candidates, as well as additional scheduling introduced in VVC. The Merge mode can be applied to any inter - frame prediction CU, not just for the skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vectors, the corresponding reference picture indices, and the reference picture list indices for each reference picture list, as well as other required information, are explicitly signaled for each CU. In addition to the inter - frame encoding / decoding features in HEVC, VVC also includes many new and improved inter - frame prediction encoding / decoding tools, listed as follows: – Extended Merge prediction – Merge mode with MVD (MMVD) – Symmetric MVD (SMVD) signaling – Affine motion compensation prediction – Sub - block - based temporal motion vector prediction (SbTMVP) – Adaptive motion vector resolution (AMVR) – Motion field storage: 1 / 16 th Luminance sample MV storage and 8x8 motion field compression – Bi - directional prediction with CU - level weights (BCW) – Bidirectional optical flow (BDOF) – Decoder - side motion vector refinement (DMVR) – Geometric partition mode (GPM) – Combined inter - intra prediction (CIIP). The following text provides details of these inter - frame prediction methods specified in VVC. 2.1.2.1. Extended Merge prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVPs from spatially neighboring CUs 2) Temporal MVPs from co - located CUs 3) History - based MVPs from the FIFO table 4) Pair - wise average MVPs 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU encoding / decoding in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit of the Merge index is encoded using context, and bypass coding is used for the other binary bits. The derivation process of each type of Merge candidate is provided in this session. As done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size region. 2.1.2.1.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the positions Figure 14 shown, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Only when one or more of the CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because the CU belongs to another stripe or slice) or are intra - coded, is position B2 considered. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs Figure 15 linked by arrows in are considered, and a candidate is added to the list only if the corresponding candidates used for the redundancy check do not have the same motion information. 2.1.2.1.2. Temporal Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vectors are derived based on the co-located CUs belonging to the co-located reference pictures. The reference picture list to be used for the derivation of the co-located CUs is signaled explicitly in the slice header. The scaled motion vectors for the temporal Merge candidates are obtained as shown by the dashed lines in Figure 16 , which are scaled from the motion vectors of the co-located CUs using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. As shown in Figure 17 , the position for the temporal candidate is selected between candidate C0 and candidate C1. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate. 2.1.2.1.3. History-based Merge Candidate Derivation The history-based MVP (HMVP) Merge candidates are added to the Merge list after the spatial MVP and TMVP. In this method, the motion information of the previously decoded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). As long as there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that at most 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, the constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if there is the same HMVP in the table. If the same HMVP is found, the same HMVP is deleted from the table, and all subsequent HMVP candidates are moved forward. The HMVP candidates can be used in the Merge candidate list construction process. The last few HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. The redundancy check regarding the HMVP candidates is applied to the spatial Merge candidates or the temporal Merge candidates. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates used for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list construction process from HMVP is aborted. 2.1.2.1.4. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the Merge indices of the Merge candidate list. The averaged motion vectors are calculated separately for each reference list. If two motion vectors are both available in a list, the two motion vectors are averaged even if they point to different reference pictures; if only one motion vector is available, that motion vector is directly used; if no motion vector is available, this list is kept invalid. When the Merge list is not full after pairwise average Merge candidates are added, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached. 2.1.2.2. Merge estimation region The Merge estimation region (MER) allows for independent derivation of the Merge candidate list for CUs within the same Merge estimation region (MER). Candidate blocks within the same MER as the current CU are not included for the generation of the current CU's Merge candidate list. Additionally, the update process for the history-based motion vector predictor candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set by log2_parallel_merge_level_minus2. 2.1.2.3. Merge mode with MVD (MMVD) In addition to the Merge mode where implicitly derived motion information is directly used for generating prediction samples of the current CU, the VVC also introduces the Merge mode with motion vector difference (MMVD). The MMVD flag is signaled immediately after the skip flag and the Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a Merge candidate is selected, the Merge candidate is further refined by the MVD information signaled. The further information includes the Merge candidate flag, an index specifying the motion amplitude, and an index indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected as the MV basis. The Merge candidate flag is signaled to specify which one is used. The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As Figure 18 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 5. Table 5 – Relationship between distance index and predefined offset The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 6. It should be noted that the meaning of the MVD sign can vary according to the information of the starting MV. When the starting MV is a non-predicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the signs in Table 6 specify the signs of the MV offsets added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), the signs in Table 6 specify the signs of the MV offsets added to the list0 MV component of the starting MV, and the sign of the list1 MV has the opposite value. Table 6 – Signs of MV offsets specified by the direction index 2.1.2.4. Bi-directional prediction with CU-level weights (BCW) In HEVC, the bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred= ((8 - w)*P0 + w*P1 + 4) >> 3 (2 - 7) In weighted average bi - prediction, 5 weights are allowed, w ∈ {-2, 3, 4, 5, 10}. For each bi - predicted CU, the weight w is determined in one of two ways: 1) For non - Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used. - At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low - latency picture, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion vector precisions. - When combined with affine, affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode. - When the two reference pictures in bi - prediction are the same, only unequal weights are conditionally checked. - When certain conditions are met, unequal weights are not searched, depending on the POC distance between the current picture and its reference picture, the coding - decoding QP, and the temporal level. The BCW weight index is decoded using one context - decoded bit followed by a bypass - decoded bit. The first context - decoded bit indicates whether equal weights are used; if unequal weights are used, additional bits are signaled using bypass decoding to indicate which unequal weight is used. Weighted Prediction (WP) is a coding and decoding tool supported by the H.264 / AVC and HEVC standards for efficiently coding and decoding video content with fading. Support for WP has also been added to the VVC standard. WP allows for the weighted parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the corresponding weights and offsets for the (multiple) reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded and decoded using the CIIP mode, the BCW index of the current CU is set to 2, for example, equal weights. 2.1.2.5. Bidirectional Optical Flow (BDOF) The VVC includes the Bidirectional Optical Flow (BDOF) tool. BDOF (previously known as BIO) was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met for the CU: - The CU is coded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order - The distance from the two reference pictures to the current picture (i.e., the POC difference) is the same - Both of the two reference pictures are short-term reference pictures. - The CU is not coded and decoded using the affine mode or the ATMVP Merge mode - The CU has more than 64 luma samples - Both the CU height and the CU width are greater than or equal to 8 luma samples - The BCW weight index indicates equal weights - WP is not enabled for the current CU - The CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bidirectional predicted sample values in the 4x4 sub-block. The following steps are applied during the BDOF process. First, the horizontal gradients and vertical gradients k = 0, 1, are calculated by directly computing the difference between two neighboring samples, i.e., where I (k) (i, j) is the sample value at the coordinates (i, j) of the predicted signal in the list k, k = 0, 1, and shift1 is calculated based on the luminance bit depth (bitDepth) as shift1 = max(6, bitDepth - 6). Then, the autocorrelations and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as where where Ω is a 6×6 window around the 4×4 sub-block, and n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then the motion refinement (v x , v y ) is derived using the following cross-correlation and autocorrelation terms:[[]] where th′ BIO = 2 max(5,BD-7) . is the floor function, and n S2 = 12. Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4×4 sub-block:[[]] Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional predicted samples as follows:[[]] predBDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + ο offset ) >> shift (2 - 13) These values are selected such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bitwidth of the intermediate parameters in the BDOF process is kept within 32 bits. To derive the gradient values, some predicted samples I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 19 shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating predicted samples outside the boundary, the predicted samples in the extended region (white positions) are generated by directly adopting the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted samples inside the CU (gray positions). These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample and gradient value outside the CU boundary are needed, the sample and gradient value are filled (i.e., repeated) from its nearest neighbor. When the width and / or height of the CU is greater than 16 luma samples, it is divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries are treated as CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 predicted samples and the L1 predicted samples is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8 * W * (H >> 1)), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated between the initial L0 predicted samples and the L1 predicted samples during the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, the BDOF is also disabled. When the CU is encoded / decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.2.6. Symmetric MVD Encoding / Decoding In VVC, in addition to the normal unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional prediction MVD signaling is also applied. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not signaled but derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is 1, then BiDirPredFlag is set equal to 0. - Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a pair of forward and backward reference pictures or a pair of backward and forward reference pictures, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise BiDirPredFlag is set to 0. 2) At the CU level, if the CU is coded / decoded by bidirectional prediction and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. Figure 20 It is a schematic diagram for the symmetric MVD mode. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes for list 0 and list 1 are respectively set equal to a pair of reference pictures. MVD1 is set equal to (-MVD0). The final motion vector is as shown in the following formula. In the encoder, the symmetric MVD motion estimation starts from the initial MV evaluation. A set of initial MV candidates, including the MVs obtained from unidirectional prediction search, the MVs obtained from bidirectional prediction search, and the MVs from the AMVP list. An MV with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.2.7. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the MVs in the Merge mode, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In the bidirectional prediction operation, the refined MVs are searched around the initial MVs in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and reference picture list L1. As Figure 21As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi - directional prediction signal. In VVC, DMVR can be applied to CUs encoded and decoded using the following patterns and features: - CU - level Merge mode with bi - directional prediction MVs - With respect to the current picture, one reference picture is past and the other reference picture is future - The distances from the two reference pictures to the current picture (i.e., POC differences) are the same - Both reference pictures are short - term reference pictures - The CU has more than 64 luma samples - The CU height and CU width are both greater than or equal to 8 luma samples - The BCW weight index indicates equal weights - WP is not enabled for the current block - The CIIP mode is not used for the current block. The refined MVs derived by the DMVR process are used to generate inter - prediction samples and are also used for temporal motion vector prediction for future picture encoding and decoding. While the original MVs are used for the de - blocking process and are also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following entries. 2.1.2.7.1. Search Scheme In DVMR, the search points are around the initial MV, and the MV offsets follow the MV difference mirroring rule. In other words, any point checked by DMVR (represented by the candidate MV pair (MV0, MV1)) follows the following two equations: MV0′ = MV0+MV_offset (2 - 15) MV1′ = MV1 - MV_offset (2 - 16) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two whole luma samples away from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage. A 25-point full search is applied to the integer sample point offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample point stage of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample point offset search stage. To reduce the penalty of the uncertainty of DMVR refinement, it is proposed to bias towards the original MV during the DMVR process. The SAD between the reference blocks referred to by the initial MV candidates reduces the SAD value by 1 / 4. The integer sample point search is followed by fractional sample point refinement. To save computational complexity, the fractional sample point refinement is derived by using the parametric error surface equation, rather than by an additional search with SAD comparison. The fractional sample point refinement is conditionally invoked based on the output of the integer sample point search stage. When the integer sample point search stage is terminated at the center with the minimum SAD in the first iteration or the second iteration search, the fractional sample point refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C(2 - 17) where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By using the cost values of the five search points to solve the above equation, (x min , y min ) is calculated as: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (2 - 18) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) (2 - 19). The values of x min and y min are automatically constrained between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to the half-peak offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain the refined incremental MV with sub-pixel accuracy. 2.1.2.7.2. Bilinear Interpolation and Sample Padding In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, with a 2-sample search range, compared to the normal motion compensation process, DVMR does not access more reference samples. After the refined MV is obtained using the DMVR search process, the normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples than the normal MC process, samples (which are not required by the interpolation process based on the original MV but are required by the interpolation process based on the refined MV) will be padded from these available samples. 2.1.2.7.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16. 2.1.2.8. Combined Inter / Intra Prediction (CIIP) In VVC, when a CU is encoded / decoded in the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width times the CU height is equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. Figure 22 The top and left neighboring blocks used in CIIP weight derivation are shown. As its name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in the CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived following the regular intra prediction process with the planar mode. Then, the intra and inter prediction signals are combined using a weighted average, where the weight values are calculated according to the encoding / decoding modes of the top and left neighboring blocks, as follows: - If the top neighbor is available and is intra-encoded, set isIntraTop to 1, otherwise set isIntraTop to 0; - If the left neighbor is available and is intra-frame coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; - If (isIntraLeft + isIntraTop) equals 2, set wt to 3; - Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2; - Otherwise, set wt to 1. The CIIP prediction is formed as follows: P CIIP = ((4 - wt)*P inter + wt*P intra + 2) >> 2 (2 - 20). 2.1.2.9. Multiple Hypothesis Prediction (MHP) At most two additional prediction factors are signaled over the inter-frame AMVP mode, the regular Merge mode, and the MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 = (1 - α n+1 )p n + α n+1 h n+1 The weighting factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For the inter-frame AMVP mode, MHP is applied only when non-uniform weights in the BCW are selected in the bi-prediction mode. 2.1.2.10. Overlapped Block Motion Compensation (OBMC) When OBMC is applied, the top and left boundary pixels of the CU are refined using the motion information of neighboring blocks with weighted prediction. The conditions for not applying OBMC are as follows: · When OBMC is disabled at the SPS level · When the current block has an intra-frame mode or an IBC mode · When the current block applies LIC · When the current luma block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blend to the top, left, bottom, and right sub-block boundary pixels using the motion information of neighboring sub-blocks. Sub-block boundary OBMC is enabled for sub-block based coding tools: · Affine AMVP mode; · Affine Merge mode and sub-block based temporal motion vector prediction (SbTMVP); · Sub-block based bilateral matching. 2.1.2.11. Local Illumination Compensation (LIC) LIC is an inter prediction technique that models the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template and the reference block template. The parameters of this function can be represented by a scaling α and an offset β, which form a linear equation, i.e., α*p[x] + β, to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When looped motion compensation is enabled, the MV must be clipped using the loop offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead for α and β is required except that the LIC flag is signaled for the AMVP mode to indicate the use of LIC. The local illumination compensation proposed in JVET-O0066 is used for unidirectional predicted inter CUs with the following modifications. · Intra neighboring samples can be used for LIC parameter derivation; · LIC is disabled for blocks with less than 32 luma samples; · For non-sub-block modes and affine modes, LIC parameter derivation is performed based on the modulo block samples corresponding to the current CU rather than based on the partial modulo block samples corresponding to the top-left first 16x16 unit; · Samples of the reference block template are generated by using MC with the block MV without rounding the samples to integer pixel precision. 2.1.2.12. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode for inter prediction is supported. The geometric partitioning mode is signaled as a Merge mode using a CU level flag, where other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h = 2 m ×2 n (where m,n ∈ {3…6} excluding 8x64 and 64x8), a total of 64 are supported by the partitioning geometric partitioning mode. When this mode is used, the CU is divided into two parts by a geometrically positioned line ( Figure 23)。The position of the dividing line is mathematically derived from the angular parameter and offset parameter of a specific division. Each part of the geometric division in the CU is inter-frame predicted using its own motion; only unidirectional prediction is allowed for each division, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure the same as traditional bidirectional prediction, and only two motion compensation predictions are required for each CU. If the geometric division mode is used for the current CU, the geometric division index indicating the division mode (angle and offset) of the geometric division and two Merge indices (one index for each division) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric division, the sample values along the geometric division edge are adjusted using a blending process with adaptive weights. This is the prediction signal for the entire CU, and the transform and quantization processes will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric division mode is stored. 2.1.2.12.1. Unidirectional Prediction Candidate List Construction The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X is equal to the parity of n) is used as the nth unidirectional prediction motion vector for the geometric division mode. These motion vectors are marked with "x" in Figure 24 . In the case where the corresponding LX motion vector of the nth extended Merge candidate does not exist, instead, the L(1 - X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric division mode. 2.1.2.12.2. Blending along the Geometric Division Edge Figure 25 Shows an exemplary generation of the blending weight w0 using the geometric division mode. After predicting each part using its own motion in each part of the geometric division, blending is applied to the two prediction signals to derive the samples around the geometric division edge. The blending weight for each position of the CU is derived based on the distance between the single position and the division edge. The distance from the position (x, y) to the division edge is derived as: where i, j are the indices for the angle and offset of the geometric division, which depend on the geometric division index signaled. ρ x,j and ρ y,j 's signs depend on the angle index i. The weight of each part of the geometric segmentation is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) (2-25) w1(x,y) = 1 - w0(x,y) (2-27) partIdx depends on the angle index i. An example of the weight w0 is shown below. 2.1.2.12.3. Motion Field Storage for Geometric Segmentation Mode Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded / decoded in the geometric segmentation mode. The type of motion vector stored for each individual position in the motion field is determined as:[[]] sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 - partIdx) : partIdx) (2-28) where motionIdx is equal to d(4x + 2, 4y + 2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType is equal to 2, the combination Mv from Mv0 and Mv2 is stored. The combination Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-predictive motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, then only the uni-predictive motion Mv2 is stored. 2.1.2.12.4. GPM with Inter-Frame and Intra-Frame Prediction (GPM Inter-Frame - Intra-Frame) Using GPM inter - intra, in addition to the Merge candidates for each non - rectangular partition region in the CU to which GPM is applied, predefined intra - prediction modes for geometric split lines can also be selected. In the proposed method, whether it is an intra - prediction mode or an inter - prediction mode is determined for each GPM - separated region using a flag from the encoder. In the inter - prediction mode, the unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, in the intra - prediction mode, the unidirectional prediction signal is generated from neighboring pixels for the intra - prediction mode specified by an index from the encoder. The variation of possible intra - prediction modes is restricted by the geometry. Finally, the two unidirectional prediction signals are mixed in the same way as in ordinary GPM. 2.1.3. Screen Content Coding and Decoding Tools 2.1.3.1. Intra - Block Copy (IBC) Intra - Block Copy (IBC) is a tool adopted in the HEVC extension for SCC. As is well known, IBC significantly improves the coding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block - level coding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current picture. The luminance block vectors of the CUs coded and decoded by IBC are in integer precision. The chrominance block vectors are also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1 - pixel motion vector precision and 4 - pixel motion vector precision. The CUs coded and decoded by IBC are regarded as a third prediction mode different from the intra - prediction mode or the inter - prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples. On the encoder side, hash - based motion estimation is performed for IBC. The encoder performs RD checking for blocks whose width or height is no greater than 16 luminance samples. For non - Merge modes, the block vector search is first performed using hash - based search. If the hash search does not return a valid candidate, a local search based on block matching will be performed. In the hash - based search, the hash - key matching (32 - bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash - key calculation for each position in the current picture is based on 4x4 sub - blocks. For a larger - sized current block, when all hash keys of all 4×4 sub - blocks match the hash keys in the corresponding reference positions, the hash key is determined to match the hash key of the reference block. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block - vector cost of each matching reference is calculated, and the one with the minimum cost is selected. In block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using flags, and the IBC mode can be signaled as the IBC AMVP mode or the IBC skip / Merge mode as follows: – IBC skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC decoded blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. The block vector prediction method uses two candidates as predictors, one from the left neighboring block and one from the upper neighboring block (if IBC decoded). When either neighboring block is not available, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.1.3.1.1. IBC reference region To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstructed part of a predefined region that includes the region of the current CTU and some regions of the left CTU. Figure 26 The reference region of the IBC mode is shown, where each block represents a 64x64 luma sample unit. Depending on the position of the current decoded CU position within the current CTU, the following applies: – If the current block falls within the upper left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, the reference samples in the lower right 64x64 block of the left CTU can be used with the CPR mode. The current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block of the left CTU and the upper right 64x64 block of the left CTU. – If the current block falls within the upper right 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (0,64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also reference the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls within the lower - left 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, if the luminance position (0, 64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to reference the reference samples in the upper - right 64x64 block and the lower - right 64x64 block of the left - hand CTU. Otherwise, the current block can also use the CPR mode to reference the reference samples in the lower - right 64x64 block of the left - hand CTU. – If the current block falls within the lower - right 64x64 block of the current CTU, then only the samples already reconstructed in the current CTU can be used in the CPR mode. This restriction allows the IBC mode to be implemented using local on - chip memory for hardware implementation. 2.1.3.1.2. Interaction between IBC and other coding tools The interaction between the IBC mode and other inter - frame coding tools in VVC (such as paired Merge candidates, history - based motion vector predictors (HMVP), intra / inter - frame joint prediction mode (CIIP), Merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM)) is as follows: – IBC can be used together with paired Merge candidates and HMVP. New paired IBC Merge candidates can be generated by averaging two IBC Merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference. – IBC cannot be combined with the following inter - frame tools: affine motion, CIIP, MMVD, and GPM. – When DUAL_TREE partitioning is used, IBC is not allowed for chrominance coding blocks. Different from the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation process of the motion vector for the IBC mode excludes all neighboring blocks in the inter - frame mode, and vice versa. The following IBC design aspects are applied: – IBC shares the same process as regular MV Merge, including paired Merge candidates and history - based motion predictors, but TMVP and zero vectors are not allowed because they are not valid for the IBC mode. – Separate HMVP caches (5 candidates each) are used for traditional MVs and IBC. – Block vector constraints are implemented in the form of bit - stream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bit - stream, and if a Merge candidate is invalid (out of range or 0), then the Merge must not be used. Such bit - stream consistency constraints are described based on a virtual cache, as described below. – For deblocking, IBC is processed as an inter mode. – If the current block is coded / decoded using an IBC prediction mode, then AMVR does not use quarter pixels; instead, AMVR is signaled to indicate only whether the MV is an inter pixel or 4 integer pixels. – The number of IBC Merge candidates can be signaled separately in the slice header from the number of regular candidates, sub-block candidates, and geometric Merge candidates. The virtual buffer concept is used to describe the allowable reference regions and valid block vectors for IBC prediction modes. Representing the CTU size as ctbSize, the virtual buffer ibcBuf has width wIbcBuf = 128x128 / ctbSize and height hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of the VPDU is min(ctbSize, 64) in each dimension, Wv = min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows: - At the start of decoding each CTU row, the entire ibcBuf is flushed with the invalid value -1. - At the start of decoding the VPDU(xVPDU, yVPDU) relative to the top-left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU % wIbcBuf, …, xVPDU % wIbcBuf + Wv - 1; y = yVPDU % ctbSize, …, yVPDU % ctbSize + Wv - 1. - After decoding a CU containing (x, y) relative to the top-left corner of the picture, set ibcBuf[x % wIbcBuf][y % ctbSize] = recSample[x][y] For a block covering the coordinates (x, y), the block vector is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, the block vector is invalid: ibcBuf[(x + bv[0]) % wIbcBuf][(y + bv[1]) % ctbSize] must not be equal to -1. 2.1.3.2. Block Differential Pulse Coding Modulation (BDPCM) VVC supports Block Differential Pulse Coding Modulation (BDPCM) for screen content encoding and decoding. At the sequence level, the BDPCM enable flag is signaled in the SPS; this flag is signaled only if the transform skip mode (described in the next section) is enabled in the SPS. When BDPCM is enabled, if the CU size is less than or equal to MaxTsSize by MaxTsSize in terms of luma samples, and if the CU is intra-coded, the flag is signaled at the CU level, where MaxTsSize is the maximum block size allowed for the transform skip mode. This flag indicates whether regular intra-coding is used or BDPCM is used. If BDPCM is used, the BDPCM prediction direction flag is signaled to indicate whether the prediction is horizontal or vertical. Then, the block is predicted using the regular horizontal or vertical intra-prediction process with unfiltered reference samples. The residuals are quantized, and the difference between each quantized residual and its predictor (i.e., the previously decoded residual at the horizontal or vertical (depending on the BDPCM prediction direction) neighboring position) is coded. For a block of size M (height) × N (width), let r i,j , 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 be the prediction residuals. Let Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 denote the quantized version of the residual r i,j . BDPCM is applied to the quantized residual values, resulting in a modified M × N array with elements where is predicted from its neighboring quantized residual values. For the vertical BDPCM prediction mode, for 0 ≤ j ≤ (N - 1), the following is used to derive For the horizontal BDPCM prediction mode, for 0 ≤ i ≤ (M - 1), the following is used to derive On the decoder side, the above process is reversed to compute Q(r i,j ), 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1, as follows: The inverse-quantized residual Q -1 (Q(r i,j )) is added to the intra-block prediction value to produce the reconstructed sample value. The predicted quantized residual values The same residual coding process as that in the transform skip mode residual coding is sent to the decoder. For lossless coding, if the slice_ts_residual_coding_disabled_flag is set to 1, the quantized residual values are sent to the decoder using the regular transform residual coding. In terms of the MPM mode for future intra mode coding, if the BDPCM prediction directions are horizontal or vertical respectively, the horizontal or vertical prediction modes are stored for the CU coded by BDPCM. For deblocking, if the blocks on both sides of the block boundary are coded by BDPCM, the specific block boundary is not deblocked. 2.1.3.3. Residual Coding for Transform Skip Mode VVC allows the transform skip mode to be used for luma blocks with a maximum size of MaxTsSize multiplied by MaxTsSize, where the value of MaxTsSize is signaled in the PPS and can be at most 32. When a CU is coded in the transform skip mode, its prediction residual is quantized and coded using the transform skip residual coding process. This process is modified from the transform coefficient coding process. In the transform skip mode, the residual of the TU is also coded in non - overlapping sub - blocks of size 4x4. For better coding efficiency, some modifications are made to customize the residual coding process towards the characteristics of the residual signal. The following summarizes the differences between the transform skip residual coding and the regular transform residual coding: – The forward scan order is applied to scan the sub - blocks within the transform block and the positions within the sub - blocks; – There is no signaling of the last (x, y) position; – When all previous flags are equal to 0, the coded_sub_block_flag is coded for each sub - block except the last one; – The sig_coeff_flag context modeling uses a simplified template, and the context model of sig_coeff_flag depends on the top and left neighboring values; – The context model of the abs_level_gt1 flag also depends on the left and top sig_coeff_flag values; – The par_level_flag uses only one context model; – Additional flags greater than 3, 5, 7, 9 are signaled to indicate the coefficient levels, one context for each flag; – The use of Rice parameter derivation with a fixed order = 1 for the binarization of the residual values; The context model of the sign flag is determined based on the neighboring values to the left and above, and the sign flag is parsed after sig_coeff_flag to keep all context-encoded and decoded bits together. For each sub-block, if coded_subblock_flag equals 1 (i.e., there is at least one non-zero quantized residual in the sub-block), the encoding and decoding of the quantized residual levels are performed in three scans (see Figure 27 ): - First scan pass: The significance flag (sig_coeff_flag), sign flag (coeff_sign_flag), flag indicating that the absolute level is greater than 1 (abs_level_gtx_flag[0]), and parity (par_level_flag) are decoded. For a given scan position, if sig_coeff_flag equals 1, then coeff_sign_flag is decoded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[0] equals 1, then par_level_flag is additionally decoded to specify the parity of the absolute level. - Scan pass greater than x: For each scan position where the absolute level is greater than 1, up to four abs_level_gtx_flag[i] (for i = 1...4) are decoded to indicate whether the absolute level at the given position is greater than 3, 5, 7, or 9, respectively. - Remaining scan passes: The remainder of the absolute level abs_remainder is decoded in bypass mode. The remainder of the absolute level is binarized using a fixed Rice parameter value of 1. The bits in scan pass #1 and scan pass #2 (the first scan pass and scan passes greater than x) are context coded until the maximum number of context-coded bits in the TU is exhausted. The maximum number of context-coded bits in the residual block is limited to 1.75 * block_width * block_height, or equivalently, limited to 1.75 context-coded bits per sample position. The bits in the last scan pass (the remaining scan passes) are bypass coded. The variable RemCcbs is first set to the maximum number of context-coded bits for the block and is decremented by 1 each time a context-coded bit is coded. When RemCcbs is greater than or equal to 4, the syntax elements in the first coding pass, which include sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag, are coded using the context-coded bits. If RemCcbs becomes less than 4 during the first coding pass, the remaining coefficients that have not been coded in the first pass are coded in the remaining scan passes (pass #3). After the first pass coding is completed, if RemCcbs is greater than or equal to 4, the syntax elements in the second coding pass, which include abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag, are coded using the context-coded bits. If RemCcbs becomes less than 4 during the second coding pass, the remaining coefficients that have not been coded in the second pass are coded in the remaining scan passes (pass #3). Figure 27 The transform skip residual coding process is shown. The asterisk marks the position where the context-coded bits are exhausted, at which point all remaining bits are coded using bypass coding. In addition, for blocks not coded in the BDPCM mode, a level mapping mechanism is applied to transform skip residual coding until the maximum number of context-coded bits is reached. The level mapping uses the top and left neighboring coefficient levels to predict the current coefficient level in order to reduce the signaling cost. For a given residual position, let absCoeff denote the absolute coefficient level before mapping and absCoeffMod denote the coefficient level after mapping. Let X0 denote the absolute coefficient level of the left neighboring position and let X1 denote the absolute coefficient level of the upper neighboring position. The level mapping is performed as follows: pred = max(X0, X1); if(absCoeff == pred) absCoeffMod = 1; else absCoeffMod = (absCoeff < pred)? absCoeff + 1 : absCoeff; Then, the absCoeffMod value is encoded and decoded as described above. After all the binary bits for context encoding and decoding are exhausted, the level mapping is disabled for all the remaining scan positions in the current block. 2.1.3.4. Palette Mode In VVC, the palette mode is used for screen content encoding and decoding in all chroma formats (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome) supported in the 4:4:4 profile. When the palette mode is enabled, if the CU size is less than or equal to 64x64 and the number of samples in the CU is greater than 16, a flag is transmitted at the CU level to indicate whether the palette mode is used. Considering that the encoding and decoding gain introduced by applying the palette mode on small CUs is not significant and it brings additional complexity to small blocks, the palette mode is disabled for CUs with less than or equal to 16 samples. The encoded and decoded coding unit (CU) by the palette is regarded as a prediction mode different from the intra prediction mode, the inter prediction mode, and the intra block copy (IBC) mode. If the palette mode is utilized, the sample values in the CU are represented by a set of representative color values. This set is called the palette. For positions where the sample values are close to the palette colors, the palette index is signaled. The out-of-palette samples can also be specified by signaling the escape symbol. For the samples within the CU encoded using the escape symbol, their component values are directly signaled using the (possibly) quantized component values. This is as Figure 28 shown. The quantized escape symbol is binary-coded using the fifth-order exponential-Golomb binary process (EG5). For the encoding and decoding of the palette, the palette predictor is maintained. For non-wavefront cases, the palette predictor is initialized to 0 at the start of each slice. For WPP cases, the palette predictor at the start of each CTU row is initialized to the predictor derived from the first CTU in the previous CTU row, such that the initialization scheme between the palette predictor and CABAC synchronization is unified. For each entry in the palette predictor, a reuse flag is signaled to indicate whether it is part of the current palette in the CU. The reuse flag is sent using zero run-length coding. Thereafter, the number of new palette entries and the component values for the new palette entries are signaled. After encoding the palette-coded CU, the palette predictor will be updated using the current palette, and the entries from the previous palette predictor that were not reused in the current palette will be added to the end of the new palette predictor until the maximum allowed size is reached. An escape flag is signaled for each CU to indicate whether an escape symbol exists in the current CU. If an escape symbol exists, the palette table is incremented by 1, and the last index is assigned to the escape symbol. In a manner similar to the coefficient groups (CGs) used in transform coefficient encoding, the CUs encoded using the palette mode are divided into multiple row-based coefficient groups, each coefficient group consisting of m samples (i.e., m = 16), where, for each CG in sequence, the index run, the palette index value, and the quantized color for the escape mode are encoded / parsed. Similar to HEVC, a horizontal or vertical traversal scan can be applied to scan the samples, as Figure 29 shown. The encoding order for palette run encoding / decoding in each slice is as follows: For each sample position, a context-encoded binary bit run_copy_flag = 0 is signaled to indicate whether the pixel has the same pattern as the previous sample position, i.e., whether both the previously scanned sample and the current sample are of run type COPY_ABOVE, or whether both the previously scanned sample and the current sample are of run type INDEX and the same index value. Otherwise, run_copy_flag = 1 is signaled. If the current sample and the previous sample have different patterns, a context-encoded binary bit copy_above_palette_indices_flag is signaled to indicate the run type of the current sample, i.e., INDEX or COPY_ABOVE. Here, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type, as the INDEX mode is used by default. In the same way, if the previously parsed run type is COPY_ABOVE, the decoder does not have to parse the run type. After the palette run encoding / decoding of the samples in one encoding / decoding pass, the index values (for INDEX mode) and the quantized escape colors are grouped and encoded / decoded using CABAC bypass encoding / decoding in another encoding / decoding pass. This separation of context-encoded binary bits and bypass-encoded binary bits can improve the throughput within each row CG. For a slice with a dual-luma / chroma tree, the palette is applied separately to luma (Y component) and chroma (Cb component and Cr component), where luma palette entries contain only Y values, and chroma palette entries contain both Cb values and Cr values. For a slice with a single tree, the palette will be applied jointly to the Y, Cb, Cr components, i.e., each entry in the palette contains a Y value, a Cb value, and a Cr value, unless when the CU is encoded using a local dual tree, in which case the encoding of luma and chroma is processed separately. In this case, if the corresponding luma block or chroma block is encoded using the palette mode, their palettes are applied in a manner similar to the dual-tree case (this is related to non-4:4:4 encoding and will be further explained in 2.1.3.4.1). For a slice encoded using a dual tree, the maximum palette prediction factor size is 63, and the maximum palette table size for encoding the current CU is 31. For a slice encoded using a dual tree, for each of the luma palette and the chroma palette, the maximum prediction factor and the palette table size are halved, i.e., the maximum prediction factor size is 31, and the maximum table size is 15. For deblocking, the palette-encoded blocks on the sides of the block boundary are not deblocked. 2.1.3.4.1. Palette Mode for Non-4:4:4 Content The palette mode in VVC supports all chroma formats in a similar way to the palette mode in HEVC SCC. For non-4:4:4 content, the following customizations are applied: 1. When signaling an escape value for a given sample position, if the sample position has only a luma component and no chroma component due to chroma subsampling, only the luma escape value is signaled. This is the same as in HEVC SCC. 2. For local dual-tree blocks, the palette mode is applied to the block in the same way as it is applied to single-tree blocks, except for the following two exceptions: a. The process of palette predictor update is slightly modified as follows. Since a local dual-tree block contains only a luma (or chroma) component, the predictor update process uses the values signaled for the luma (or chroma) component and fills in the "missing" chroma (or luma) component by setting it to the default value (1<<(component bit depth - 1)). b. The maximum palette predictor size is kept at 63 (since the strip is coded / decoded using a single tree), but the maximum palette table size for luma / chroma blocks is kept at 15 (since the block is coded / decoded using a separate palette). 3. For the palette mode in monochromatic formats, the number of color components in the palette-coded block is set to 1 instead of 3. 2.1.3.4.2. Encoder Algorithm for Palette Mode On the encoder side, the following steps are used to generate the palette table for the current CU: 1. First, to derive the initial entries in the palette table for the current CU, simplified K-means clustering is applied. The palette table for the current CU is initialized as an empty table. For each sample position in the CU, the SAD between the sample and each palette table entry is calculated, and the minimum SAD among all palette table entries is obtained. If the minimum SAD is less than a predefined error limit (errorLimit), the current sample and the palette table entry with the minimum SAD are clustered together. Otherwise, a new palette table entry is created. The threshold errorLimit is QP-dependent and is retrieved from a lookup table containing 57 elements covering the entire QP range. After all samples in the current CU have been processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entry after the 31st entry is discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroid of each cluster from step 1 or using one of the palette colors in the palette predictor. The option with the lower rate-distortion cost is selected as the final color of the palette table. If a cluster has only a single sample point and the corresponding palette entry is not in the palette predictor, the corresponding sample point is converted to an escape symbol in the next stage. 3. The palette table generated in this way contains some new entries from the centroids of the clusters in step 1 and some entries from the palette predictor. Therefore, the table is reordered again so that all new entries (i.e., the centroids) are placed at the start of the table, followed by the entries from the palette predictor. Given the palette table of the current CU, the encoder selects the palette index for each sample point position in the CU. For each sample point position, the encoder checks the RD cost of all index values corresponding to the palette table entries and the index representing the escape symbol, and selects the index with the minimum RD cost using the following equation: RD cost = distortion × (isChroma? 0.8:1) + lambda × bits bypassed in coding / decoding (2 - 33) After deciding the index map of the current CU, each entry in the palette table is checked to see if it is used by at least one sample point position in the CU. Any unused palette entry will be deleted. After the index map of the current CU is decided, grid RD optimization is applied to find the best values of run_copy_flag and run type for each sample point position by comparing the RD costs of three options: the same as the previously scanned position, run type COPY_ABOVE, or run type INDEX. When calculating the SAD value, the sample values are scaled down to 8 bits, unless the CU is coded / decoded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. Additionally, in the case of lossless coding / decoding, the rate is only used in the above rate-distortion optimization step (since lossless coding / decoding does not produce distortion). 2.1.3.5. Adaptive Color Transform In the HEVC SCC extension, Adaptive Color Transform (ACT) is applied to reduce the redundancy between the three color components in the 4:4:4 chroma format. ACT is also adopted in the VVC standard to improve the coding / decoding efficiency of 4:4:4 chroma format coding / decoding. Similar to HEVC SCC, ACT performs a loop color space transform in the prediction residual domain by adaptively transforming the residual from the input color space to the YCgCo space. Figure 30The decoding flowchart in the case where ACT is applied is shown. By transmitting an ACT flag at the CU level through signaling, two color spaces are adaptively selected. When the flag is equal to 1, the residual of the CU is encoded and decoded in the YCgCo space; otherwise, the residual of the CU is encoded and decoded in the original color space. Additionally, similar to the HEVC ACT design, for inter-frame CUs and IBC CUs, ACT is only enabled when there is at least one non-zero coefficient in the CU. For intra-frame CUs, ACT is only enabled when the chrominance component selects the same intra-prediction mode (i.e., DM mode) as the luma component. 2.1.3.5.1. ACT Mode In the HEVC SCC extension, ACT supports both lossless encoding and decoding and lossy encoding and decoding based on the lossless flag (i.e., cu_transquant_bypass_flag). However, there is no flag signaled in the bitstream to indicate whether lossy encoding and decoding is applied or lossless encoding and decoding is applied. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown as follows. Since the YCgCo-R transform is not normalized. To compensate for the dynamic range change of the residual signal before and after the color transform, QP adjustments of (-5, 1, 3) are applied to the transform residuals of the Y, Cg, and Co components respectively. The adjusted quantization parameter only affects the quantization and inverse quantization of the residuals in the CU. For other encoding and decoding processes (such as deblocking), the original QP is still applied. Additionally, since the forward color transform and the inverse color transform require access to the residuals of all three components, the ACT mode is always disabled for separate tree partitioning and ISP modes where the prediction block sizes of different color components are different. When ACT is applied, the transform skip (TS) extended to encode and decode chrominance residuals and the block differential pulse code modulation (BDPCM) are also enabled. 2.1.3.5.2. ACT Fast Encoding Algorithm To avoid brute-force R-D search in both the original color space and the transformed color space, when ACT is enabled, the following fast encoding algorithm is applied in the VTM reference software to reduce the encoder complexity. – The order of the RD check for enabling / disabling ACT depends on the original color space of the input video. For RGB videos, the RD cost of the ACT mode is checked first; for YCbCr videos, the RD cost of the non-ACT mode is checked first. The RD cost of the second color space is only checked when there is at least one non-zero coefficient in the first color space. – When a CU is obtained through different splitting paths, the same ACT enable / disable decision is reused. Specifically, when the CU is coded / decoded for the first time, the selected color space for coding the residuals of the CU is stored. Then, when the same CU is obtained through another splitting path, the stored color space decision is directly reused instead of checking the RD cost of the two spaces. – The RD cost of the parent CU is used to determine whether to check the RD cost of the second color space for the current CU. For example, if the RD cost of the first color space for the parent CU is less than the RD cost of the second color space, then for the current CU, the second color space is not checked. – To reduce the number of coding / decoding modes to be tested, the selected coding / decoding modes are shared between the two color spaces. Specifically, for intra modes, the preselected intra mode candidates based on SATD-based intra mode selection are shared between the two color spaces. For inter modes and IBC modes, the block vector search or motion estimation is only performed once. The block vectors and motion vectors are shared by the two color spaces. 2.1.3.6. Intra Template Matching (IntraTMP) Intra Template Matching Prediction (IntraTM) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches in the reconstructed part of the current frame for the template that is most similar to the current template and uses the corresponding block as the prediction block. Then the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is generated by matching the L-shaped causal neighbor of the current block with Figure 31 another block in the predefined search region consisting of the following: R1: The current CTU R2: The top-left CTU R3: The upper CTU R4: The left CTU. SAD is used as the cost function. Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses the block corresponding to that template as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH where 'a' is a constant that controls the gain / complexity trade-off. In practice, 'a' is equal to 5. The intra-template matching tool is enabled for CUs with dimensions of width and height less than or equal to 64. The maximum CU size for intra-template matching is configurable. When DIMD is not used for the current CU, the intra-template matching prediction mode is signaled at the CU level via a dedicated flag. 2.1.3.6.1. Using the Block Vector Derived from IntraTMP for IBC The block vector (BV) derived from intra-template matching prediction (IntraTMP) is used for intra-block copy (IBC). In the IBC candidate list construction, the stored IntraTMP BV of neighboring blocks is used as a spatial BV candidate together with the IBC BV. 2.1.3.6.2. Direct Block Vector (DBV) Mode for Chrominance Prediction For the chrominance component, when the chrominance dual tree is activated in an intra slice, if one of the five luminance blocks is encoded / decoded using MODE_IBC, its block vector bvL is used and scaled to derive the chrominance block vector bvC. The scaling factor depends on the chrominance format sampling structure. Then, by using the position (xCb, yCb) of the current chrominance block and the bvC of this chrominance block, the corresponding offset position (xCb + bvC[0], yCb + bvC[1]) is determined, and block copy prediction is performed. Figure 32 Five positions in the reconstructed luminance samples are shown. Figure 33 The prediction process of the DBV mode is shown. A CU-level flag is signaled to indicate whether the proposed DBV mode is applied, as shown in Table 7. Table 7 Binary Processing of intra_chroma_pred_mode in the Proposed Method 2.1.4. Transformation and Quantization 2.1.4.1. Large Block Size Transformation with High-Frequency Zeroing In VVC, large block size transforms with a maximum size of 64×64 are enabled, which are mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are zeroed out, so that only the lower frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of the transform coefficients are retained. Similarly, when N equals 64, only the first 32 rows of the transform coefficients are retained. When the transform skip mode is used for large blocks, the entire block is used without zeroing out any values. Additionally, in the transform skip mode, the transform shift is removed. VTM also supports configurable maximum transform sizes in the SPS, enabling the encoder to flexibly select transform sizes of up to 32 lengths or 64 lengths according to the needs of a specific implementation. 2.1.4.2. Multiple Transform Selection (MTS) for Core Transform In addition to DCT-II adopted in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding of both inter-coded and intra-coded blocks. This multiple transform selection scheme uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 8 shows the basis functions of the selected DST / DCT. Table 8 – Transform Basis Functions of DCT-II / VIII and DSTVII for N-Point Input To maintain the orthogonality of the transform matrix, the transform matrix is quantized more precisely than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients must have 10 bits after horizontal transform and after vertical transform. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter respectively. When MTS is enabled in the SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is only applied to luminance. When one of the following conditions is applied, the MTS signaling is skipped. – The position of the last significant coefficient for the luminance TB is less than 1 (i.e., only DC) – The last significant coefficient of the luminance TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform type for the horizontal direction and for the vertical direction, respectively. The transform and signaling mapping table is shown in Table 9. Unified transform selection and implicit MTS for ISP are used by removing the intra mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra block and both intra-display MTS and inter-explicit MTS are turned on, only DST7 is used for the horizontal transform core and the vertical transform core. When it comes to the transform matrix precision, 8-bit primary transform cores are used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7; 8-point, 16-point, and 32-point DCT-2. In addition, other transform cores, including 64-point DCT-2; 4-point DCT-8; 8-point, 16-point, 32-point DST-7 and DCT-8, use 8-bit primary transform cores. Table 9 - Transform and Signaling Mapping Table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16x16 low-frequency region are retained. Similar to HEVC, the residual of a block can be encoded and decoded using the transform skip mode. To avoid redundancy in syntax encoding and decoding, when the CU-level MTS_CU_flag is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. When MTS is enabled for an inter-coded block, the implicit MTS can still be enabled. 2.1.4.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, LFNST is applied between the forward primary transform and quantization (at the encoder) and between the de-quantization and the inverse primary transform (at the decoder side). In LFNST, a 4x4 non-separable transform or an 8x8 non-separable transform is applied according to the block size. For example, 4x4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8x8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). Figure 34 The low-frequency non-separable transform (LFNST) process is shown. Using the input as an example, the application of the non-separable transform used in LFNST is described as follows. To apply 4x4 LFNST, a 4x4 input block X is first represented as a vector The non-separable transform is calculated as where represents the transform coefficient vector, and T is a 16x16 transform matrix. Subsequently, the 16x1 coefficient vector is reorganized into 4x4 blocks using the scan order (horizontal, vertical, or diagonal) for the block. Coefficients with smaller indices will be placed at positions in the 4x4 coefficient block with smaller scan indices. 2.1.4.3.1. Simplified Non-Separable Transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method such that the non-separable transform is implemented in a single pass without multiple iterations. However, the non-separable transform matrix dimension needs to be reduced to minimize the computational complexity as well as the storage space for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in the LFNST. The main idea of the reduced non-separable transform is to map an N (for 8x8 NSST, N is usually equal to 64) -dimensional vector to an R -dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix is not an NxN matrix but an R×N matrix as follows: The transformed R rows are the R bases of the N-dimensional space. The inverse transformation matrix for RT is the transpose of its forward transformation. For the 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix of the conventional 8x8 non-separable transformation matrix size is reduced to a 16x48 direct matrix. Thus, the 48×16 inverse RST matrix is used at the decoder side to generate the core (primary) transformation coefficients in the upper left 8x8 region. When the 16x48 matrix is applied instead of the 16x64 matrix with the same transformation set configuration, each 16x48 matrix obtains 48 input data from three 4x4 blocks in the upper left 8x8 block excluding the lower right 4x4 block. With the reduced dimension, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB with a reasonable performance degradation. To reduce complexity, LFNST is restricted to be applied only when all coefficients outside the first coefficient subgroup are not significant. Thus, when LFNST is applied, all only primary transformation coefficients must be zero. This allows the LFNST index signaling to be conditioned on the last significant position, thus avoiding the additional coefficient scan in the current LFNST design that is only used to check for significant coefficients at specific positions. The worst-case processing of LFNST (in terms of multiplications per pixel) is limited to 8x16 and 8x48 transformations for the non-separable transformations of 4x4 and 8x8 blocks, respectively. In these cases, when LFNST is applied, for other sizes less than 16, the last significant scan position must be less than 8. For blocks of shape 4xN and Nx4 with N>8, the proposed restriction means that now LFNST is applied only once and only to the upper left 4x4 region. Since all only primary coefficients are zero when LFNST is applied, the number of operations required for the primary transformation is reduced in this case. From the encoder's perspective, when the LFNST transformation is tested, the quantization of the coefficients is significantly simplified. The rate-distortion optimized quantization must be performed for at most the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.1.4.3.2. LFNST Transform Selection A total of 4 transform sets are used, and 2 inseparable transform matrices (kernels) are used for each transform set in the LFNST. As shown in Table 10, the mapping from the intra prediction mode to the transform set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chrominance block. For each transform set, the selected inseparable quadratic transform candidate is further specified by the LFNST index signaled through the signal. This index is signaled once in the bitstream for each intra CU after the transform coefficients. Table 10 - Transform Selection Table IntraPredMode Tr. set index IntraPredMode<0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode <= 80 1 81 <= IntraPredMode <= 83 0 2.1.4.3.3. LFNST Index Signaling and Interaction with Other Tools Since the LFNST is restricted to be applied only when all coefficients outside the first coefficient subgroup are not significant, the LFNST index encoding and decoding depends on the position of the last significant coefficient. Additionally, the LFNST index is context - encoded but does not depend on the intra prediction mode, and only the first binary bit is context - encoded. Furthermore, the LFNST applies to intra CUs in both intra slices and inter slices, and applies to both luminance and chrominance. If the dual - tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter slices (dual - tree disabled), a single LFNST index is signaled and used for both luminance and chrominance. Considering that due to the existing maximum transform size limit (64x64), large CUs larger than 64x64 are implicitly partitioned (TU tiling), the LFNST index search can quadruple the data cache for a certain number of decoding pipeline stages. Therefore, the maximum size allowed by the LFNST is limited to 64x64. Note that the LFNST only utilizes DCT2 enabled. The LFNST index signaling is placed before the MTS index signaling. The use of the scaling matrix for perceptual quantization does not imply that the scaling matrix specified for the primary matrix can be used for the LFNST coefficients. Therefore, the use of a scaling matrix for the LFNST coefficients is not allowed. For the single - tree splitting mode, the chrominance LFNST is not applied. 2.1.4.4. Sub - block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only a sub-part of the residual block is coded / decoded for the CU. When the inter-predicted CU has cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the whole residual block or a sub-part of the residual block is coded / decoded. In the former case, the MTS information is further parsed inter-frame to determine the transform type of the CU. In the latter case, a part of the residual block is coded / decoded using the presumed adaptive transform while the other part of the residual block is set to zero. When SBT is applied to an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. As Figure 35 shown, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to a binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to an asymmetric binary tree (ABT) partition. In the ABT partition, only the small region contains non-zero residuals. If one dimension of the CU is 8 in luma samples, the 1:3 / 3:1 partition along that dimension is not allowed. There are at most 8 SBT modes for a CU. Position-dependent transform core selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). Two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, in Figure 35 it, the horizontal transform and vertical transform for each SBT position are specified. For example, the horizontal transform and vertical transform for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Thus, the sub-block transform jointly specifies the TU tiling of the residual block, cbf, and the horizontal and vertical core transform types. SBT is not applied to CUs coded using the inter-intra combined mode. 2.1.4.5. Maximum Transform Size and Zeroing of Transform Coefficients Both the CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the maximum intra-coded block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the primary transform process, no canonical zeroing operation is applied to the transform coefficients. However, if LFNST is applied, the primary transform coefficients outside the LFNST region are canonically zeroed. 2.1.4.6. Enhanced MTS for Intra Coding and Decoding In the current VVC design [1], for MTS, only the DST7 and DCT8 transform kernels for intra coding and inter coding are utilized. Additional primary transforms including DCT5, DST4, DST1, and the identity transform (IDT) are adopted. The MTS set is also made to depend on the TU size and intra mode information. Sixteen different TU sizes are considered, and for each TU size, five different categories are considered according to the intra mode information. For each category, four different transform pairs are considered, which is the same as in VVC. Note that although a total of 80 different categories are considered, some of these different categories often share exactly the same transform set. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For angular modes, joint symmetry on the TU shape and intra prediction is considered. Thus, mode i (i > 34) with TU shape A×B will be mapped to the same category corresponding to mode j = (68 - i) with TU shape B×A. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For wide-angle modes, the nearest regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to mode 80. The MTS index [0, 3] is signaled by using 2-bit fixed-length coding and decoding. 2.1.4.7. Second Transform: LFNST Extension with Large Kernels The LFNST design in VVC is extended as follows: · The number of the LFNST set (S) and candidates (C) is extended to S = 35 and C = 3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: o For predModeIntra < 2, lfnstTrSetIdx is equal to 2 o For predModeIntra in [0, 34], lfnstTrSetIdx = predModeIntra o For predModeIntra in [35, 66], lfnstTrSetIdx = 68 – predModeIntra · Three different kernels (LFNST4, LFNST8, and LFNST16) are defined to indicate the LFNST kernel set, which is applied to 4xN / Nx4 (N≥4), 8xN / Nx8 (N≥8), and MxN (M, N≥16), respectively. The kernel dimensions are specified by the following formula: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96). The forward LFNST is applied to the upper-left low-frequency region, which is called the region of interest (ROI). When the LFNST is applied, the primary transform coefficients existing in the regions other than the ROI are set to zero, which remains unchanged compared with the VVC standard. The ROI for LFNST16 is as Figure 36 shown. The ROI consists of six 4x4 sub-blocks, which are consecutive in the scanning order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be Rx96. In this contribution, R is selected as 32, and 32 coefficients (two 4x4 sub-blocks) are correspondingly generated from the forward LFNST16, and these coefficients are placed in the coefficient scanning order. The ROI for LFNST8 is as Figure 37 shown. The forward LFNST8 matrix can be Rx64, and R is selected as 32. The generated coefficients are positioned in the same way as LFNST16. The mapping from the intra prediction mode to these sets is shown in the following table. Table 11. Mapping of Intra Prediction Mode to LFNST Set Index Intra prediction mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST set index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST set index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 LFNST set index 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNsT set index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.1.4.8. Non-separable Primary Transform (NSPT) for Intra Coding and Decoding DCT-II + LFNST is replaced by NSPT for block sizes 4x4, 4x8, 8x4, and 8x8. NSPT follows the design of LFNST (i.e., 3 candidates and 35 sets), which is selected based on the intra mode. The kernel sizes are as follows: · NSPT4x4: 16x16; · NSPT4x8 / NSPT8x4: 32x20; · NSPT8x8: 64x32. Therefore, 12 coefficients and 32 coefficients are set to zero for NSPT4x8 / NSPT8x4 and NSPT8x8, respectively. 2.1.4.9. Sign Prediction The basic idea of the coefficient sign prediction method (JVET-D0031 and JVET-J0021) is to calculate the reconstruction residuals for both negative and positive sign combinations for applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal sign, the cost function is defined as Figure 38 the discontinuity measurement across block boundaries shown. The cost function is measured for all hypotheses, and the hypothesis with the minimum cost is selected as the predictor for the coefficient sign. The cost function is defined as the sum of the absolute second derivatives in the residual domain for the upper row and the left column as follows: where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. The term (-R -1 +2R0-P1) can be calculated only once per block, and only the residual hypotheses are subtracted. 3. Problems In ECM-7.0, OBMC can be applied to inter-frame codec blocks, whether they are inter-frame AMVP-coded or inter-frame Merge-coded. For inter-frame AMVP-coded blocks, a syntax flag can be signaled at the block level indicating whether OBMC is applied. For inter-frame Merge-coded blocks, it is implicitly assumed that OBMC is applied regardless of the block characteristics and the codec information of neighboring blocks. However, there may be cases where some inter-frame Merge-coded blocks do not prefer the OBMC mode. For example, blocks containing sharp edges, or a small amount of gradient, or a small amount of color may not prefer the OBMC mode. Block-level adaptive OBMC that considers the prediction mode of neighboring blocks can bring higher codec gain. 4. Detailed Solutions The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term "video unit" or "codec unit" or "block" can represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. In this disclosure, regarding "blocks encoded / decoded using mode N", "mode N" here can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or an encoding / decoding technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter - inter, GPM intra - intra, GPM inter - intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.). It should be noted that the following terms are not limited to the specific terms defined in the existing standards. Any variants of the encoding / decoding tools are also applicable. 4.1. In one example, whether OBMC is applied to the current block can depend on the prediction mode of the spatial / temporal neighboring blocks adjacent / non - adjacent to the current block. 1) For example, the current block can be encoded / decoded by inter - frame Merge. 2) For example, the current block can be encoded / decoded by inter - frame AMVP. 3) For example, it can be based on whether there are neighboring blocks encoded / decoded using IBC. 4) For example, it can be based on whether there are neighboring blocks encoded / decoded using PLT. 5) For example, it can be based on whether there are neighboring blocks encoded / decoded using intraTMP. 6) For example, it can be based on whether there are neighboring blocks encoded / decoded using BDPCM. 7) For example, it can be based on whether there are neighboring blocks encoded / decoded using transform skip. 8) For example, the neighboring block can be adjacent to the current block. 9) For example, the neighboring block can be non - adjacent to the current block. 10) For example, the neighboring block can be a spatial neighboring block within the current picture. 11) For example, the neighboring block can be a temporal block in the reference picture. 12) For example, the neighboring block can be a sub - block smaller than the current block (e.g., 4x4 or 8x8). 13) For example, the neighboring block can be a video unit larger than or equal to the current block. 14) For example, the neighboring block can be a sample position. 15) For example, a series of adjacent neighboring blocks / sub-blocks to the left and / or above the current block can be examined one by one (e.g., following a predefined position and a predefined checking order). a. For example, if there is a neighbor encoded / decoded using a specific pattern, the process is terminated and it is considered that OBMC is not applied to the current block. b. For example, if there is a neighbor encoded / decoded using the INTER mode, it is further checked whether its reference block is encoded / decoded using a specific pattern (e.g., the reference block is identified by adding the motion vector associated with the INTER-encoded neighbor and the position of the INTER-encoded neighbor), and if the reference block is encoded / decoded using a specific pattern, the process is terminated and it is considered that OBMC is not applied to the current block. i. For example, the reference block is in the reference picture. c. For example, if there is a neighbor encoded / decoded using intraTMP, it can be further checked whether its reference block is encoded / decoded using a specific pattern (e.g., the reference block is identified by adding the block vector associated with the intraTMP-encoded neighbor and the position of the intraTMP-encoded neighbor), and if the reference block is encoded / decoded using a specific pattern, the process is terminated and it is considered that OBMC is not applied to the current block. i. For example, the reference block is in the current picture. d. For example, the specific pattern can be IBC and / or PLT. e. For example, the specific pattern can be intraTMP. f. For example, the specific pattern can be BDPCM. g. For example, the specific pattern can be transform skip. 16) For example, a series of non-adjacent neighboring blocks / sub-blocks in the already encoded / decoded area of the current picture can be examined one by one (e.g., following a predefined position and a predefined checking order). h. For example, if there is a neighbor encoded / decoded using a specific pattern, the process is terminated and it is considered that OBMC is not applied to the current block. i. For example, if there is a neighbor encoded / decoded using the INTER mode, it is checked whether its reference block is encoded / decoded using a specific pattern (e.g., the reference block is identified by adding the motion vector associated with the INTER-encoded neighbor and the position of the INTER-encoded neighbor), and if the reference block is encoded / decoded using a specific pattern, the process is terminated and it is considered that OBMC is not applied to the current block. i. For example, the reference block is in the reference picture. j. For example, if there is a neighbor decoded using intraTMP, whether its reference block is decoded using a specific mode (e.g., the reference block is identified by adding the block vector associated with the neighbor decoded using intraTMP and the position of the neighbor decoded using intraTMP), and if the reference block is decoded using a specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. i. For example, the reference block is in the current picture. k. For example, the specific mode can be IBC and / or PLT. l. For example, the specific mode can be intraTMP. m. For example, the specific mode can be BDPCM. n. For example, the specific mode can be transform skip. 17) For example, a series of temporal blocks / sub-blocks in the reference picture can be checked one by one (e.g., following a predefined position and order). o. For example, if there is a temporal block decoded using a specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. p. For example, if there is a temporal block decoded using INTER mode, whether its reference block is decoded using a specific mode (e.g., the reference block is identified by adding the motion vector associated with the temporal block decoded using INTER and the position of the temporal block decoded using INTER) is further checked, and if the reference block is decoded using a specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. q. For example, if there is a temporal block decoded using intraTMP, whether its reference block is decoded using a specific mode (e.g., the reference block is identified by adding the block vector associated with the temporal block decoded using intraTMP and the position of the temporal block decoded using intraTMP) is further checked, and if the reference block is decoded using a specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. r. For example, the specific mode can be IBC and / or PLT. s. For example, the specific mode can be intraTMP. t. For example, the specific mode can be BDPCM. u. For example, the specific mode can be transform skip. 4.2. In one example, whether OBMC is applied to the current block can depend on the prediction mode of the reference block. 1) For example, the current block can be decoded using inter-frame Merge. 2) For example, the current block may be coded / decoded by inter AMVP. 3) For example, the reference block may be a block / sub-block identified based on adding a displacement (e.g., predefined or based on a motion vector or a block vector) to the position of the first block. a. For example, the first block may be the current block. b. For example, the first block may be a neighboring block adjacent to the current block. c. For example, the first block may be a neighboring block non - adjacent to the current block. d. For example, the first block may be a reference block of the current block. e. For example, the first block may be a reference block of a neighboring block. f. For example, the reference block may be identified based on the position of the current block coded / decoded in INTER mode and its motion information (e.g., motion vector and reference index) associated with the current block. i. For example, in this case, the reference block is in the reference picture. g. For example, the reference block may be identified based on the position of the current block coded / decoded in IntraTMP mode and its motion information (e.g., block vector) associated with the current block. i. For example, in this case, the reference block is in the current picture. h. For example, the reference block may be identified based on the position of the neighboring block coded / decoded in INTER mode and its motion information (e.g., motion vector and reference index) associated with the INTER - coded neighboring block. i. For example, in this case, the reference block is in the reference picture. i. For example, the reference block may be identified based on the position of the neighboring block coded / decoded in IntraTMP mode and its motion information (e.g., block vector) associated with the IntraTMP - coded neighboring block. i. For example, in this case, the reference block is in the current picture. j. For example, the reference block may be identified based on the position of the reference block coded / decoded in INTER mode and its motion information (e.g., motion vector and reference index) associated with the INTER - coded reference block. i. For example, in this case, the reference block is in another reference picture (rather than the reference picture where the INTER - coded reference block is located). k. For example, the reference block may be identified based on the position of the reference block coded / decoded in IntraTMP mode and its motion information (e.g., block vector) associated with the IntraTMP - coded reference block. i. For example, in this case, the reference block is in the same reference picture as the reference block located in the reference block encoded / decoded in IntraTMP mode. 4) For example, in the case where neighboring blocks are encoded / decoded using INTER mode, the reference block can be examined. a. If the neighbor is inter-frame (INTER) encoded / decoded, the reference block is identified by adding the motion vector associated with the INTER-encoded / decoded neighbor and the position of the INTER-encoded / decoded neighbor, and if the reference block is encoded / decoded using a specific mode, it is considered that OBMC is not applied to the current block. b. For example, the specific mode can be IBC and / or PLT. c. For example, the specific mode can be intraTMP. d. For example, the specific mode can be BDPCM. e. For example, the specific mode can be transform skip. 5) For example, in the case where neighboring blocks are encoded / decoded using IntraTMP mode, the reference block can be examined. a. If the neighbor is intraTMP encoded / decoded, the reference block is identified by adding the block vector associated with the intraTMP-encoded / decoded neighbor and the position of the intraTMP-encoded / decoded neighbor, and if the reference block is encoded / decoded using a specific mode, it is considered that OBMC is not applied to the current block. b. For example, the specific mode can be IBC and / or PLT. c. For example, the specific mode can be BDPCM. d. For example, the specific mode can be transform skip. 4.3. In one example, whether OBMC is applied to the current block can depend on the prediction mode of the reference block of the reference block. 1) For example, the current block can be inter-frame Merge encoded / decoded. 2) For example, the current block can be inter-frame AMVP encoded / decoded. 3) For example, the reference block of the reference block can be identified by adding the motion vector associated with the reference block encoded / decoded in INTER mode and the position of the INTER-encoded / decoded reference block. 4) For example, the reference block of the reference block can be identified by adding the block vector associated with the reference block encoded / decoded in IntraTMP mode and the position of the IntraTMP-encoded / decoded reference block. 5) For example, the historical / propagated prediction mode can be stored in the cache. a. For example, if the block itself is encoded / decoded using a specific mode or it has a reference block that was encoded / decoded using a specific mode, the specific mode can be stored in a cache associated with the block information, indicating the information about the block having a history / propagation of the specific mode. 6) For example, the specific mode can be IBC and / or PLT. 7) For example, the specific mode can be intraTMP. 8) For example, the specific mode can be BDPCM. 9) For example, the specific mode can be transform skip. 4.4. In one example, whether a block in the current picture is encoded / decoded in a specific mode is stored in the cache. 1) For example, the specific mode can be IBC. 2) For example, the specific mode can be PLT. 3) For example, the specific mode can be intraTMP. 4) For example, the specific mode can be BDPCM. 5) For example, the specific mode can be transform skip. 6) For example, whether the block is encoded / decoded using IBC or PLT can be stored using a shared parameter / cache. a. For example, the storage may require a single parameter / cache. 7) For example, whether the block is encoded / decoded using IBC or PLT or intraTMP or BDPCM can be stored using separate parameters / caches. b. For example, the storage may require multiple parameters / caches. 8) For example, the storage can be done at the granularity of MxM (e.g., M = 4 or 8) sub - blocks. 4.5. In one example, whether the block and / or its reference block is encoded / decoded in a specific mode can be stored in the cache. 1) For example, the specific mode can be IBC. 2) For example, the specific mode can be PLT. 3) For example, the specific mode can be intraTMP. 4) For example, the specific mode can be BDPCM. 5) For example, the specific mode can be transform skip. 6) For example, if the block or its reference block is encoded / decoded using IBC or PLT, a parameter equal to true (e.g., indicating that the block or reference block is a historical / propagated screen content block) can be stored in the cache. a. On the other hand, alternatively, a parameter equal to false (e.g., indicating that the block or reference block is not a historical / propagated screen content block) can be stored in the cache. 7) For example, whether a block is encoded / decoded using IBC or PLT and whether the reference block of the block is encoded / decoded using IBC or PLT are stored as separate parameters and stored in separate caches. 8) For example, storage can be performed at the granularity of MxM (e.g., M = 4 or 8) sub-blocks. 4.6. In one example, whether OBMC is enabled can be coupled with whether a specific tool is enabled. 1) For example, the specific tool can be IBC. 2) For example, the specific tool can be PLT. 3) For example, the specific tool can be intraTMP. 4) For example, the specific tool can be BDPCM. 5) For example, the specific tool can be transform skip. 6) In one example, whether to apply a specific tool can be controlled by a first syntax element (SE), e.g., controlled in VPS / SPS / PPS / slice header / CTU / CU / etc. 7) In one example, whether to apply OBMC can be controlled by a second syntax element (SE), e.g., controlled in VPS / SPS / PPS / slice header / CTU / CU / etc. 8) In one example, it can be constrained that if the first SE indicates that the specific tool is enabled, the second SE must indicate that OBMC is to be disabled. 9) In one example, it can be set at the encoder that if the first SE indicates that the specific tool is enabled, the second SE must indicate that OBMC is to be disabled. 4.7. Whether and / or how to apply the methods disclosed above can be signaled at the sequence level / group of pictures level / picture level / slice level / slice group level, e.g., signaled in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 4.8. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / picture / sub-picture / other types of regions containing more than one sample or pixel. 4.9. Whether and / or how to apply the methods disclosed above can depend on the encoded / decoded information, e.g., block size, color format, single / double-tree segmentation, color component, slice / picture type.

[0099] The term "video unit" or "coding unit" or "block" may represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. In the present disclosure, regarding "a block coded using mode N", "mode N" here may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame - inter-frame, GPM intra-frame - intra-frame, GPM inter-frame - intra-frame, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.).

[0100] Figure 39 A flowchart of a method 3900 for video processing according to an embodiment of the present disclosure is shown. Method 3900 is implemented during the conversion between a target video block of a video and a bitstream of the video.

[0101] At block 3910, for the conversion between a video unit of a video and a bitstream of the video unit, it is determined based on at least one of the following whether overlapping block motion compensation (OBMC) is applied to the current block of the video unit: the prediction mode of a neighboring block, the prediction mode of a reference block, the prediction mode of a first reference block of a second reference block, or whether a target coding tool is enabled for the current block.

[0102] At block 3920, the conversion is performed based on the determination. In some embodiments, the conversion may include encoding a video unit from the bitstream. Alternatively or additionally, the conversion may include decoding a video unit from the bitstream. In this way, block-level adaptive OBMC considering the prediction mode of neighboring blocks can bring higher coding gain and improve coding efficiency.

[0103] In some embodiments, the neighboring block includes at least one of the following: a spatial neighboring block adjacent to the current block, a spatial neighboring block non-adjacent to the current block, a temporal neighboring block adjacent to the current block, or a temporal neighboring block non-adjacent to the current block.

[0104] In some embodiments, the current block is coded by inter-frame Merge. Alternatively, the current block is coded by inter-frame advanced motion vector prediction (AMVP).

[0105] In some embodiments, whether OBMC is applied to a current block is based on whether there are neighboring blocks that are encoded / decoded using Intra Block Copy (IBC). In some embodiments, whether OBMC is applied to a current block is based on whether there are neighboring blocks that are encoded / decoded using a palette. In some embodiments, whether OBMC is applied to a current block is based on whether there are neighboring blocks that are encoded / decoded using Intra Template Matching Prediction (intraTMP). In some embodiments, whether OBMC is applied to a current block is based on whether there are neighboring blocks that are encoded / decoded using Block Differential Pulse Code Modulation (BDPCM). In some embodiments, whether OBMC is applied to a current block is based on whether there are neighboring blocks that are encoded / decoded using transform skip.

[0106] In some embodiments, a neighboring block is adjacent to the current block. In some embodiments, a neighboring block is not adjacent to the current block. In some embodiments, a neighboring block is a spatial neighboring block within the current picture. In some embodiments, a neighboring block is a temporal block in a reference picture. In some embodiments, a neighboring block is a sub-block smaller than the current block (e.g., 4x4 or 8x8). In some embodiments, a neighboring block is a video unit larger than or equal to the current block. In some embodiments, a neighboring block is a sample position.

[0107] In some embodiments, an inspection process is applied to a series of adjacent neighboring blocks or sub-blocks to the left and / or above the current block to determine whether OBMC is applied to the current block. 15) For example, a series of adjacent neighboring blocks / sub-blocks to the left and / or above the current block can be inspected one by one (e.g., following a predefined position and a predefined inspection order). In some embodiments, if there is a neighboring block encoded / decoded using a target mode, the inspection process is terminated, and it is considered that OBMC is not applied to the current block.

[0108] In some embodiments, if there is a neighboring block encoded / decoded using an inter mode, it is further checked whether the reference block of the neighboring block encoded / decoded using the inter mode is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the inspection process is terminated, and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the motion vector associated with the neighboring block encoded / decoded using the inter mode and the position of the neighboring block encoded / decoded using the inter mode. In some embodiments, the reference block is in a reference picture.

[0109] In some embodiments, if there is a neighbor that is encoded / decoded using intraTMP, it is further checked whether the reference block of the neighbor block encoded / decoded using intraTMP is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the block vector associated with the neighbor block encoded / decoded using intraTMP and the position of the neighbor block encoded / decoded using intraTMP. In some embodiments, the reference block is in the current picture. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0110] In some embodiments, the checking process is applied to a series of non - adjacent neighbor blocks or sub - blocks in the already - encoded region of the current picture one by one to determine whether OBMC is applied to the current block. For example, a series of non - adjacent neighbor blocks / sub - blocks in the already - encoded region of the current picture can be checked one by one (e.g., following a predefined position and a predefined checking order). In some embodiments, if there is a neighbor block encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

[0111] In some embodiments, if there is a neighbor block encoded / decoded using an inter mode, it is further checked whether the reference block of the neighbor block encoded / decoded using the inter mode is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the motion vector associated with the neighbor block encoded / decoded using the inter mode and the position of the neighbor block encoded / decoded using the inter mode. In some embodiments, the reference block is in the reference picture.

[0112] In some embodiments, if there is a neighbor that is encoded / decoded using intraTMP, it is further checked whether the reference block of the neighbor block encoded / decoded using intraTMP is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the block vector associated with the neighbor block encoded / decoded using intraTMP and the position of the neighbor block encoded / decoded using intraTMP. In some embodiments, the reference block is in the current picture. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0113] In some embodiments, the checking process is applied to a series of temporal blocks or sub - blocks in a reference picture one by one to determine whether OBMC is applied to the current block. For example, a series of temporal blocks / sub - blocks in the reference picture can be checked one by one (e.g., following a predefined position and order). In some embodiments, if there is a temporal block encoded / decoded using a target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

[0114] In some embodiments, if there is a temporal block encoded / decoded using an inter - frame mode, it is further checked whether the reference block of the temporally - block encoded / decoded by the inter - frame mode is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the motion vector associated with the temporally - block encoded / decoded by the inter - frame mode and the position of the temporally - block encoded / decoded by the inter - frame mode.

[0115] In some embodiments, if there is a temporal block encoded / decoded using intraTMP, it is further checked whether the reference block of the temporally - block encoded / decoded by intraTMP is encoded / decoded using the target mode, and if the reference block is encoded / decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block. In some embodiments, the reference block is identified by adding the block vector associated with the temporally - block encoded / decoded by intraTMP and the position of the temporally - block encoded / decoded by intraTMP. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0116] In some embodiments, whether OBMC is applied to the current block depends on the prediction mode of the reference block. In some embodiments, the current block is inter - frame Merge encoded / decoded. Alternatively, the current block is inter - frame AMVP encoded / decoded.

[0117] In some embodiments, the reference block is a block or sub - block identified by adding a displacement to the position of the first block. For example, the reference block can be a block / sub - block identified by adding a displacement (e.g., predefined or based on a motion vector or based on a block vector) to the position of the first block.

[0118] In some embodiments, the first block is the current block. In some embodiments, the first block is a neighboring block adjacent to the current block. In some embodiments, the first block is a neighboring block non - adjacent to the current block. In some embodiments, the first block is the reference block of the current block. In some embodiments, the first block is the reference block of a neighboring block.

[0119] In some embodiments, a reference block is identified based on the position of a current block encoded or decoded in an inter-frame mode and motion information associated with the current block (e.g., a motion vector and a reference index). In some embodiments, the reference block is in a reference picture.

[0120] In some embodiments, a reference block is identified based on the position of a current block encoded or decoded in an IntraTMP mode and motion information associated with the current block (e.g., a block vector). In some embodiments, the reference block is in the current picture.

[0121] In some embodiments, a reference block is identified based on the position of a neighboring block encoded or decoded in an inter-frame mode and motion information associated with the neighboring block encoded or decoded in an INTER mode (e.g., a motion vector and a reference index). In some embodiments, the reference block is in a reference picture.

[0122] In some embodiments, a reference block is identified based on the position of a neighboring block encoded or decoded in an IntraTMP mode and motion information associated with the neighboring block encoded or decoded in an IntraTMP mode (e.g., a block vector). In some embodiments, the reference block is in the current picture.

[0123] In some embodiments, a reference block is identified based on the position of a reference block encoded or decoded in an inter-frame mode and motion information associated with the reference block encoded or decoded in an inter-frame mode (e.g., a motion vector and a reference index). In some embodiments, the reference block is in another reference picture, rather than the reference picture in which the reference block encoded or decoded in an inter-frame mode is located.

[0124] In some embodiments, a reference block is identified based on the position of a reference block encoded or decoded in an IntraTMP mode and motion information associated with the reference block encoded or decoded in an IntraTMP mode (e.g., a block vector). In some embodiments, the reference block is in the same reference picture in which the reference block encoded or decoded in an IntraTMP mode is located.

[0125] In some embodiments, the reference block is checked in the case where a neighboring block is encoded or decoded using an INTER mode. In some embodiments, if a neighboring block is encoded or decoded using an inter-frame mode, the reference block is identified by adding the motion vector associated with the neighboring block encoded or decoded in an inter-frame mode and the position of the neighboring block encoded or decoded in an inter-frame mode, and if the reference block is encoded or decoded using a target mode, it is considered that OBMC is not applied to the current block. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0126] In some embodiments, when a neighboring block is coded or decoded using the IntraTMP mode, the reference block is checked. In some embodiments, if a neighboring block is coded or decoded using the IntraTMP mode, the reference block is identified by adding the block vector associated with the IntraTMP-coded neighboring block and the position of the IntraTMP-coded neighboring block, and if the reference block is coded or decoded using the target mode, it is considered that OBMC is not applied to the current block. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, BDPCM mode, or transform skip mode.

[0127] In some embodiments, whether OBMC is applied to the current block depends on the prediction mode of the first reference block of the second reference block. In some embodiments, the current block is inter-frame Merge coded. Alternatively, the current block is inter-frame AMVP coded.

[0128] In some embodiments, the first reference block of the second reference block is identified by adding the motion vector associated with the second reference block, which is a reference block coded in the inter-frame mode, and the position of the second reference block. In some embodiments, the first reference block of the second reference block is identified by adding the block vector associated with the second reference block, which is a reference block coded in the IntraTMP mode, and the position of the second reference block.

[0129] In some embodiments, the historical prediction mode or the propagated prediction mode is stored in a cache. In some embodiments, if a block is coded or decoded using the target mode or the reference block of the block is coded or decoded using the target mode, the target mode is stored in the cache associated with the block information, which indicates that the block has historical information or propagated information using the target mode. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0130] In some embodiments, whether a block in the current picture is coded or decoded using the target mode is stored in a cache. The block can be one of the following: the current block, a neighboring block, or a reference block. In some embodiments, the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0131] In some embodiments, whether a block is coded or decoded using IBC or palette is stored using a shared parameter or a cache. In some embodiments, a single parameter or cache is required to store whether a block is coded or decoded using IBC or palette.

[0132] In some embodiments, whether a block is encoded or decoded using IBC or palette or intraTMP or BDPCM is stored using separate parameters or caches. In some embodiments, multiple parameters or caches are required to store whether a block is encoded or decoded using IBC or palette or intraTMP or BDPCM.

[0133] In some embodiments, whether a block in the current picture is encoded or decoded in a target mode is stored at the granularity of MxM sub-blocks. In some embodiments, M is equal to 4 or 8.

[0134] In some embodiments, whether at least one of a block or a reference block of the block is encoded or decoded in a target mode is stored in a cache. The block can be one of the following: a current block, a neighboring block, or a second reference block. In some embodiments, the target mode is at least one of the following: an IBC mode, a palette mode, an intraTMP mode, a BDPCM mode, or a transform skip mode.

[0135] In some embodiments, if a block or a reference block of the block is encoded or decoded using IBC or palette, a parameter equal to true is stored in the cache. For example, if a block or its reference block is encoded or decoded using IBC or PLT, a parameter equal to true (e.g., indicating that the block or its reference block is a historical / propagated screen content block) can be stored in the cache. In some embodiments, if a block or a reference block of the block is not encoded or decoded using IBC or palette, a parameter equal to false (e.g., indicating that the block or reference block is not a historical / propagated screen content block) is stored in the cache.

[0136] In some embodiments, whether a block is encoded or decoded using IBC or PLT and whether a reference block of the block is encoded or decoded using IBC or palette are stored as separate parameters and stored in separate caches. In some embodiments, whether at least one of a block or a reference block of the block is encoded or decoded in a target mode is stored at the granularity of MxM sub-blocks. In some embodiments, M is equal to 4 or 8.

[0137] In one example, whether OBMC is enabled can be coupled with whether a specific tool is enabled. In some embodiments, the target encoding or decoding tools include at least one of the following: IBC, palette, intraTMP, BDPCM, or transform skip.

[0138] In some embodiments, whether to apply a target codec tool is controlled by a first syntax element (SE). Alternatively or additionally, whether to apply OBMC can be controlled by a second syntax element. In some embodiments, the first syntax element is indicated in one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), slice header, Coding Tree Unit (CTU), or Coding Unit (CU). In some embodiments, the second syntax element is indicated in one of the following: VPS, SPS, PPS, slice header, CTU, or CU.

[0139] In some embodiments, if the first syntax element indicates that the target codec tool is enabled, the second syntax element indicates that OBMC is disabled. In some embodiments, if the first syntax element indicates that the target codec tool is enabled, the second syntax element indicates that OBMC will be disabled at the encoder setting.

[0140] In some embodiments, an indication of whether and / or how to determine whether OBMC is applied to a current block is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0141] In some embodiments, an indication of whether and / or how to determine whether OBMC is applied to a current block is indicated in one of the following: sequence header, picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, or slice group header. In some embodiments, an indication of whether and / or how to determine whether OBMC is applied to a current block is included in one of the following: Prediction Block (PB), Transform Block (TB), Coding Block (CB), Prediction Unit (PU), Transform Unit (TU), Coding Unit (CU), Virtual Pipeline Data Unit (VPDU), Coding Tree Unit (CTU), CTU row, slice, tile, sub-picture, or a region containing more than one sample or pixel.

[0142] In some embodiments, method 3900 further includes: determining whether and / or how to determine whether OBMC is applied to a current block based on the coded information of a video unit, the coded information including at least one of the following: block size, color format, single-tree segmentation and / or dual-tree segmentation, color component, slice type, or picture type.

[0143] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, and the bitstream of the video is generated by a method executed by a device for video processing. The method includes: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; and generating a bitstream based on the determination.

[0144] According to still some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0145] Embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.

[0146] Item 1. A method for video processing, including: for the conversion between a video unit of a video and a bitstream of the video unit, determining whether overlapping block motion compensation (OBMC) is applied to a current block of the video unit based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; and performing the conversion based on the determination.

[0147] Item 2. The method according to Item 1, wherein the neighboring block includes at least one of the following: a spatial neighboring block adjacent to the current block, a spatial neighboring block non-adjacent to the current block, a temporal neighboring block adjacent to the current block, or a temporal neighboring block non-adjacent to the current block.

[0148] Item 3. The method according to Item 2, wherein the current block is inter-frame Merge coded or the current block is inter-frame advanced motion vector prediction (AMVP) coded.

[0149] Item 4. The method according to Item 2, wherein whether the OBMC is applied to the current block is based on whether there is a neighboring block coded by using intra block copy (IBC).

[0150] Item 5. The method according to Item 2, wherein whether the OBMC is applied to the current block is based on whether there are neighboring blocks encoded using a palette.

[0151] Item 6. The method according to Item 2, wherein whether the OBMC is applied to the current block is based on whether there are neighboring blocks encoded using intra-template matching prediction (intraTMP).

[0152] Item 7. The method according to Item 2, wherein whether the OBMC is applied to the current block is based on whether there are neighboring blocks encoded using block differential pulse codec modulation (BDPCM).

[0153] Item 8. The method according to Item 2, wherein whether the OBMC is applied to the current block is based on whether there are neighboring blocks encoded using transform skip.

[0154] Item 9. The method according to any one of Items 1-8, wherein the neighboring block is adjacent to the current block.

[0155] Item 10. The method according to any one of Items 1-8, wherein the neighboring block is non-adjacent to the current block.

[0156] Item 11. The method according to any one of Items 1-8, wherein the neighboring block is a spatial neighboring block within the current picture.

[0157] Item 12. The method according to any one of Items 1-8, wherein the neighboring block is a temporal block in a reference picture.

[0158] Item 13. The method according to any one of Items 1-8, wherein the neighboring block is a sub-block smaller than the current block.

[0159] Item 14. The method according to any one of Items 1-8, wherein the neighboring block is a video unit greater than or equal to the current block.

[0160] Item 15. The method according to any one of Items 1-8, wherein the neighboring block is a sample position.

[0161] Item 16. The method according to Item 1, wherein the checking process is applied one by one to a series of adjacent neighboring blocks or sub-blocks to the left and / or above the current block to determine whether the OBMC is applied to the current block.

[0162] Item 17. The method according to Item 16, wherein if there is a neighboring block encoded using a target mode, the checking process is terminated, and it is considered that the OBMC is not applied to the current block.

[0163] Item 18. The method according to Item 16, wherein if there is a neighboring block encoded or decoded using an inter mode, it is further checked whether a reference block of the neighboring block encoded or decoded using the inter mode is encoded or decoded using a target mode, and if the reference block is encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0164] Item 19. The method according to Item 18, wherein the reference block is identified by adding a motion vector associated with the neighboring block encoded or decoded using the inter mode and the position of the neighboring block encoded or decoded using the inter mode.

[0165] Item 20. The method according to Item 18, wherein the reference block is in a reference picture.

[0166] Item 21. The method according to Item 16, wherein if there is a neighbor encoded or decoded using intraTMP, it is further checked whether a reference block of the neighboring block encoded or decoded using intraTMP is encoded or decoded using a target mode, and if the reference block is encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0167] Item 22. The reference block is identified by adding a block vector associated with the neighboring block encoded or decoded using intraTMP and the position of the neighboring block encoded or decoded using intraTMP.

[0168] Item 23. The method according to Item 21, wherein the reference block is in the current picture.

[0169] Item 24. The method according to any one of Items 17 - 23, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0170] Item 25. The method according to Item 1, wherein the checking process is applied to a series of non - adjacent neighboring blocks or sub - blocks in an already encoded or decoded area of the current picture one by one to determine whether OBMC is applied to the current block.

[0171] Item 26. The method according to Item 25, wherein if there is a neighboring block encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0172] Item 27. The method according to Item 25, wherein if there is a neighboring block encoded or decoded using an inter mode, whether a reference block of the neighboring block encoded or decoded using the inter mode is encoded or decoded using a target mode is further checked, and if the reference block is encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0173] Item 28. The method according to Item 27, wherein the reference block is identified by adding a motion vector associated with the neighboring block encoded or decoded using the inter mode and the position of the neighboring block encoded or decoded using the inter mode.

[0174] Item 29. The method according to Item 27, wherein the reference block is in a reference picture.

[0175] Item 30. The method according to Item 25, wherein if there is a neighbor encoded or decoded using intraTMP, whether a reference block of the neighboring block encoded or decoded using intraTMP is encoded or decoded using a target mode is further checked, and if the reference block is encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0176] Item 31. The method according to Item 30, wherein the reference block is identified by adding a block vector associated with the neighboring block encoded or decoded using intraTMP and the position of the neighboring block encoded or decoded using intraTMP.

[0177] Item 32. The method according to Item 30, wherein the reference block is in the current picture.

[0178] Item 33. The method according to any one of Items 25 - 32, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0179] Item 34. The method according to Item 1, wherein the checking process is applied to a series of temporal blocks or sub - blocks in a reference picture one by one to determine whether OBMC is applied to the current block.

[0180] Item 35. The method according to Item 34, wherein if there is a temporal block encoded or decoded using the target mode, the checking process is terminated, and it is considered that OBMC is not applied to the current block.

[0181] Item 36. The method according to Item 34, wherein if there is a temporal block encoded or decoded using an inter-frame mode, it is further checked whether a reference block of the temporal block encoded or decoded using the inter-frame mode is encoded or decoded using a target mode, and if the reference block is encoded or decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

[0182] Item 37. The method according to Item 36, wherein the reference block is identified by adding a motion vector associated with the temporal block encoded or decoded using the inter-frame mode and the position of the temporal block encoded or decoded using the inter-frame mode.

[0183] Item 38. The method according to Item 34, wherein if there is a temporal block encoded or decoded using intraTMP, it is further checked whether a reference block of the temporal block encoded or decoded using intraTMP is encoded or decoded using a target mode, and if the reference block is encoded or decoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

[0184] Item 39. The method according to Item 38, wherein the reference block is identified by adding a block vector associated with the temporal block encoded or decoded using intraTMP and the position of the temporal block encoded or decoded using intraTMP.

[0185] Item 40. The method according to any one of Items 34 - 39, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0186] Item 41. The method according to Item 1, wherein whether OBMC is applied to the current block depends on the prediction mode of the reference block.

[0187] Item 42. The method according to Item 41, wherein the current block is encoded or decoded using inter-frame Merge, or wherein the current block is encoded or decoded using inter-frame AMVP.

[0188] Item 43. The method according to Item 41, wherein the reference block is a block or sub-block identified by adding a displacement to the position of a first block.

[0189] Item 44. The method according to Item 43, wherein the first block is the current block.

[0190] Item 45. The method according to Item 43, wherein the first block is a neighboring block adjacent to the current block.

[0191] Item 46. The method according to item 43, wherein the first block is a neighboring block that is non - adjacent to the current block.

[0192] Item 47. The method according to item 43, wherein the first block is a reference block of the current block.

[0193] Item 48. The method according to item 43, wherein the first block is a reference block of a neighboring block.

[0194] Item 49. The method according to item 43, wherein the reference block is identified based on the position of the current block encoded / decoded in an inter - frame mode and the motion information associated with the current block.

[0195] Item 50. The method according to item 49, wherein the reference block is in a reference picture.

[0196] Item 51. The method according to item 43, wherein the reference block is identified based on the position of the current block encoded / decoded in an IntraTMP mode and the motion information associated with the current block.

[0197] Item 52. The method according to item 51, wherein the reference block is in the current picture.

[0198] Item 53. The method according to item 43, wherein the reference block is identified based on the position of a neighboring block encoded / decoded in an inter - frame mode and the motion information associated with the neighboring block encoded / decoded in an INTER mode.

[0199] Item 54. The method according to item 53, wherein the reference block is in a reference picture.

[0200] Item 55. The method according to item 43, wherein the reference block is identified based on the position of a neighboring block encoded / decoded in an IntraTMP mode and the motion information associated with the neighboring block encoded / decoded in an IntraTMP mode.

[0201] Item 56. The method according to item 55, wherein the reference block is in the current picture.

[0202] Item 57. The method according to item 43, wherein the reference block is identified based on the position of a reference block encoded / decoded in an inter - frame mode and the motion information associated with the reference block encoded / decoded in an inter - frame mode.

[0203] Item 58. The method according to item 57, wherein the reference block is in another reference picture, rather than in the reference picture in which the reference block encoded / decoded in an inter - frame mode is located.

[0204] Item 59. The method according to item 43, wherein the reference block is identified based on the position of the reference block encoded and decoded in IntraTMP mode and the motion information associated with the reference block encoded and decoded in IntraTMP mode.

[0205] Item 60. The method according to item 59, wherein the reference block is in the same reference picture as the reference block in which the IntraTMP mode encoded and decoded is located.

[0206] Item 61. The method according to item 41, wherein the reference block is checked when the neighboring block is encoded and decoded using the INTER mode.

[0207] Item 62. The method according to item 61, wherein if the neighboring block is encoded and decoded using an inter-frame mode, the reference block is identified by adding the motion vector associated with the inter-frame encoded neighboring block and the position of the inter-frame encoded neighboring block, and if the reference block is encoded and decoded using the target mode, it is considered that the OBMC is not applied to the current block.

[0208] Item 63. The method according to item 62, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0209] Item 64. The method according to item 41, wherein the reference block is checked when the neighboring block is encoded and decoded using the IntraTMP mode.

[0210] Item 65. The method according to item 64, wherein if the neighboring block is encoded and decoded using the IntraTMP mode, the reference block is identified by adding the block vector associated with the IntraTMP encoded neighboring block and the position of the IntraTMP encoded neighboring block, and if the reference block is encoded and decoded using the target mode, it is considered that the OBMC is not applied to the current block.

[0211] Item 66. The method according to item 65, wherein the target mode is at least one of the following: IBC mode, palette mode, BDPCM mode, or transform skip mode.

[0212] Item 67. The method according to item 1, wherein whether the OBMC is applied to the current block depends on the prediction mode of the first reference block of the second reference block.

[0213] Item 68. The method according to Item 67, wherein the current block is inter-frame Merge coded or the current block is inter-frame AMVP coded.

[0214] Item 69. The method according to Item 67, wherein the first reference block of the second reference block is identified by adding a motion vector associated with the second reference block that is coded in an inter-frame mode and the position of the second reference block.

[0215] Item 70. The method according to Item 67, wherein the first reference block of the second reference block is identified by adding a block vector associated with the second reference block that is coded in an IntraTMP mode and the position of the second reference block.

[0216] Item 71. The method according to Item 67, wherein a history prediction mode or a propagated prediction mode is stored in a cache.

[0217] Item 72. The method according to Item 71, wherein if a block is coded using a target mode or a reference block of the block is coded using a target mode, the target mode is stored in a cache associated with block information, and the block information indicates that the block has historical information or propagated information using the target mode.

[0218] Item 73. The method according to Item 72, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0219] Item 74. The method according to any one of Items 1-73, wherein whether a block in a current picture is coded using a target mode is stored in a cache, and wherein the block is one of the following: the current block, a neighboring block, or a reference block.

[0220] Item 75. The method according to Item 74, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0221] Item 76. The method according to Item 74, wherein whether the block is coded using IBC or palette is stored using shared parameters or a cache.

[0222] Item 77. The method according to Item 76, wherein a single parameter or cache is required to store whether the block is coded using IBC or palette.

[0223] Item 78. The method according to item 74, wherein whether the block is encoded or decoded using IBC or palette or intraTMP or BDPCM is stored using separate parameters or caches.

[0224] Item 79. The method according to item 78, wherein multiple parameters or caches are required to store whether the block is encoded or decoded using IBC or palette or intraTMP or BDPCM.

[0225] Item 80. The method according to item 74, wherein whether the block in the current picture is encoded in a target mode is stored at the granularity of MxM sub-blocks.

[0226] Item 81. The method according to item 80, wherein M is equal to 4 or 8.

[0227] Item 82. The method according to any one of items 1 - 73, wherein whether at least one of the block or the reference block of the block is encoded in a target mode is stored in a cache, and wherein the block is one of the following: the current block, the neighboring block, or the second reference block.

[0228] Item 83. The method according to item 82, wherein the target mode is at least one of the following: IBC mode, palette mode, intraTMP mode, BDPCM mode, or transform skip mode.

[0229] Item 84. The method according to item 82, wherein if the block or the reference block of the block is encoded using IBC or palette, a parameter equal to true is stored in the cache.

[0230] Item 85. The method according to item 82, wherein if the block or the reference block of the block is not encoded using IBC or palette, a parameter equal to false is stored in the cache.

[0231] Item 86. The method according to item 82, wherein whether the block is encoded using IBC or PLT and whether the reference block of the block is encoded using IBC or palette are stored as separate parameters and stored in separate caches.

[0232] Item 87. The method according to item 82, wherein whether at least one of the block or the reference block of the block is encoded in a target mode is stored at the granularity of MxM sub-blocks.

[0233] Item 88. The method according to item 87, wherein M is equal to 4 or 8.

[0234] Item 89. The method according to Item 1, wherein the target codec tool comprises at least one of the following: IBC, palette, intraTMP, BDPCM, or transform skip.

[0235] Item 90. The method according to Item 89, wherein whether the target codec tool is applied is controlled by a first syntax element (SE), and / or wherein whether OBMC is applied is controlled by a second syntax element.

[0236] Item 91. The method according to Item 90, wherein if the first syntax element indicates that the target codec tool is enabled, the second syntax element indicates that the OBMC will be disabled.

[0237] Item 92. The method according to Item 90, wherein if the first syntax element indicates that the target codec tool is enabled, the second syntax element indicates that the OBMC will be disabled and is set at the encoder.

[0238] Item 93. The method according to any one of Items 90-92, wherein the first syntax element is indicated in one of the following: video parameter set (VPS), sequence parameter set (SPS), picture parameter set (PPS), slice header, coding tree unit (CTU), or coding unit (CU), or wherein the second syntax element is indicated in one of the following: VPS, SPS, PPS, slice header, CTU, or CU.

[0239] Item 94. The method according to any one of Items 1-93, wherein the indication of whether and / or how it is determined whether the OBMC is applied to the current block is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.

[0240] Item 95. The method according to any one of Items 1-93, wherein the indication of whether and / or how it is determined whether the OBMC is applied to the current block is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

[0241] Item 96. The method according to any one of Items 1-93, wherein an indication of whether and / or how it is determined whether the OBMC is applied to the current block is included in one of the following: a prediction block (PB), a transform block (TB), a coding / decoding block (CB), a prediction unit (PU), a transform unit (TU), a coding / decoding unit (CU), a virtual pipeline data unit (VPDU), a coding / decoding tree unit (CTU), a CTU row, a slice, a picture, a sub-picture, or a region containing more than one sample or pixel.

[0242] Item 97. The method according to any one of Items 1-93, further comprising: determining whether and / or how it is determined whether the OBMC is applied to the current block based on the coded / decoded information of the video unit, the coded / decoded information including at least one of the following: block size, color format, single and / or double tree segmentation, color component, slice type, or picture type.

[0243] Item 98. The method according to any one of Items 1-97, wherein the conversion includes encoding the video unit into the bitstream.

[0244] Item 99. The method according to any one of Items 1-97, wherein the conversion includes decoding the video unit from the bitstream.

[0245] Item 100. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1-99.

[0246] Item 101. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1-99.

[0247] Item 102. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method executed by an apparatus for video processing, wherein the method comprises: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target coding / decoding tool is enabled for the current block; and generating the bitstream based on the determination.

[0248] Item 103. A method for storing a bitstream of a video, comprising: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a prediction mode of a neighboring block, a prediction mode of a reference block, a prediction mode of a first reference block of a second reference block, or whether a target codec tool is enabled for the current block; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium. Example device

[0249] Figure 40 FIG. shows a block diagram of a computing device 4000 in which various embodiments of the present disclosure may be implemented. The computing device 4000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).

[0250] It should be understood that Figure 40 the computing device 4000 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.

[0251] As Figure 40 shown, the computing device 4000 includes a general-purpose computing device 4000. The computing device 4000 may include at least one or more processors or processing units 4010, a memory 4020, a storage unit 4030, one or more communication units 4040, one or more input devices 4050, and one or more output devices 4060.

[0252] In some embodiments, the computing device 4000 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 4000 may support any type of interface to the user (such as a "wearable" circuitry, etc.).

[0253] The processing unit 4010 can be a physical processor or a virtual processor, and can implement various processes based on the programs stored in the memory 4020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4000. The processing unit 4010 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0254] The computing device 4000 generally includes various computer storage media. Such media can be any media accessible by the computing device 4000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4020 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 4030 can be any removable or non-removable media, and can include machine-readable media, such as a memory, a flash drive, a magnetic disk, or other media that can be used to store information and / or data and can be accessed in the computing device 4000.

[0255] The computing device 4000 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 40 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In this case, each drive can be connected to a bus (not shown) via one or more data media interfaces.

[0256] The communication unit 4040 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4000 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Thus, the computing device 4000 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.

[0257] The input device 4050 can be one or more input devices among various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and so on. The output device 4060 can be one or more output devices among various output devices, such as a display, a speaker, a printer, and so on. With the aid of the communication unit 4040, the computing device 4000 can also communicate with one or more external devices (not shown), such as a storage device and a display device, the computing device 4000 can also communicate with one or more devices that enable a user to interact with the computing device 4000, or if needed, the computing device 4000 can also communicate with any device (such as a network card, a modem, etc.) that enables the computing device 4000 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0258] In some embodiments, some or all components of the computing device 4000 can also be arranged in a cloud computing architecture instead of being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which will not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be combined or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or directly or otherwise installed on a client device.

[0259] In an embodiment of the present disclosure, the computing device 4000 can be used to implement video encoding / decoding. The memory 4020 can include one or more video codec modules 4025 having one or more program instructions. These modules are accessible and executable by the processing unit 4010 to perform the functions of various embodiments described herein.

[0260] In an example embodiment of performing video encoding, an input device 4050 may receive video data as an input 4070 to be encoded. The video data may be processed, for example, by a video codec module 4025 to generate an encoded bitstream. The encoded bitstream may be provided as an output 4080 via an output device 4060.

[0261] In an example embodiment of performing video decoding, an input device 4050 may receive an encoded bitstream as an input 4070. The encoded bitstream may be processed, for example, by a video codec module 4025 to generate decoded video data. The decoded video data may be provided as an output 4080 via an output device 4060.

[0262] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Accordingly, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For the conversion between a video unit of a video and the bitstream of the video unit, determining whether overlapping block motion compensation (OBMC) based on at least one of the following is applied to a current block of the video unit: The prediction mode of neighboring blocks, The prediction mode of reference blocks, The prediction mode of a first reference block of a second reference block, or Whether a target codec tool is enabled for the current block; And Performing the conversion based on the determination.

2. The method according to claim 1, wherein the neighboring blocks include at least one of the following: Spatial neighboring blocks adjacent to the current block, Spatial neighboring blocks non - adjacent to the current block, Temporal neighboring blocks adjacent to the current block, or Temporal neighboring blocks non - adjacent to the current block.

3. The method according to claim 2, wherein the current block is inter - frame Merge - coded or the current block is inter - frame advanced motion vector prediction (AMVP) - coded.

4. The method according to claim 2, wherein whether OBMC is applied to the current block is based on whether there are neighboring blocks coded using intra - block copy (IBC).

5. The method according to claim 2, wherein whether OBMC is applied to the current block is based on whether there are neighboring blocks coded using a palette.

6. The method according to claim 2, wherein whether OBMC is applied to the current block is based on whether there are neighboring blocks coded using intra - template matching prediction (intraTMP).

7. The method according to claim 2, wherein whether OBMC is applied to the current block is based on whether there are neighboring blocks coded using block differential pulse codec modulation (BDPCM).

8. The method according to claim 2, wherein whether OBMC is applied to the current block is based on whether there are neighboring blocks coded using transform skip.

9. The method according to any one of claims 1 - 8, wherein the neighboring blocks are adjacent to the current block.

10. The method according to any one of claims 1 - 8, wherein the neighboring blocks are non - adjacent to the current block.

11. The method according to any one of claims 1 - 8, wherein the neighboring blocks are spatial neighboring blocks within the current picture.

12. The method according to any one of claims 1 - 8, wherein the neighboring blocks are temporal blocks in a reference picture.

13. The method according to any one of claims 1 - 8, wherein the neighboring blocks are sub - blocks smaller than the current block.

14. The method according to any one of claims 1 - 8, wherein the neighboring blocks are video units greater than or equal to the current block.

15. The method according to any one of claims 1 - 8, wherein the neighboring blocks are sample positions.

16. The method according to claim 1, wherein the checking process is applied one by one to a series of adjacent neighboring blocks or sub - blocks on the left and / or above the current block to determine whether OBMC is applied to the current block.

17. The method according to claim 16, wherein if there is a neighboring block encoded or decoded using the target mode, the checking process is terminated and it is considered that the OBMC is not applied to the current block.

18. The method according to claim 16, wherein if there is a neighboring block encoded using the inter mode, it is further checked whether the reference block of the neighboring block encoded by the inter mode is encoded using the target mode, and if the reference block is encoded using the target mode, the checking process is terminated and it is considered that the OBMC is not applied to the current block.

19. The method according to claim 18, wherein the reference block is identified by adding the motion vector associated with the neighboring block encoded by the inter mode and the position of the neighboring block encoded by the inter mode.

20. The method according to claim 18, wherein the reference block is in the reference picture.

21. The method according to claim 16, wherein if there is a neighbor encoded using intraTMP, it is further checked whether the reference block of the neighboring block encoded by intraTMP is encoded using the target mode, and if the reference block is encoded using the target mode, the checking process is terminated and it is considered that the OBMC is not applied to the current block.

22. The method according to claim 21, wherein the reference block is identified by adding the block vector associated with the neighboring block encoded by intraTMP and the position of the neighboring block encoded by intraTMP.

23. The method according to claim 21, wherein the reference block is in the current picture.

24. The method according to any one of claims 17 - 23, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

25. The method according to claim 1, wherein the checking process is applied to a series of non - adjacent neighboring blocks or sub - blocks in the already encoded region of the current picture one by one to determine whether the OBMC is applied to the current block.

26. The method according to claim 25, wherein if there is a neighboring block encoded using the target mode, the checking process is terminated and it is considered that the OBMC is not applied to the current block.

27. The method according to claim 25, wherein if there is a neighboring block encoded using the inter mode, it is further checked whether the reference block of the neighboring block encoded by the inter mode is encoded using the target mode, and if the reference block is encoded using the target mode, the checking process is terminated and it is considered that the OBMC is not applied to the current block.

28. The method according to claim 27, wherein the reference block is identified by adding the motion vector associated with the neighboring block encoded by the inter mode and the position of the neighboring block encoded by the inter mode.

29. The method according to claim 27, wherein the reference block is in a reference picture.

30. The method according to claim 25, wherein if there is a neighbor decoded / encoded by intraTMP, it is further checked whether the reference block of the neighbor block decoded / encoded by intraTMP is decoded / encoded using the target mode, and if the reference block is decoded / encoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

31. The method according to claim 30, wherein the reference block is identified by adding a block vector associated with the neighbor block decoded / encoded by intraTMP and the position of the neighbor block decoded / encoded by intraTMP.

32. The method according to claim 30, wherein the reference block is in the current picture.

33. The method according to any one of claims 25 - 32, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

34. The method according to claim 1, wherein the checking process is applied to a series of temporal blocks or sub - blocks in a reference picture one by one to determine whether OBMC is applied to the current block.

35. The method according to claim 34, wherein if there is a temporal block decoded / encoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

36. The method according to claim 34, wherein if there is a temporal block decoded / encoded using an inter - frame mode, it is further checked whether the reference block of the temporal block decoded / encoded using the inter - frame mode is decoded / encoded using the target mode, and if the reference block is decoded / encoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

37. The method according to claim 36, wherein the reference block is identified by adding a motion vector associated with the temporal block decoded / encoded using the inter - frame mode and the position of the temporal block decoded / encoded using the inter - frame mode.

38. The method according to claim 34, wherein if there is a temporal block decoded / encoded by intraTMP, it is further checked whether the reference block of the temporal block decoded / encoded by intraTMP is decoded / encoded using the target mode, and if the reference block is decoded / encoded using the target mode, the checking process is terminated and it is considered that OBMC is not applied to the current block.

39. The method according to claim 38, wherein the reference block is identified by adding a block vector associated with the temporal block decoded / encoded by intraTMP and the position of the temporal block decoded / encoded by intraTMP.

40. The method according to any one of claims 34 - 39, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

41. The method according to claim 1, wherein whether the OBMC is applied to the current block depends on the prediction mode of the reference block.

42. The method according to claim 41, wherein the current block is inter-frame Merge coded, or wherein the current block is inter-frame AMVP coded.

43. The method according to claim 41, wherein the reference block is a block or sub-block identified by adding a displacement to the position of the first block.

44. The method according to claim 43, wherein the first block is the current block.

45. The method according to claim 43, wherein the first block is a neighboring block adjacent to the current block.

46. The method according to claim 43, wherein the first block is a neighboring block non-adjacent to the current block.

47. The method according to claim 43, wherein the first block is a reference block of the current block.

48. The method according to claim 43, wherein the first block is a reference block of a neighboring block.

49. The method according to claim 43, wherein the reference block is identified based on the position of the current block coded in an inter-frame mode and the motion information associated with the current block.

50. The method according to claim 49, wherein the reference block is in a reference picture.

51. The method according to claim 43, wherein the reference block is identified based on the position of the current block coded in an IntraTMP mode and the motion information associated with the current block.

52. The method according to claim 51, wherein the reference block is in the current picture.

53. The method according to claim 43, wherein the reference block is identified based on the position of a neighboring block coded in an inter-frame mode and the motion information associated with the neighboring block coded in the INTER mode.

54. The method according to claim 53, wherein the reference block is in a reference picture.

55. The method according to claim 43, wherein the reference block is identified based on the position of a neighboring block coded in an IntraTMP mode and the motion information associated with the neighboring block coded in the IntraTMP mode.

56. The method according to claim 55, wherein the reference block is in the current picture.

57. The method according to claim 43, wherein the reference block is identified based on the position of a reference block coded in an inter-frame mode and the motion information associated with the reference block coded in the inter-frame mode.

58. The method according to claim 57, wherein the reference block is in another reference picture, rather than in the reference picture where the reference block coded in the inter-frame mode is located.

59. The method according to claim 43, wherein the reference block is identified based on the position of a reference block coded in an IntraTMP mode and the motion information associated with the reference block coded in the IntraTMP mode.

60. The method according to claim 59, wherein the reference block is in the same reference picture as the reference block in which the IntraTMP mode-encoded / decoded reference block is located.

61. The method according to claim 41, wherein the reference block is checked when the neighboring block is encoded / decoded using the INTER mode.

62. The method according to claim 61, wherein if the neighboring block is encoded / decoded using the inter-frame mode, the reference block is identified by adding the motion vector associated with the inter-frame-encoded neighboring block and the position of the inter-frame-encoded neighboring block, and if the reference block is encoded / decoded using the target mode, it is considered that the OBMC is not applied to the current block.

63. The method according to claim 62, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

64. The method according to claim 41, wherein the reference block is checked when the neighboring block is encoded / decoded using the IntraTMP mode.

65. The method according to claim 64, wherein if the neighboring block is encoded / decoded using the IntraTMP mode, the reference block is identified by adding the block vector associated with the IntraTMP-encoded neighboring block and the position of the IntraTMP-encoded neighboring block, and if the reference block is encoded / decoded using the target mode, it is considered that the OBMC is not applied to the current block.

66. The method according to claim 65, wherein the target mode is at least one of the following: IBC mode, Palette mode, BDPCM mode, or Transform skip mode.

67. The method according to claim 1, wherein whether the OBMC is applied to the current block depends on the prediction mode of the first reference block of the second reference block.

68. The method according to claim 67, wherein the current block is inter-frame Merge-encoded or the current block is inter-frame AMVP-encoded.

69. The method according to claim 67, wherein the first reference block of the second reference block is identified by adding the motion vector associated with the second reference block as a reference block encoded / decoded using the inter-frame mode and the position of the second reference block.

70. The method according to claim 67, wherein the first reference block of the second reference block is identified by adding the block vector associated with the second reference block as a reference block encoded / decoded using the IntraTMP mode and the position of the second reference block.

71. The method according to claim 67, wherein the historical prediction mode or the propagated prediction mode is stored in the cache.

72. The method according to claim 71, wherein if a block is encoded or decoded using a target mode, or a reference block of the block is encoded or decoded using the target mode, the target mode is stored in a cache associated with block information, and the block information indicates that the block has historical information or propagated information using the target mode.

73. The method according to claim 72, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

74. The method according to any one of claims 1-73, wherein whether a block in the current picture is encoded or decoded using a target mode is stored in a cache, and wherein the block is one of the following: the current block, the neighboring block, or the reference block.

75. The method according to claim 74, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

76. The method according to claim 74, wherein whether the block is encoded or decoded using IBC or palette is stored using a shared parameter or cache.

77. The method according to claim 76, wherein a single parameter or cache is required to store whether the block is encoded or decoded using IBC or palette.

78. The method according to claim 74, wherein whether the block is encoded or decoded using IBC or palette or intraTMP or BDPCM is stored using separate parameters or caches.

79. The method according to claim 78, wherein multiple parameters or caches are required to store whether the block is encoded or decoded using IBC or palette or intraTMP or BDPCM.

80. The method according to claim 74, wherein whether the block in the current picture is encoded or decoded using a target mode is stored at the granularity of MxM sub-blocks.

81. The method according to claim 80, wherein M is equal to 4 or 8.

82. The method according to any one of claims 1-73, wherein whether at least one of the block or the reference block of the block is encoded or decoded using a target mode is stored in a cache, and wherein the block is one of the following: the current block, the neighboring block, or the second reference block.

83. The method according to claim 82, wherein the target mode is at least one of the following: IBC mode, Palette mode, intraTMP mode, BDPCM mode, or Transform skip mode.

84. The method according to claim 82, wherein if the block or the reference block of the block is encoded or decoded using IBC or palette, a parameter equal to true is stored in the cache.

85. The method according to claim 82, wherein if the block or the reference block of the block is not encoded or decoded using IBC or palette, a parameter equal to false is stored in the cache.

86. The method according to claim 82, wherein whether the block is encoded or decoded using IBC or PLT and whether the reference block of the block is encoded or decoded using IBC or a palette are stored as separate parameters and stored in a separate cache.

87. The method according to claim 82, wherein whether at least one of the block or the reference block of the block is encoded in a target mode is stored at the granularity of MxM sub-blocks.

88. The method according to claim 87, wherein M is equal to 4 or 8.

89. The method according to claim 1, wherein the target encoding / decoding tool includes at least one of the following: IBC, palette, intraTMP, BDPCM, or transform skip.

90. The method according to claim 89, wherein whether the target encoding / decoding tool is applied is controlled by a first syntax element (SE), and / or wherein whether OBMC is applied is controlled by a second syntax element.

91. The method according to claim 90, wherein if the first syntax element indicates that the target encoding / decoding tool is enabled, the second syntax element indicates that the OBMC will be disabled.

92. The method according to claim 90, wherein if the first syntax element indicates that the target encoding / decoding tool is enabled, the second syntax element indicates that the OBMC will be disabled and is set at the encoder.

93. The method according to any one of claims 90-92, wherein the first syntax element is indicated in one of the following: Video Parameter Set (VPS), Sequence Parameter Set (SPS), Picture Parameter Set (PPS), slice header, Coding Tree Unit (CTU), or Coding Unit (CU), or wherein the second syntax element is indicated in one of the following: VPS, SPS, PPS, slice header, CTU, or CU.

94. The method according to any one of claims 1-93, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.

95. The method according to any one of claims 1-93, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated in one of the following: sequence header, picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), slice header, or slice group header.

96. The method according to any one of claims 1-93, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is included in one of the following: prediction block (PB), transformation block (TB), coding block (CB), prediction unit (PU), transformation unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), Coding Tree Unit (CTU), CTU row, slice, tile, sub-picture, or A region containing more than one sample or pixel.

97. The method according to any one of claims 1-93, further comprising: Based on the transcoded information of the video unit, determining whether and / or how to determine whether OBMC is applied to the current block, the transcoded information including at least one of the following: Block size, Color format, Single and / or dual-tree segmentation, Color component, Strip type, or Picture type.

98. The method according to any one of claims 1-97, wherein the conversion includes encoding the video unit into the bitstream.

99. The method according to any one of claims 1-97, wherein the conversion includes decoding the video unit from the bitstream.

100. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-99.

101. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-99.

102. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method executed by an apparatus for video processing, wherein the method includes: Determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: The prediction mode of neighboring blocks, The prediction mode of reference blocks, The prediction mode of a first reference block of a second reference block, or Whether a target codec tool is enabled for the current block; And Generating the bitstream based on the determination.

103. A method for storing a bitstream of a video, comprising: Determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: The prediction mode of neighboring blocks, The prediction mode of reference blocks, The prediction mode of a first reference block of a second reference block, or Whether a target codec tool is enabled for the current block; Generating the bitstream based on the determination; And Storing the bitstream in a non-transitory computer-readable medium.