Method and device for video processing and medium

By determining the OBMC application based on the inheritance of motion vector candidates between the video unit and the bitstream of the video unit, the problem of improving encoding and decoding efficiency in the existing video encoding and decoding technology is solved, and higher encoding and decoding gain and efficiency are achieved.

CN120457686APending Publication Date: 2025-08-08DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089821.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-28
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

There is room for improvement in the encoding and decoding efficiency of existing video encoding and decoding technologies, especially in video compression technologies, especially for video encoding/decoding standards such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264/MPEG-4 AVC, ITU-T H.265 HEVC and multi-function video encoding and decoding (VVC) standards.

Method used

In the conversion between the video unit and the bitstream of the video unit, it is determined whether to apply motion compensation (OBMC) based on overlapping sub-blocks (OBMC) to achieve block-level adaptive OBMC, and improve encoding and decoding efficiency.

Benefits of technology

Improve the codec gain of video encoding and codec and improve the codec efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120457686A_ABST
    Figure CN120457686A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method includes determining, for a conversion between a video unit of a video and a bitstream of the video unit, whether overlapping sub-block based motion compensation (OBMC) is applied to a current block of the video unit based on inheritance from a motion vector candidate; and performing a conversion based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to overlapped sub-block based motion compensation (OBMC) flag inheritance. Background Art

[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is provided. The method includes: determining, for conversion between a video unit and a bitstream of the video unit, whether overlapped sub-block based motion compensation (OBMC) should be applied to a current block of the video unit based on inheritance from motion vector candidates; and performing conversion based on the determination. In this way, block-level adaptive OBMC, which inherits OBMC parameters (e.g., OBMC on / off flags) from neighboring blocks, can achieve higher codec gain and improve codec efficiency.

[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; and generating a bitstream based on the determination.

[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;

[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;

[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;

[0014] Figure 4 Intra prediction mode is shown;

[0015] Figure 5 The reference samples used for wide-angle intra prediction are shown;

[0016] Figure 6 The problem of discontinuity is shown when the orientation exceeds 45°;

[0017] Figure 7A Schematic diagram showing the definition of sample points used by PDPC for diagonal top right mode applied to diagonal and adjacent angle intra modes;

[0018] Figure 7B Schematic diagram showing the definition of sample points used by PDPC in the diagonal bottom left mode applied to the diagonal and adjacent angle intra modes;

[0019] Figure 7C Schematic diagram showing the definition of sample points used by PDPC applied to diagonal and adjacent angle intra modes;

[0020] Figure 7D Schematic diagram showing the definition of sample points used by PDPC for adjacent diagonal bottom left mode applied to diagonal and adjacent angle intra modes;

[0021] Figure 8 An example of four reference rows of neighboring prediction blocks is shown;

[0022] Figure 9A A schematic diagram showing the process of sub-division according to block size;

[0023] Figure 9B A schematic diagram showing the process of sub-division according to block size;

[0024] Figure 10 shows the matrix-weighted intra prediction process;

[0025] Figure 11 The spatial GPM candidates are shown;

[0026] Figure 12 A GPM template is shown;

[0027] Figure 13 GPM mixing is shown;

[0028] Figure 14 Shows the location of the spatial merge candidate;

[0029] Figure 15 shows candidate pairs considered for redundancy check of spatial merge candidates;

[0030] Figure 16 A diagram showing motion vector scaling for temporal Merge candidates is shown;

[0031] Figure 17 The candidate positions for the time domain Merge candidates C0 and C1 are shown;

[0032] Figure 18 The MMVD search point is shown;

[0033] Figure 19 shows the extended CU area used in BDOF;

[0034] Figure 20 A diagram for a symmetric MVD pattern is shown;

[0035] Figure 21 Decoding side motion vector refinement is shown;

[0036] Figure 22 The top and left neighboring blocks used in the derivation of CIIP weights are shown;

[0037] Figure 23An example of GPM partitioning grouped at the same angle is shown;

[0038] Figure 24 Unidirectional prediction MV selection for geometric partitioning mode is shown;

[0039] Figure 25 shows an exemplary generation of warp weights w0 using a geometric segmentation pattern;

[0040] Figure 26 The current CTU processing order and its available reference samples in the current and left CTUs are shown;

[0041] Figure 27 The residual encoding and decoding passes for transforming skipped blocks are shown;

[0042] Figure 28 An example of a block encoded and decoded in palette mode is shown;

[0043] Figure 29 Sub-block based index map scanning for a palette is shown, on the left for horizontal scanning and on the right for vertical scanning;

[0044] Figure 30 Shown is a decoding flow chart with ACT;

[0045] Figure 31 The intra-frame template matching search area used is shown;

[0046] Figure 32 Five locations in the reconstructed luminance samples are shown;

[0047] Figure 33 The prediction process of DBV mode is shown;

[0048] Figure 34 The low-frequency non-separable transform (LFNST) process is shown;

[0049] Figure 35 The SBT position, type and transformation type are shown;

[0050] Figure 36 The ROI for LFNST16 is shown;

[0051] Figure 37 The ROI for LFNST8 is shown;

[0052] Figure 38 Discontinuity measurements are shown;

[0053] Figure 39 A flowchart showing a method for video processing according to an embodiment of the present disclosure is shown; and

[0054] Figure 40 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.

[0055] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION

[0056] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.

[0057] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0058] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.

[0059] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0060] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment

[0061] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0062] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.

[0063] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.

[0064] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.

[0065] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.

[0066] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.

[0067] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0068] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0069] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0070] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.

[0071] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0072] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0073] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.

[0074] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.

[0075] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0076] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0077] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.

[0078] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0079] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0080] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0081] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0082] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0083] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.

[0084] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.

[0085] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0086] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0087] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0088] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0089] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.

[0090] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0091] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.

[0092] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.

[0093] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.

[0094] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.

[0095] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.

[0096] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0097] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0098] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present disclosure relates to video coding and decoding technologies. Specifically, it relates to overlapping sub-block motion compensation (OBMC) and related technologies in image / video coding and decoding. It can be applied to existing video coding and decoding standards such as HEVC and VVC. It is also applicable to future video coding and decoding standards or video codecs. 2. Introduction Video codec standards evolve primarily through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly. The new video codec standard was officially named the Versatile Video Codec (VVC) at the April 2018 JVET meeting, and the first version of the VVC Test Model (VTM) was released at that time. The VVC working draft and the VTM test model are updated after each meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. 2.1. Existing Codec Tools Intra-frame prediction 2.1.1.1. Intra-mode codec with 67 intra-prediction modes To capture arbitrary edge directions present in natural videos, the number of directional intra modes in VVC is extended to 65 from the 33 used in HEVC. Figure 4 New directional modes not in HEVC are depicted as red dashed arrows, and planar and DC modes remain unchanged. These more dense directional intra prediction modes apply to all block sizes and both luma and chroma intra prediction. In VVC, for non-square blocks, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid division operations for DC prediction, only the longer sides are used to calculate the average value of non-square blocks. 2.1.1.2. Intra-mode encoding and decoding In order to keep the complexity of the most probable mode (MPM) list generation low, a 6MPM intra mode encoding and decoding method is adopted considering two available adjacent intra modes. The following three aspects are considered to construct the MPM list: – Default intra mode; – Neighborhood intra mode; – Derived intra mode. Regardless of whether MRL and ISP codecs are applied, a unified 6-MPM list is used for intra blocks. The MPM list is constructed based on the intra modes of the neighboring blocks on the left and above. Assuming the mode on the left is denoted as Left and the mode of the block above is denoted as Above, the unified MPM list is constructed as follows: – When a neighboring block is not available, its intra mode is set to planar by default. – If both Left and Above modes are non-angle modes: –MPM list → {plane, DC, V, H, V-4, V+4}. – If one of the Left and Above modes is an angular mode and the other is a non-angular mode: – Set mode Max as the larger mode in Left and Above –MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2}. – If Left and Above are both angles, and they are different: – Set mode Max as the larger mode in Left and Above – If the difference between Left mode and Above mode is in the range of 2 to 62 (inclusive) – MPM list → {Plane, Left, Above, DC, Max-1, Max+1} -otherwise – MPM list → {Flat, Left, Above, DC, Max-2, Max+2}. – If Left and Above are both angles, and they are the same: – MPM list → {plane, left, left-1, left+1, DC, left-2}. In addition, the first binary bit of the mpm index codeword is the CABAC context codec. A total of three contexts are used, corresponding to whether the current intra block is MRL-enabled, ISP-enabled, or a regular intra block. During the 6MPM list generation process, deduplication is used to remove repeated patterns so that only unique patterns can be included in the MPM list. For entropy coding of the 61 non-MPM modes, truncated binary code (TBC) is used. 2.1.1.3. Wide-angle intra prediction for non-square blocks Conventional angular intra prediction directions are defined as 45 degrees to -135 degrees clockwise. In VVC, for non-square blocks, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the wide-angle mode index after parsing. The total number of intra prediction modes remains unchanged at 67, and the intra mode encoding and decoding method remains unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are provided as Figure 5 As shown, it is defined. The number of alternative modes in the wide-angle direction mode depends on the aspect ratio of the block. The alternative intra prediction modes are shown in Table 1. Table 1 - Intra prediction modes replaced by Wide mode like Figure 6 As shown in Figure 2, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the increased gap Δp. α negative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, and the 8 modes are [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted using these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the non-fractional modes in the traditional prediction mode and the wide-angle mode. In VVC, 4:2:2 and 4:4:4 as well as 4:2:0 chroma formats are supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, which expanded the number of entries from 35 to 67 to keep pace with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the mapping table entries to more accurately convert the prediction angles of the chroma blocks. 2.1.1.4. Mode-dependent intra-frame smoothing (MDIS) A four-tap intra-frame interpolation filter is used to improve the accuracy of directional intra-frame prediction. In HEVC, a two-tap linear interpolation filter is used to generate intra-frame prediction blocks in directional prediction modes (i.e., excluding planar and DC prediction values). In VVC, a simplified 6-bit 4-tap Gaussian interpolation filter is used only for directional intra-frame modes. The non-directional intra-frame prediction process is unchanged. The selection of the 4-tap filter is performed according to the MDIS condition for providing non-fractional shifted directional intra-frame prediction modes, i.e., all directional modes excluding the following items: 2, HOR_IDX, DIA_IDX, VER_IDX, 66. Depending on the intra prediction mode, the following reference sample processing is performed: – Directional intra prediction modes are classified into one of the following groups: – vertical mode or horizontal mode (HOR_IDX, VER_IDX), – Diagonal mode, indicating that the angle is a multiple of 45 degrees (2, DIA_IDX, VDIA_IDX), – Other directional modes; – If the directional intra prediction mode is classified as belonging to group A, no filter is applied to the reference samples to generate the predicted samples; – Otherwise, if the mode belongs to group B, the [1,2,1] reference sample filter may be applied (depending on the MDIS condition) to the reference samples to further copy these filtered values into the intra prediction values according to the selected direction, but no interpolation filter is applied; Otherwise, if the mode is classified as belonging to group C, only the intra reference sample interpolation filter is applied to the reference samples to generate predicted samples at fractional or integer positions that fall between the reference samples according to the selected direction (no reference sample filtering is performed). 2.1.1.5. Position-dependent intra prediction combination In VVC, the results of intra prediction for DC, planar, and several angular modes are further modified by the Position-Dependent Intra Prediction Combining (PDPC) method. PDPC is an intra prediction method that uses a combination of HEVC-style intra prediction without filtering boundary reference samples and with filtering boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, bottom-left angular mode and its eight neighboring angular modes, and top-right angular mode and its eight neighboring angular modes. The prediction sample pred(x',y') is predicted using a linear combination of the intra prediction mode (DC, planar, angular) and the reference sample according to the following equation 3-8: pred(x',y')=(wL×R -1,y’ +wT×Rx’,-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x',y')+32)>>6 (2-1) where R x,-1 ,R -1,y Respectively represent the reference sample points at the top and left boundaries of the current sample point (x, y), and R -1,-1 Indicates the reference sample located at the upper left corner of the current block. If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filters are required as in the DC mode boundary filters or horizontal / vertical mode edge filters in HEVC. The PDPC process for DC mode and planar mode is the same, and clipping operations are avoided. For angular mode, the PDPC scaling factor is adjusted so that no range check is required, and the angle condition for enabling PDPC is removed (scaling >= 0 is used). In addition, the PDPC weights are based on 32 in all angular mode cases. The PDPC weights depend on the prediction mode, as shown in Table 2. PDPC is applied to blocks with width and height both greater than or equal to 4. Figures 7A-7D The reference sample points (R x,-1 ,R -1,y and R -1,-1 ). The prediction sample point pred(x',y') is located at (x',y') in the prediction block. For example, for the diagonal mode, the reference sample point R x,-1 The coordinate x is given by: x = x' + y' + 1, and the reference point R -1,y The coordinate y of is similarly given by: y = x' + y' + 1. For other annular patterns, the reference point R x,-1 and R -1,y Can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position is used. Table 2 - Example of PDPC weights according to prediction mode 2.1.1.6. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 8 In

[15] , an example of 4 reference lines is depicted, where the samples of segments A and F are not extracted from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines are used (reference line 1 and reference line 3). The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row indices greater than 0, only the additional reference row modes are included in the MPM list, and only the mpm index is signaled without the remaining modes. The reference row index is signaled before the intra prediction mode, and if a non-zero reference row index is signaled, the intra prediction mode does not include planar mode. MRL is disabled for the first row of a block within a CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC will be disabled when additional rows are used. For MRL mode, the derivation of DC values in DC intra prediction mode for non-zero reference row indices is aligned with the derivation for reference row index 0. MRL requires storage of the CTU and 3 adjacent luma reference rows to generate the prediction. The Cross Component Linear Model (CCLM) tool also requires 3 adjacent luma reference rows for its downsampling filter. The definition of MRL using the same 3 rows is aligned with CCLM to reduce the storage requirements of the decoder. 2.1.1.7. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size of ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into four sub-partitions. It can be noted that M×128 (M≤64) and 128×N (N≤64) ISP blocks may generate potential problems for 64×64 VDPU. For example, an M×128 CU in the single-tree case has M×128 luma TBs and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into 4 M×32TBs (only horizontal division is possible), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divisible. Therefore, the size of both chroma components will be larger than a 32×32 block. Similarly, a similar situation can be created using a 128×NCU with ISP. Therefore, these two situations are problems for a 64×64 decoder pipeline. Therefore, the CU size that can use ISP is limited to a maximum of 64×64. Figure 9A and Figure 9B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples. In ISP, 1xN / 2xN sub-block predictions are not allowed to rely on reconstructed values from previously decoded 1xN / 2xN sub-blocks of the codec block, resulting in a minimum prediction width of four samples for each sub-block. For example, an 8xN (N>4) codec block encoded using ISP with vertical partitioning is partitioned into two prediction regions of size 4xN and four transforms of size 2xN. Furthermore, a 4xN codec block encoded using ISP with vertical partitioning is predicted using a full 4xN block; each of the four transforms is used. While 1xN and 2xN transform sizes are allowed, it is asserted that transforms for these blocks in the 4xN region can be performed in parallel. For example, when a 4xN prediction region contains four 1xN transforms, no transform is performed horizontally; the transform in the vertical direction can be performed as a single 4xN transform in the vertical direction. Similarly, when a 4xN prediction region contains two 2xN transform blocks, the transform operations for the two 2xN blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, processing these smaller blocks does not increase latency compared to processing intra blocks for 4x4 regular codecs. Table 3 - Entropy codec coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization and inverse transformation. Therefore, the reconstructed sample values of each sub-partition can be used to generate the prediction of the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the upper left sample containing the CU, and then continues to the sub-partition downward (horizontal partitioning) or to the right (vertical partitioning). Therefore, the reference samples used to generate the sub-partition prediction signal are only located on the left and upper sides of the row. All sub-partitions share the same intra-frame mode. The following is a summary of the interaction of ISP with other codec tools. - Multiple Reference Line (MRL): If the MRL index of a block is not 0, the ISP codec mode will be inferred to be 0, and thus the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: As shown in Table 3, the size of the entropy coding sub-block has been modified to have 16 samples in all possible cases. Note that the new size only affects blocks generated by ISP where one dimension is less than 4 samples. In all other cases, the coefficient group remains 4×4. –CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions and the first n-1 subpartitions have yielded zero CBF, the CBF of the nth subpartition is assumed to be 1. –MPM usage: The MPM flag will be inferred as the one in blocks coded by ISP mode, and the MPM list is modified to exclude DC mode and give priority to horizontal intra mode for ISP horizontal partitioning and vertical intra mode for ISP vertical partitioning. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. –PDPC: When the CU uses ISP codec mode, the PDPC filter will not be applied to the final sub-split. –MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests for the different available transforms for each final sub-split. The transform selection for ISP mode will instead be fixed and will be selected based on the intra mode used, processing order and block size. Therefore, no signaling is required. For example, let t H and t V The horizontal and vertical transforms are selected for the w×h sub-partition, where w is the width and h is the height. The transforms are then selected according to the following rules: If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w=2 or w>32, t H =DCT-II – If h=2 or h>32, t V =DCT-II – Otherwise, the transformation is selected as shown in Table 4. Table 4 - Transform selection depends on intra mode In ISP mode, all 67 intra modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the conditions for intra interpolation filter selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.1.1.8. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a new intra prediction technique added to VVC. To predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, they are generated as in traditional intra prediction. Figure 10 As shown, the generation of the prediction signal is based on the following three steps, which are averaging, matrix-vector multiplication and linear interpolation. ●Average of neighboring points Among the boundary samples, four samples or eight samples are selected by averaging based on the block size and shape. Specifically, input boundary bdry top and bdry left The adjacent boundary samples are reduced to smaller boundaries by averaging them according to predefined rules that depend on the block size. and Then, the two reduced boundaries and is spliced to the reduced boundary vector bdry red , so the size of the boundary vector for block reduction of shape 4×4 is four, and the size of the boundary vector for block reduction of all other shapes is eight. If mode refers to a MIP mode, then this splicing is defined as follows: Matrix multiplication The averaged samples are taken as input, matrix-vector multiplication is performed, and then the offset is added. The result is a reduced prediction signal on the downsampled set of samples in the original block. From the reduced input vector bdry red The reduced prediction signal pred red, is generated, the reduced prediction signal pred red is the width W red and height H red Here, W red and H red is defined as: Reduced prediction signal pred red It is calculated by taking the matrix-vector product and adding the offset: pred red =A·bdry red +b. Here, A is a matrix with W red ·H red rows, and if W=H=4, then it has 4 columns, in all other cases, it has 8 columns. b is the size W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, S2. The index idx = idx(W,H) is defined as follows: Here, each coefficient of the matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices Each matrix has 16 rows and 4 columns, and 16 offset vectors Each offset vector has a size of 16. The set of matrices and offset vectors is used for blocks of size 4×4. Set S1 consists of 8 matrices Each matrix has 16 rows and 8 columns, and 8 offset vectors The size of each offset vector is 16. Set S2 consists of 6 matrices Each matrix has 64 rows and 8 columns, and 6 offset vectors of size 64 Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the downsampled set by linear interpolation, which is a single-step linear interpolation in each direction. Interpolation is performed first in the horizontal direction and then in the vertical direction, regardless of block shape or block size. ●MIP mode signaling and coordination with other codec tools For each codec unit (CU) in intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transposed flag (isTransposed) that determines whether the mode is transposed, and the MIP mode Id (modeId) that determines which matrix to use for a given MIP mode are derived as follows: isTransposed=predModeIntra&1 modeId=predModeIntra>>1(2-6). The MIP codec mode is coordinated with other codec tools by taking into account the following aspects: – LFNST is enabled for MIPs on large blocks. Here the planar LFNST transform is used. – Reference sample derivation for MIP is performed like for conventional intra prediction modes. – For the downsampling step used in MIP prediction, the original reference samples are used instead of the downsampled samples. – Clipping is performed before upsampling rather than after upsampling. – Regardless of the maximum transform size, MIP is allowed to be up to 64×64. - The number of MIP modes is 32 for sizeld=0, 16 for sizeld=1, and 12 for sizeld=2. 2.1.1.9. Airspace GPM (SGPM) In the spatial domain GPM, a candidate list including partitioning and two intra prediction modes is constructed. Up to 11 intra prediction modes of MPM are used to form a combination, and the length of the candidate list is set to be equal to 16. The selected candidate index is transmitted through the signal. List Usage Figure 11 The templates shown in are re-ranked. The GPM blending process is not used in the templates, and the SAD between the prediction of the template and the reconstruction of the template is used for ranking. Figure 12 A GPM template is shown. Figure 13 GPM segmentation boundaries are shown. The SGPM mode is applied to blocks whose width and height satisfy the same restrictions as in inter-frame GPM. The following projects are considered: ●Airspace GPM segmentation mode: 26 predefined modes Adaptive inference algorithm based on horizontal and vertical gradient ratio ●Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation mode, an IPM list is derived for each part using intra-inter GPM list derivation. The IPM list size is 3. In the list, TIMD derived patterns are replaced by 2 derived patterns with horizontal and vertical orientations (using top or left template) or TIMD derived patterns are excluded. MPM list: A unified MPM list, with up to 11 elements, is used for all segmentation modes. Template size (left and top): 1 or 4 ●Extended block size: Spatial GPM is extended to be further applied to 4x8, 8x4, 4x16 and 16x4 blocks, which can be described as 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. Adaptive Hybrid: Adaptive mixing is tested for spatial GPM, where the mixing depth τ is derived as follows: ■If min(width, height) == 4, then 1 / 2τ is selected ■ Otherwise, if min(width, height) == 8, then τ is selected ■ Otherwise, if min(width, height) == 16, then 2τ is selected ■ Otherwise, if min(width, height) == 32, then 4τ is selected ■Otherwise, 8τ is selected. 2.1.2. Inter-frame prediction For each inter-predicted CU, the motion parameters include motion vector, reference picture index and reference picture list usage index, as well as additional information required by the new coding features of VVC that will be used for inter-prediction sample generation. The motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no encoded and decoded motion vector increments or reference picture indices. Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, which include spatial candidates and temporal candidates, as well as additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, as well as other required information, are explicitly signaled for each CU. In addition to the inter-frame coding and decoding functions in HEVC, VVC also includes some new and refined inter-frame prediction coding and decoding tools, as listed below: – Extended Merge Forecast –Merge mode with MVD (MMVD) – Symmetrical MVD (SMVD) signaling – Affine motion compensated prediction – Sub-block based temporal motion vector prediction (SbTMVP) – Adaptive Motion Vector Resolution (AMVR) – Motion field storage: 1 / 16 luminance sample MV storage and 8x8 motion field compression – Bidirectional prediction with CU-level weights (BCW) – Bidirectional Optical Flow (BDOF) –Decoder-side motion vector refinement (DMVR) – Geometric Partitioning Mode (GPM) – Joint Intra-frame and Inter-frame Prediction (CIIP). Details of those inter prediction methods specified in VVC are provided below. 2.1.2.1. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Airspace MVP from airspace adjacent CU 2) Temporal MVP from the same CU 3) History-based MVP from FIFO table 4) Paired Average MVP 5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU code in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded with context, while bypass coding is used for the remaining bins. The derivation process of Merge candidates for each category is provided in this section. As in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a certain size. 2.1.2.1.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Figure 14 Among the candidates at the positions shown, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1 and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1 and A1 are not available (for example, because it belongs to another slice or piece) or is intra-coded. After the candidate at position A1 is added, a redundancy check is performed on the addition of the remaining candidates. The redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 15 The pairs are linked by arrows in , and a candidate is added to the list only if the corresponding candidates used for redundancy check do not have the same motion information. 2.1.2.1.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. In particular, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Figure 16 As shown by the dotted line in , the scaled motion vector for the temporal merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distance, tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. like Figure 17As shown, the position of the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate. 2.1.2.1.3. History-based Merge Candidate Derivation History-based MVP (HMVP) Merge candidates are added to the Merge list after spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-subblock inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The last few HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. Redundancy check is applied to HMVP candidates to spatial or temporal Merge candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N<=4) × M: (8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process from HMVP is terminated. 2.1.2.1.4. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the number represents the merge index of the merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, the list remains invalid. When the Merge list is not full after pairwise average Merge candidates are added, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached. 2.1.2.2.Merge Estimation Region Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). For the generation of the Merge candidate list for the current CU, candidate blocks in the same MER as the current CU are not included. In addition, the update process for the history-based motion vector prediction candidate list is updated only when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2parMrglevel is greater than (yCb>>Log2ParMrgLevel), and where (xCb, yCb) is the top left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.1.2.3. Merge Mode with MVD (MMVD) In addition to the Merge mode where the implicitly derived motion information is directly used for prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the Skip flag and Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it is further refined using signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected to be used as the MV basis. The merge candidate flag is signaled to specify which candidate is used. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 18 As shown in Table 5, the offset is added to the horizontal component or vertical component of the starting MV. Table 5 - Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions, as shown in Table 6. Note that the meaning of the MVD symbol may vary depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV, where both lists point to the same side of the current picture (i.e., both referenced POCs are greater than the POC of the current picture or both are less than the POC of the current picture), the symbol in Table 6 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV, where both MVs point to different sides of the current picture (i.e., one referenced POC is greater than the POC of the current picture and the other referenced POC is less than the POC of the current picture), the symbol in Table 6 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 6 — Sign of MV offset specified by direction index Direction Index 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + – 2.1.2.4. Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, a bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-7) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CU, the weight index is transmitted by signal after the motion vector difference; 2) For Merge CU, the weight index is inferred from the neighboring block based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing encoder complexity. These algorithms are summarized below. When combined with AMVR, the unequal weighting of 1-pixel and 4-pixel motion vector accuracy is only conditionally checked if the current picture is a low-latency picture. When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Unequal weights are not searched when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using a context-coded bit followed by a bypass-coded bit. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficient encoding and decoding of video content in fading conditions. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be transmitted via signal for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not transmitted via signal, and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index of a CU using constructed affine merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded or decoded in CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.1.2.5. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, formerly known as BIO, is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met: – The CU is coded using “true” bi-prediction mode, i.e. one of the two reference pictures precedes the current picture in display order, and the other of the two reference pictures follows the current picture in display order; – The distances from the two reference pictures to the current picture (i.e., POC difference) are the same; – Both reference images are short-term reference images; –CU is not encoded or decoded using Affine mode or ATMVP Merge mode; –CU has more than 64 luminance samples; – CU height and CU width are both greater than or equal to 8 luminance samples; – BCW weight index indicates equal weight; –WP is not enabled for the current CU; –CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As its name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 prediction samples and the L1 prediction samples. Motion refinement is then used to adjust the sample values of the bidirectional prediction in the 4x4 sub-block. The following steps are applied in the BDOF process. First, by directly calculating the difference between two adjacent samples, the horizontal gradient and vertical gradient of the two predicted signals, and is calculated, i.e., Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k, k=0, 1, and shift1 is calculated based on the luma bit depth bitDepth as shift1=max(6, bitDepth-6). Then, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as follows: in, where Ω is a 6×6 window around a 4×4 sub-block, and n a and n b The values of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8), respectively. Then use the following method, motion refinement (v x ,v y ) is derived using the cross-correlation and autocorrelation terms: in th′ BIO =2 max(5,BD-7) . is the round-down (floor) function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4×4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: predBDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )>>shift (2-13). These values are chosen so that the multiplier in the BDOF process does not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, some predicted sample points I outside the current CU boundary in list k (k = 0, 1) (k) (i,j) needs to be generated. Figure 19 As depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and a conventional 8-tap motion compensated interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculations. For the rest of the steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, they are filled (i.e. repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W represents the sub-block width and H represents the sub-block height. In order to avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded in symmetric MVD mode or CIIP mode. 2.1.2.6. Symmetric MVD encoding and decoding In VVC, in addition to the conventional unidirectional prediction and bidirectional prediction mode MVD signal transmission, a symmetric MVD mode for bidirectional prediction MVD signal transmission is applied. In symmetric MVD mode, motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not transmitted through the signal but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. Figure 20 This is an illustration for symmetric MVD mode. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The resulting motion vector is shown below. In the encoder, symmetric MVD motion estimation starts with an initial MV estimate. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The one with the lowest distortion cost is selected as the initial MV for symmetric MVD motion search. 2.1.2.7. Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of Merge mode MV, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In the bidirectional prediction operation, the refined MV is searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. Figure 21As shown, based on each MV candidate around the initial MV, the SAD between the red blocks is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs that are coded or decoded using the following modes and features: – CU-level Merge mode with bi-predictive MV – One reference image is in the past and the other is in the future relative to the current image – The distance from the two reference images to the current image (i.e., POC difference) is the same – Both reference images are short-term reference images –CU has more than 64 luminance samples –CU height and CU width are both greater than or equal to 8 luminance samples –BCW weight index indicates equal weight –WP is not enabled for the current block – CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction for future picture encoding and decoding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.1.2.7.1. Search Scheme In DVMR, the search point is centered around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point represented by a candidate MV pair (MV0, MV1) and examined by DMVR follows the following two equations: MV0′=MV0+MV_offset (2-15) MV1′=MV1-MV_offset (2-16) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample phase of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation rather than performing an additional search using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase ends with the center with the minimum SAD in either the first or second iteration of the search. In the sub-pixel offset estimation based on the parameter error surface, the center position cost and the costs of the four neighboring positions from the center are used to fit the two-dimensional parabolic error surface equation of the following form E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (2-17) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-18) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-19). x min and y min The value of is automatically clamped between -8 and 8, since all cost values are positive, and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16-pixel MV precision in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain a sub-pixel accurate refinement delta MV. 2.1.2.7.2. Bilinear interpolation and sample filling In VVC, the resolution of the video image (MV) is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the original fractional pixel MV with integer sample offsets, so these fractional samples need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the DMVR search process. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples within the 2-sample search range compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To prevent the normal motion compensation process from accessing more reference samples, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV are filled from the available samples. 2.1.2.7.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.1.2.8. Joint Intra-Frame and Inter-Frame Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is transmitted by signal to indicate whether the intra / inter joint prediction (CIIP) mode is applied to the current CU. Figure 22 The top and left neighboring blocks used in the CIIP weight derivation are shown. As its name suggests, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter is derived using the same inter-frame prediction process applied to the conventional Merge mode; and the intra-frame prediction signal P is derived after the conventional intra-frame prediction process of the planar mode. intra Then, the intra and inter prediction signals are combined using a weighted average, where the weight values depend on the coding mode of the top and left neighboring blocks and are calculated as follows: – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntratop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntralLeft to 0; – If (isIntraLeft + isIntraTop) is equal to 2, then wt is set to 3; – Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, wt is set to 2; – Otherwise, set wt to 1. The CIIP forecast is constructed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (2-20). 2.1.2.9. Multiple Hypothesis Prediction (MHP) In inter-AMVP mode, normal Merge mode and MMVD mode, up to two additional prediction values are signaled. The final overall prediction signal is iteratively accumulated using each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. 2.1.2.10. Overlapped Sub-Block Motion Compensation (OBMC) When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information of neighboring blocks and weighted prediction. The conditions under which OBMC should not be applied are as follows: When OBMC is disabled at the SPS level ●When the current block has intra mode or IBC mode ●When the current block applies LIC ●When the current luminance block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. It is enabled for sub-block based codecs: ●Affine AMVP mode; Affine Merge mode and sub-block based temporal motion vector prediction (SbTMVP); ●Sub-block based bilateral matching. 2.1.2.11. Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its prediction block based on the local illumination variation between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, which form a linear equation, i.e., α*p[x]+β, to compensate for illumination variation, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When surround motion compensation is enabled, the MV must be clipped to account for the surround offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except that the LIC flag is signaled for AMVP mode to indicate the use of LIC. The local illumination compensation proposed in JVET-O0066 is used for unidirectionally predicted inter CUs with the following modifications. ● Neighboring samples within the frame can be used in LIC parameter derivation; LIC is disabled for blocks with fewer than 32 luma samples; For both non-subblock and affine modes, LIC parameter derivation is performed based on the template block samples corresponding to the current CU instead of the partial template block samples corresponding to the first top-left 16x16 unit; • The samples of the reference block template are generated by using MC with the block MV without rounding to integer pixel precision. 2.1.2.12. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter-frame prediction. The geometric partitioning mode is transmitted through the signal using the CU level flag as a Merge mode. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode and sub-block Merge mode. For each possible CU size w×h=2 m ×2 n , where m,n∈{3…6} and excluding 8x64 and 64x8, the geometric segmentation mode supports a total of 64 segmentations. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line (e.g. Figure 23 (as shown in Figure 2). The position of the dividing line is mathematically derived from the angle and offset parameters for the specific partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has a motion vector and a reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, only two motion-compensated predictions are required for each CU. If geometric partitioning mode is used for the current CU, a geometric partitioning index indicating the partitioning mode (angle and offset) of the geometric partitioning and two Merge indices (one for each partition) are further transmitted through the signal. The number of maximum GPM candidate sizes is explicitly transmitted through the signal in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, the sample values along the edges of the geometric partitioning are adjusted using a blending process with adaptive weights. This is the prediction signal for the entire CU, and like in other prediction modes, the transform and quantization process will be applied to the entire CU. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.2.12.1. One-way prediction candidate list construction The unidirectional prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode, where X is equal to the parity of n. These motion vectors are Figure 24 In the case where the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used instead as the unidirectional prediction motion vector for the geometric partition mode. 2.1.2.12.2. Blending Along Geometric Partition Edges Figure 25 An example generation of warping weights w0 using geometric partitioning mode is shown. After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edge. The blending weight for each position of the CU is derived based on the distance between the individual position and the partition edge. The distance from position (x,y) to the segmentation edge is derived as: where i,j are the angle and offset indices of the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j The sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-25) w1(x,y)=1-w0(x,y) (2-27). partIdx depends on the angle index i. An example of weight w0 is shown below. 2.1.2.12.3. Motion Field Storage for Geometric Partitioning Mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-28) where motionIdx is equal to d(4x+2,4y+2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.1.2.12.4. GPM with Inter and Intra Prediction (GPM Inter-Intra) In GPM inter-intra, in addition to the Merge candidates for each non-rectangular partitioned area in the CU to which GPM is applied, a predefined intra prediction mode for the geometric partition line can be selected. In the proposed method, for each GPM separation area, a flag from the encoder is used to determine whether it is intra prediction mode or inter prediction mode. When it is inter prediction mode, a unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, when it is intra prediction mode, the unidirectional prediction signal is generated from the neighboring pixels used for the intra prediction mode, which is specified by the index from the encoder. The variation of possible intra prediction modes is limited by the geometric shape. Finally, the two unidirectional prediction signals are mixed in the same way as ordinary GPM. 2.1.3. Screen content encoding and decoding tools 2.1.3.1. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture. The luminance block vector of the CU encoded and decoded by IBC is integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode in addition to the intra prediction mode or inter prediction mode. The IBC mode is applicable to CUs with a width and height that are less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4×4 sub-block. For a current block of larger size, the hash key is determined to match the hash key of the reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, which can be signaled in IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if IBC is used). When either neighbor is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.1.3.1.1.IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined area, which includes the area of the current CTU and some areas of the left CTU. Figure 26 The reference area of the IBC mode is shown, where each block represents a 64x64 luma sample unit. Depending on the location of the current codec CU position within the current CTU, the following applies: – If the current block falls into the upper left 64x64 block of the current CTU, in addition to the reconstructed samples in the current CTU, it can also use the CPR mode to refer to the reference samples in the lower right 64x64 block of the left CTU. The current block can also use the CPR mode to refer to the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU. – If the current block falls into the upper right 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (0, 64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls into the lower left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU. Otherwise, the current block can also use the CPR mode to refer to the reference samples in the lower right 64x64 block of the left CTU. If the current block falls into the lower right 64x64 block of the current CTU, it can only use the CPR mode to refer to the samples that have been reconstructed in the current CTU. This restriction allows the IBC mode to be implemented using local on-chip memory for hardware implementation. 2.1.3.1.2.Interaction between IBC and other codecs The interaction between IBC mode and other inter-frame codec tools in VVC (such as paired merge candidates, history-based motion vector predictor (HMVP), intra-frame inter-frame joint prediction (CIIP), merge mode with motion vector difference (MMVD) and geometric partition mode (GPM)) is as follows: – IBC can be used with paired merge candidates and HMVP. A new paired IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, IBC motion is inserted into the history buffer for future reference. – IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM. – When DUAL_TREE partitioning is used, IBC is not allowed for chroma codec blocks. Unlike the HEVC screen content codec extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation process of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects are applied: – IBC shares the same process as in regular MV Merge, including the use of paired Merge candidates and history-based motion prediction values, but TMVP and zero vectors are not allowed because they are invalid for IBC mode. – Separate HMVP caches (5 candidates each) are used for traditional MV and IBC. – Block vector constraints are implemented as bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream, and if the merge candidate is invalid (out of range or 0), the merge shall not be used. Such bitstream consistency constraints are expressed in terms of a virtual cache as follows. – For deblocking, IBC is handled as inter mode. If the current block is coded using IBC prediction mode, AMVR does not use quarter pels; instead, AMVR is signaled to only indicate whether the MV is inter-pel or 4-integer-pel. - The number of IBC Merge candidates may be signaled in the slice header separately from the number of regular, sub-block, and geometry Merge candidates. The concept of a virtual buffer is used to describe the allowed reference regions and valid block vectors for IBC prediction mode. Denoting the CTU size as ctbSize, the virtual buffer ibcBuf has a width of wIbcBuf = 128x128 / ctbSize and a height of hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows. – At the beginning of decoding each CTU line, flush the entire ibcBuf with an invalid value of -1. – At the start of decoding the VPDU (xVPDU, yVPDU) relative to the upper left corner of the picture, set ibcBuf[x][y]=-1, where x=xVPDU%wIbcBuf,...,xVPDU%wIbcBuf+W v -1; y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1. – After decoding the CU containing (x, y) relative to the top left corner of the picture, set ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]. For a block covering coordinates (x, y), if the following is true for the block vector bv = (bv[v[0], bv[1]), then it is valid; otherwise, it is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1. 2.1.3.2. Block Differential Pulse Code Modulation (BDPCM) VVC supports Block Differential Pulse Code Modulation (BDPCM) for screen content encoding and decoding. At the sequence level, the BDPCM enable flag is signaled in the SPS; this flag is signaled only when transform skip mode (described in the next section) is enabled in the SPS. When BDPCM is enabled, a flag is sent at the CU level if the CU size in terms of luma samples is less than or equal to MaxTsSize multiplied by MaxTsSize and if the CU is intra coded, where MaxTsSize is the maximum block size for which transform skip mode is allowed. This flag indicates whether regular intra codec is used or BDPCM is used. If BDPCM is used, a BDPCM prediction direction flag is sent to indicate whether the prediction is horizontal or vertical. The block is then predicted using a regular horizontal or vertical intra prediction process that samples unfiltered reference samples. The residuals are quantized and the difference between each quantized residual and its predicted value (i.e., the previously coded residual of a neighboring position, either horizontally or vertically (depending on the BDPCM prediction direction)) is coded. For a block of size M (height) × N (width), let r i,j (0≤i≤M-1, 0≤j≤N-1) is the prediction residual. Let Q(r i,j )(0≤i≤M-1,0≤j≤N-1) represents the residual r i,j BDPCM is applied to the quantized residual value, generating a quantized version of The modified M×N matrix in is predicted from its neighboring quantized residual values. For vertical BDPCM prediction mode, when 0≤j≤(N-1), the following is used to derive For the horizontal BDPCM prediction mode, when 0≤i≤(M-1), the following is used to derive At the decoder side, the above process is reversed to calculate Q(r i,k ), 0≤i≤M-1, 0≤j≤N-1, as follows: If vertical BDPCM is used (2-31) If horizontal BDPCM is used (2-32). Dequantized residual Q -1 (Q(r i,j )) is added to the intra block prediction value to generate the reconstructed sample value. Predicted quantized residual value The residual is sent to the decoder using the same residual coding process as in transform skip mode residual coding. For lossless codecs, if slice_ts_residual_coding_disabled_flag is set to 1, the quantized residual value is sent to the decoder using regular transform residual coding. In terms of MPM modes for future intra mode codecs, for BDPCM-coded CUs, the horizontal prediction mode or the vertical prediction mode is stored if the BDPCM prediction direction is horizontal or vertical, respectively. For deblocking, if both blocks on either side of a block boundary are coded using BDPCM, then that particular block boundary is not deblocked. 2.1.3.3. Residual Codec for Transform Skip Mode VVC allows transform skip mode to be used for luma blocks of size up to MaxTsSize times MaxTsSize, where the value of MaxTsSize is signaled in the PPS and can be at most 32. When a CU is encoded and decoded in transform skip mode, its prediction residual is quantized and encoded using the transform skip residual encoding and decoding process. This process is modified from the transform coefficient encoding and decoding process. In transform skip mode, the residual of the TU is also encoded and decoded in units of non-overlapping sub-blocks of size 4x4. For better coding and decoding efficiency, some modifications are made to customize the residual coding and decoding process to the residual signal characteristics. The following summarizes the differences between transform skip residual coding and conventional transform residual coding: – Forward scan order is applied to scan sub-blocks within a converted block and positions within a sub-block; – no signaling of the final (x, y) position; – When all previous flags are equal to 0, coded_sub_block_flag is coded for each sub-block except the last sub-block; –sig_coeff_flag context modeling uses a simplified template, and the context model of sig_coeff_flag depends on the top and left neighboring values; – The context model of the abs_level_gt1 flag also depends on the left sig_coeff_flag value and the top sig_co-eff_flag value; –par_level_flag uses only one context model; – Additional flags greater than 3, 5, 7, and 9 are signaled to indicate coefficient magnitudes, one context per flag; – For the binarization of the remainder values, the Rice parameter derivation uses a fixed order (order = 1); – The context model for the sign flag is determined based on the left and above neighboring values, and the sign flag is parsed after sig_coeff_flag to keep all context-coded bins together. For each subblock, if coded_subblock_flag is equal to 1 (i.e., there is at least one non-zero quantized residual in the subblock), the encoding and decoding of the quantized residual magnitude is performed in three scanning passes (see Figure 27 ): - First scan pass: The significance flag (sig_coeff_flag), the sign flag (coeff_sign_flag), the absolute magnitude greater than 1 flag (abs_level_gtx_flag[0]), and the parity (par_level_flag) are coded. For a given scan position, if sig_coeff_flag is equal to 1, then coeff_sign_flag is coded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute magnitude is greater than 1). If abs_level_gtx_flag[0] is equal to 1, then par_level_flag is additionally coded to specify the parity of the absolute magnitude. – Greater than x scan passes: For each scan position where the absolute magnitude is greater than 1, up to four abs_level_gtx_flag[i] (i=1...4) are encoded to indicate whether the absolute magnitude at the given position is greater than 3, 5, 7 or 9, respectively. – Remainder scan pass: The absolute magnitude remainder abs_remainder is encoded and decoded in bypass mode. The absolute magnitude remainder is binarized using a fixed Rice parameter value of 1. The bins in scan pass #1 and scan pass #2 (the first scan pass and greater than x scan passes) are context-coded until the maximum number of context-coded bins in the TU has been exhausted. The maximum number of context-coded bins in the residual block is limited to 1.75*block_width*block_height, or equivalently, an average of 1.75 context-coded bins per sample position. The bins in the last scan pass (the remainder scan pass) are bypass-coded. The variable RemCcbs is first set to the maximum number of context-coded bins for the block and is decremented by 1 each time a context-coded bin is coded. When RemCcbs is greater than or equal to 4, the syntax elements in the first codec pass are coded using context-coded bins, including sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag. If RemCcbs becomes smaller than 4 when encoding and decoding the first pass, the remaining coefficients that have not been encoded in the first pass are encoded in the remainder scan pass (pass #3). After the first pass of encoding and decoding is completed, if RemCcbs is greater than or equal to 4, the syntax elements in the second encoding and decoding pass are encoded using the bins of context encoding, including abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag. If RemCcbs becomes less than 4 during the second pass of encoding and decoding, the remaining coefficients that have not been encoded in the second pass are encoded in the remainder scan pass (pass #3). Figure 27 The transform skip residual coding process is shown. The stars mark the positions where the context coded bits are exhausted, at which point all remaining bits are coded using the bypass codec. In addition, for blocks that are not coded in BDPCM mode, a magnitude mapping mechanism is applied to transform skip residual coding until the maximum number of bins for context coding has been reached. Magnitude mapping uses the top and left neighboring coefficient magnitudes to predict the current coefficient magnitude in order to reduce signaling overhead. For a given residual position, denote absCoeff as the absolute coefficient magnitude before mapping, and absCoeffMod as the coefficient magnitude after mapping. Let X0 denote the absolute coefficient magnitude of the left neighboring position, and let X1 denote the absolute coefficient magnitude of the upper neighboring position. Magnitude mapping is performed as follows: pred=max(X0,X1); if(absCoeff == pred) absCoeffMod = 1; else absCoeffMod=(absCoeff <pred)?absCoeff+1:absCoeff. The absCoeffMod value is then encoded as described above.After all context encoded bins have been exhausted, magnitude mapping is disabled for all remaining scan positions in the current block. 2.1.3.4. Palette Mode In VVC, palette mode is used for screen content encoding and decoding in all chroma formats supported in 4:4:4 profile (i.e., 4:4:4, 4:2:0, 4:2:2 and monochrome). When palette mode is enabled, if the CU size is less than or equal to 64x64, a flag is sent at the CU level and the number of samples in the CU is greater than 16 to indicate whether palette mode is used. Considering that applying palette mode on small CUs introduces insignificant codec gain and brings additional complexity on small blocks, palette mode is disabled for CUs with less than or equal to 16 samples. The codec unit (CU) for palette encoding is treated as a prediction mode different from intra prediction, inter prediction and intra block copy (IBC) mode. If palette mode is used, the sample values in the CU are represented by a set of representative color values. This set is called a palette. For positions with sample values close to the palette colors, the palette index is transmitted through the signal. It is also possible to specify samples outside the palette by signaling an escape symbol. For samples encoded and decoded using the escape symbol within the CU, their component values are (possibly) directly transmitted through the signal using quantized component values. This is in Figure 28 The quantized escaped symbols are binarized using a fifth-order exponential Golomb binarization process (EG5). For palette encoding and decoding, palette prediction values are maintained. For the non-wavefront case, the palette prediction values are initialized to 0 at the beginning of each slice. For the WPP case, the palette prediction values at the beginning of each CTU row are initialized to the prediction values derived from the first CTU in the previous CTU row, so that the initialization scheme between the palette prediction values and CABAC synchronization is unified. For each entry in the palette prediction value, the reuse flag is transmitted via a signal to indicate whether it is part of the current palette in the CU. The reuse flag is sent using run-length coding of zeros. Thereafter, the number of new palette entries and the component values for the new palette entries are transmitted via a signal. After encoding a CU for palette encoding and decoding, the palette prediction value will be updated using the current palette, and entries from the previous palette prediction value that are not reused in the current palette will be added at the end of the new palette prediction value until the maximum allowed size is reached. The escape flag is transmitted for each CU through a signal to indicate whether there is an escape symbol in the current CU. If there is an escape symbol, the palette table is expanded by one entry and the last index is assigned to the escape symbol. In a manner similar to the coefficient groups (CGs) used in transform coefficient coding, a CU coded using palette mode is divided into multiple row-based coefficient groups, each consisting of m samples (i.e., m=16), where for each CG, index runs, palette index values, and quantized colors for escape mode are sequentially coded / parsed. Figure 29 As shown in , as in HEVC, horizontal or vertical traversal scanning can be applied to scan the samples. The coding order for palette run-length coding in each segment is as follows: For each sample position, one context-coded binary bit run_copy_flag=0 is transmitted by signal to indicate whether the pixel has the same mode as the previous sample position, that is, whether the previously scanned sample and the current sample are both of run type COPY_ABOVE, or whether the previously scanned sample and the current sample are both of run type INDEX and have the same index value. Otherwise, run_copy_flag=1 is transmitted by signal. If the current sample has a different mode from the previous sample, one context-coded binary bit copy_above_palette_indices_flag is transmitted by signal to indicate the run type of the current sample, that is, INDEX or COPY_ABOVE. Here, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to parse the run type because INDEX mode is used by default. In the same way, if the previously parsed run type is COPY_ABOVE, the decoder does not need to parse the run type. After palette run encoding and decoding of samples in one codec pass, the index values (for INDEX mode) and quantized escape colors are grouped and encoded in another codec pass using CABAC bypass codec. This separation of context-encoded bits and bypass-encoded bits can improve the throughput within each row of CG. For slices with dual luma / chroma trees, the palette is applied separately to luma (Y component) and chroma (Cb component and Cr component), where the luma palette entries contain only Y values and the chroma palette entries contain both Cb and Cr values. For slices with a single tree, the palette is applied jointly on the Y, Cb, Cr components, i.e., each entry in the palette contains a Y value, a Cb value, a Cr value, unless the CU is encoded using a local dual tree, in which case the encoding and decoding of luma and chroma are handled separately. In this case, if the corresponding luma block or chroma block is encoded using palette mode, their palettes are applied in a manner similar to the dual tree case (this is relevant for non-4:4:4 codecs and will be further explained in 2.1.3.4.1). For slices utilizing dual-tree codec, the maximum palette predictor size is 63, and the maximum palette table size for the codec of the current CU is 31. For slices utilizing dual-tree codec, the maximum predictor and palette table sizes are halved, i.e., for each of the luma palette and chroma palettes, the maximum predictor size is 31 and the maximum table size is 15. For deblocking, palette-coded blocks on either side of a block boundary are not deblocked. 2.1.3.4.1. Palette Mode for Non-4:4:4 Content The palette modes in VVC are supported for all chroma formats in a similar way to the palette modes in HEVC SCC. For non-4:4:4 content, the following customizations are applied: 1. When signaling an escape value for a given sample position, if that sample position has only a luma component but no chroma components due to chroma downsampling, only the luma escape value is signaled. This is the same as in HEVC SCC. 2. For local dual-tree blocks, the palette mode is applied to the block in the same way as the palette mode applied to single-tree blocks, with two exceptions: a. The process of updating the palette prediction value is slightly modified as follows. Since the local dual-tree block only contains luminance (or chrominance) component, so the prediction value update process uses the signaled value of the luminance (or chrominance) component and sets it to the default value (1<<(component bit depth-1)) to fill the "missing" chrominance (or luminance) component. b. The maximum palette prediction value size is kept at 63 (because slices are coded using a single tree), but the maximum palette table size for luma / chroma blocks is kept at 15 (because blocks are coded using separate palettes). 3. For palette mode in monochrome format, the number of color components in a palette-encoded block is set to 1 instead of 3. 2.1.3.4.2. Encoder Algorithm for Palette Mode On the encoder side, the following steps are used to generate the palette table of the current CU. 1. First, in order to derive the initial entries in the palette table of the current CU, a simplified K-means clustering is applied. The palette table of the current CU is initialized to an empty table. For each sample position in the CU, the SAD between this sample and each palette table entry is calculated, and the minimum SAD among all palette table entries is obtained. If the minimum SAD is less than the predefined error limit errorLimit, the current sample is clustered with the palette table entry with the minimum SAD. Otherwise, a new palette table entry is created. The threshold errorLimit is QP-dependent and is retrieved from a lookup table containing 57 elements covering the entire QP range. After all samples of the current CU have been processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entries after the 31st entry are discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroid of each cluster from step 1 or using one of the palette colors in the palette prediction value. The option with the lower rate-distortion cost is selected as the final color of the palette table. If a cluster has only a single sample and the corresponding palette entry is not in the palette prediction value, the corresponding sample is converted to an escape symbol in the next step. 3. The palette table thus generated contains some new entries from the centroids of the clusters in step 1, and some entries from the palette predictions. Therefore, the table is reordered again so that all new entries (i.e., centroids) are placed at the beginning of the table, followed by entries from the palette predictions. Given the palette table of the current CU, the encoder selects a palette index for each sample position in the CU. For each sample position, the encoder checks the RD cost of all index values corresponding to the palette table entries and the RD cost of the index representing the escaped symbol, and selects the index with the minimum RD cost using the following equation: RD cost = distortion × (isChroma?0.8:1) + λ × bits of bypass coding (2-33). After determining the index map for the current CU, each entry in the palette table is checked to see if it is used by at least one sample position in the CU. Any unused palette entry is removed. After the index map of the current CU is determined, a grid RD optimization is applied to find the optimal run_copy_flag value and run type value for each sample position by comparing the RD costs of three options: the same as the previously scanned position, run type COPY_ABOVE, or run type INDEX. When calculating the SAD value, the sample value is scaled down to 8 bits unless the CU is encoded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. In addition, in the case of lossless codecs, only the bit rate is used in the above rate-distortion optimization step (because lossless codecs do not cause distortion). 2.1.3.5. Adaptive Color Transformation In the HEVC SCC extension, Adaptive Color Transform (ACT) is applied to reduce redundancy between the three color components in the 444 chroma format. ACT is also adopted in the VVC standard to enhance the codec efficiency of the 444 chroma format. As in HEVC SCC, ACT performs an in-loop color space conversion in the prediction residual domain by adaptively converting the residual from the input color space to the YCgCo space. Figure 30The decoding flow chart for applying ACT is shown. The two color spaces are adaptively selected by signaling an ACT flag at the CU level. When the flag is equal to 1, the residual of the CU is encoded and decoded in the YCgCo space; otherwise, the residual of the CU is encoded and decoded in the original color space. In addition, similar to the HEVC ACT design, for inter and IBC CUs, ACT is only enabled when there is at least one non-zero coefficient in the CU. For intra CUs, ACT is only enabled when the chroma components select the same intra prediction mode (i.e., DM mode) as the luma component. 2.1.3.5.1.ACT mode In the HEVC SCC extension, ACT supports both lossless and lossy codecs based on the lossless flag (i.e., cu_transquant_bypass_flag). However, the flag indicating whether the lossy or lossless codec is applied is not signaled in the bitstream. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown below. Since the YCgCo-R transform is not normalized, to compensate for the dynamic range variation of the residual signal before and after color conversion, QP adjustments of (-5, 1, 3) are applied to the transformed residuals of the Y, Cg, and Co components, respectively. The adjusted quantization parameters only affect the quantization and inverse quantization of the residual in the CU. For other codec processes (such as deblocking), the original QP is still applied. In addition, because the forward and inverse color transforms require access to the residuals of all three components, ACT mode is always disabled for separate-tree partition and ISP mode, where the prediction block sizes of different color components are different. When ACT is applied, transform skip (TS) and block differential pulse code modulation (BDPCM) extended to the coded chroma residual are also enabled. 2.1.3.5.2.ACT Fast Encoding Algorithm To avoid brute force RD searches in both the original and converted color spaces, the following fast encoding algorithm is applied in the VTM reference software to reduce encoder complexity when ACT is enabled. – The order of enabling / disabling RD checking for ACT depends on the original color space of the input video. For RGB video, the RD cost of ACT mode is checked first; for YCbCr video, the RD cost of non-ACT mode is checked first. The RD cost of the second color space is only checked if there is at least one non-zero coefficient in the first color space. – When a CU is obtained through different split paths, the same ACT enable / disable decision is reused. Specifically, when the CU is first encoded and decoded, the selected color space used for residual encoding of a CU will be stored. Then, when the same CU is obtained through another split path, instead of checking the RD cost of the two spaces, the stored color space decision will be directly reused. The RD cost of the parent CU is used to determine whether to check the RD cost of the second color space of the current CU. For example, if the RD cost of the first color space is less than the RD cost of the second color space of the parent CU, the second color space is not checked for the current CU. To reduce the number of codec modes tested, the selected codec mode is shared between the two color spaces. Specifically, for intra mode, pre-selected intra mode candidates based on SATD-based intra mode selection are shared between the two color spaces. For inter mode and IBC mode, block vector search or motion estimation is performed only once. Block vectors and motion vectors are shared between the two color spaces. 2.1.3.6. Intra-frame Template Matching (Intra-frame TMP) Intra Template Matching (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is obtained by combining the L-shaped causal neighbor of the current block with Figure 31 The search area is generated by matching another block in a predefined search area. The search area consists of the following components: R1: Current CTU R2: Upper left CTU R3: Upper CTU R4: left CTU. SAD is used as the cost function. In each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set proportionally to the block dimensions (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. That is: SearchRange_w=a*BlkW SearchRange_h=a*BlkH Where "a" is a constant that controls the gain / complexity tradeoff. In practice, "a" is equal to 5. The intra template matching tool is enabled for CUs with size less than or equal to 64 in width and height. The maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.1.3.6.1. Using Block Vectors Derived from Intra-TMP for IBC The block vector (BV) derived from intra template matching prediction (intra TMP) is used for intra block copy (IBC). The stored intra TMP BV and IBC BV of the neighboring blocks are used as spatial BV candidates in the IBC candidate list construction. 2.1.3.6.2. Direct Block Vector (DBV) Mode for Chroma Prediction For chroma components, when the chroma dual tree is activated in an intra slice, if one of the luma blocks (five positions) is coded in MODE_IBC, its block vector bvL is used and scaled to derive the chroma block vector bvC. The scaling factor depends on the chroma format sampling structure. Then, by using the position of the current chroma block (xCb, yCb) and its bvC, the corresponding offset position (xCb+bvC[0], yCb+bvC[1]) is determined, and block copy prediction is performed. Figure 32 Five positions in the reconstructed luminance samples are shown. Figure 33 The prediction process of the DBV model is shown. A CU level flag is signaled to indicate whether the proposed DBV mode is applied, as shown in Table 7. Table 7 - Binarization process for intra_chroma_pred_mode in the proposed method 2.1.4. Transformation and Quantization 2.1.4.1. Large Block Size Transformation with High-Frequency Zeroing In VVC, large block size transforms of up to 64x64 are enabled, which are mainly used for higher resolution video, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are cleared so that only the lower-frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M is equal to 64, only the left 32 columns of transform coefficients are retained. Similarly, when N is equal to 64, only the top 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without clearing any values. In addition, transform displacement is removed in transform skip mode. VTM also supports configurable maximum transform size in SPS, giving the encoder the flexibility to select up to 32-length or 64-length transform size according to the needs of a specific implementation. 2.1.4.2. Multiple Transformation Selection (MTS) for Kernel Transformations In addition to the DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of both inter-frame and intra-frame codec blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 8 shows the selected DST / DCT basis functions. Table 8 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain orthogonality of the transform matrix, the transform matrix is quantized more accurately than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within 16 bits, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS applies only to luma. MTS signaling is skipped when one of the following conditions applies: —The position of the last significant coefficient of the luminance TB is less than 1 (i.e. only DC) —The last significant coefficient of the luminance TB is located within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, the other two flags are additionally transmitted via a signal to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 9. By eliminating the intra mode and block shape dependency, a unified transform selection for ISP and implicit MTS is used. If the current block is in ISP mode or if the current block is an intra block and both intra and inter explicit MTS are enabled, only DST7 is used for both horizontal and vertical transform kernels. In terms of transform matrix accuracy, an 8-bit main transform kernel is used. Therefore, all transform kernels used in HEVC are kept the same, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8) all use an 8-bit main transform kernel. Table 9 - Transformation and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, high-frequency transform coefficients are cleared to zero for DST-7 and DCT-8 blocks with size (width or height, or both) equal to 32. Only coefficients in the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, implicit MTS can still be enabled. 2.1.4.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, LFNST is applied between the forward main transform and quantization (at the encoder) and between dequantization and inverse main transform (at the decoder side). In LFNST, either a 4x4 non-separable transform or an 8x8 non-separable transform is applied depending on the block size. For example, 4x4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and 8x8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). Figure 34 The low frequency non-separable transform (LFNST) process is shown. The application of the non-separable transform used in LFNST is described below using the input as an example. To apply 4x4 LFNST, the 4x4 input block X First, it is represented as a vector The inseparable transform is calculated as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. The 16x1 coefficient vector Subsequently, it is reorganized into 4x4 blocks using the scan order for the block (horizontal, vertical, or diagonal). Coefficients with smaller indices will be placed in the 4x4 coefficient block together with smaller scan indices. 2.1.4.3.1. Reduced Inseparable Transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method so that it can be implemented in a single pass without multiple iterations. However, the dimension of the non-separable transform matrix needs to be reduced to minimize the computational complexity and the memory space to store the transform coefficients. Therefore, the Reduced Inseparable Transform (or RST) method is used in the LFNST. The main idea of the reduced inseparable transform is to map an N-dimensional vector (for 8x8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an NxN matrix, the RST matrix becomes an R×N matrix as follows: The R rows of the transform are the R basis of the N-dimensional space. The inverse transform matrix of the RT is the transpose of its forward transform. For 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix (which is the size of a conventional 8x8 non-separable transform matrix) is reduced to a 16x48 direct matrix. Therefore, a 48x16 inverse RST matrix is used on the decoder side to generate the core (main) transform coefficients in the 8x8 upper left region. When a 16x48 matrix is applied instead of a 16x64 with the same transform set configuration, each of them takes 48 input data from three 4x4 blocks in the upper left 8x8 block, excluding the lower right 4x4 block. With the reduced dimensionality, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, with a reasonable performance degradation. To reduce complexity, LFNST is limited to being applicable only when all coefficients outside the first coefficient subgroup are insignificant. Therefore, when LFNST is applied, all pure main transform coefficients must be zero. This allows for LFNST index signaling to be adjusted at the last significant position, thus avoiding the extra coefficient scans in current LFNST designs, which only require checking significant coefficients at specific positions. The worst-case processing of LFNST (in terms of per-pixel multiplications) limits the non-separable transform to 8x16 and 8x48 transforms for 4x4 and 8x8 blocks, respectively. In these cases, when LFNST is applied, the last significant scan position must be less than 8, and less than 16 for other sizes. For blocks with 4xN and Nx4 (N>8) shapes, the proposed restrictions mean that LFNST is now applied only once, and only to the top-left 4x4 region. Since all pure main transform coefficients are zero when LFNST is applied, the number of operations required for the main transform is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when the LFNST transform is tested. Rate-distortion-optimized quantization only needs to be performed on the first 16 coefficients (in scan order); the remaining coefficients are forced to zero. 2.1.4.3.2. LFNST Transform Selection In LFNST, a total of 4 transform sets are used, each containing 2 inseparable transform matrices (kernels). As shown in Table 10, the mapping from intra prediction modes to transform sets is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an LFNST index that is explicitly signaled. This index is transmitted only once in the bitstream after the transform coefficients of each intra CU. Table 10 - Transformation selection table IntraPredMode Transform Set Index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0 2.1.4.3.3.LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to being applicable only when all coefficients outside the first coefficient subgroup are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-encoded, but does not depend on the intra prediction mode, and only the first bin is context-encoded. In addition, LFNST is applied to intra CUs in intra and inter slices, and to both luma and chroma. If dual-tree is enabled, the LFNST indexes for luma and chroma are signaled separately. For inter slices (dual-tree disabled), a single LFNST index is signaled and used for both luma and chroma. Considering that large CUs larger than 64x64 are implicitly partitioned (TU slicing) due to the existing maximum transform size limit (64x64), LFNST index searches can quadruple the data buffering for a given number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST is limited to 64x64. Note that LFNST is only enabled with DCT2. LFNST index signaling is placed before MTS index signaling. The use of scaling matrices for perceptual quantization is not explicit. The scaling matrices specified for the master matrix can be used for LFNST coefficients. Therefore, the use of scaling matrices for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied. 2.1.4.4. Sub-block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When an inter-predicted CU has cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded and decoded. For the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, part of the residual block is encoded and decoded using the inferred adaptive transform, while the other part of the residual block is cleared to zero. When SBT is used for inter-coded CUs, the SBT type and SBT location information are signaled in the bitstream. Figure 35As shown, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is similar to the binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is similar to the asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 (in units of luminance samples), 1:3 / 3:1 partitioning along this dimension is not allowed. The CU has a maximum of 8 SBT modes. In SBT-V and SBT-H, position-dependent transform kernel selection is applied to the luma transform block (chroma TBs always use DCT-2). Different kernel transforms are associated with the two positions, SBT-H and SBT-V. More specifically, the horizontal and vertical transforms for each SBT position are Figure 35 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DCT-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the horizontal and vertical kernel transform types for the TU slice, cbf, and residual block. SBT is not applied to CUs coded in the joint intra and inter modes. 2.1.4.5. Maximum Transform Size and Zeroing of Transform Coefficients The CTU size and the maximum transform size (i.e., all MTS transform kernels) are both extended to 256, where the largest intra-coded block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, there is no normalized zeroing operation applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are normalized to zero. 2.1.4.6. Enhanced MTS for intra-frame coding and decoding In the current VVC design, for MTS, only DST7 and DCT8 transform kernels are used, which are used for intra and inter coding. Additional main transforms including DCT5, DST4, DST1 and identity transform (IDT) are adopted. At the same time, the MTS set is configured based on TU size and intra-frame mode information. 16 different TU sizes are considered, and for each TU size, 5 different categories are considered based on intra-frame mode information. For each category, as with VVC, 4 different transform pairs are considered. Note that although a total of 80 different categories are considered, some of these different categories typically share exactly the same transform set. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For angle modes, the joint symmetry of TU shape and intra prediction is taken into account. Therefore, mode i (i>34) with TU shape AxB will be mapped to the same category corresponding to mode j=(68-i) with TU shape BxA. However, for each transform pair, the order of horizontal and vertical transform kernels is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical and horizontal transform kernels are swapped. For wide-angle mode, the closest conventional angle mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. The MTS index [0,3] is signaled using a 2-bit fixed-length codec. 2.1.4.7. Secondary Transform: Extension of LFNST with Large Kernel The LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is extended to S=35 and C=3, and for a given intra-frame model The LFNST set (lfnstTrSetIdx) of formula (predModeIntra) is derived according to the following formula: ○ For predModeIntra<2, lfnstTrSetIdx is equal to 2 ○ For predModeIntra in [0,34], lfnstTrSetIdx = predModeIntra o For predModeIntra in [35,66], lfnstTrSetIdx = 68 - predModeIntra. • Three different kernels LFNST4, LFNST8, and LFNST16 are defined to indicate LFNST kernel sets applied to 4xN / Nx4 (N≥4), 8xN / Nx8 (N≥8), and MxN (M, N≥16), respectively. The kernel dimension is given by the following formula: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96). Forward LFNST is applied to the upper left low-frequency region called the Region-Of-Interest (ROI). When LFNST is applied, the main transform coefficients present in the region other than the ROI are cleared, which does not change the VVC standard. The ROI of LFNST16 is Figure 36 It is shown in . It consists of six 4x4 sub-blocks, which are consecutive in scan order. Since the number of input samples is 96, the transform matrix for forward LFNST16 can be Rx96. In this paper, R is chosen to be 32, and accordingly 32 coefficients (two 4x4 sub-blocks) are generated from forward LFNST16, which are placed according to the coefficient scan order. The ROI of LFNST8 is Figure 37 The forward LFNST8 matrix can be Rx64, and R is chosen to be 32. The generated coefficients are positioned in the same way as LFNST16. The mapping from intra prediction modes to these sets is shown in the table below. Table 11 Mapping of intra prediction modes to LFNST set indexes Intra prediction mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST collection index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST collection index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra prediction mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST collection index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.1.4.8. Non-separable primary transform (NSPT) for intra-frame coding and decoding DCT-II+LFNST is replaced by NSPT for block sizes 4x4, 4x8, 8x4, and 8x8. NSPT follows the design of LFNST, i.e., 3 candidates and 35 sets, selected based on intra mode. The kernel sizes are as follows: NSPT4x4: 16x16; NSPT4x8 / NSPT8x4: 32x20; NSPT8x8: 64x32. Therefore, 12 and 32 coefficients are cleared to zero for NSPT4x8 / NSPT8x4 and NSPT8x8, respectively. 2.1.4.9. Symbol Prediction The basic idea of the coefficient sign prediction methods (JVET-D0031 and JVET-J0021) is to compute the reconstruction residuals for negative and positive sign combinations for the applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal symbol, the cost function is defined as Figure 38The discontinuity measure across block boundaries is shown above. All hypotheses are measured, and the hypothesis with the smallest cost is chosen as the predicted value of the coefficient sign. The cost function is defined as the sum of the absolute second-order derivatives in the residual domain with respect to the upper rows and left columns as follows: Where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. (-R -1 The term +2R0-P1) can be calculated only once per block and only the residual assumption is subtracted. 3. Question In ECM-7.0, OBMC can be applied to inter-coded blocks, regardless of whether they are inter-AMVP coded or inter-MERGE coded. For inter-AMVP coded blocks, a syntax flag can be signaled at the block level to indicate whether OBMC is applied. For inter-MERGE coded blocks, OBMC is implicitly inferred to be applied regardless of the block characteristics and the codec information of neighboring blocks. However, there may be some cases where inter-MERGE coded blocks do not prefer OBMC mode. For example, blocks containing sharp edges, or few gradients, or few colors may not prefer OBMC mode. Block-level adaptive OBMC that inherits the OBMC on / off flag from neighbors can bring higher coding gain. 4. Detailed solution The following detailed embodiments should be considered as examples to explain the general concept. These embodiments should not be interpreted in a narrow manner. In addition, these embodiments can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In the present disclosure, regarding "blocks encoded and decoded in mode N", "mode N" here can be a prediction mode (for example, MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a coding and decoding technology (for example, AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame-inter-frame, GPM intra-frame-intra-frame, GPM inter-frame-intra-frame, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intra-frame TMP, ALF, deblocking, SAO, bilateral filter, LMCS and their corresponding variants, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 4.1. In one example, whether OBMC is applied to the current block may be inherited from the motion vector candidate. 1) For example, the motion vector candidate may be a candidate in the inter-frame Merge list. a. For example, it can be in the inter-frame regular Merge list. b. For example, it can be in the Inter-TM Merge list. c. For example, it can be in the inter-frame BM Merge list. d. For example, it can be in the Inter-frame GEO Merge list. e. For example, it can be in the inter-frame CIIP Merge list. f. For example, it can be in the Inter-MMVD Merge list. g. For example, it can be in the inter-frame affine Merge list. h. For example, it can be in the inter-frame SbTMVP Merge list. 2) For example, the motion vector candidate may be a candidate in the inter-AMVP list. a. For example, it can be in the inter-frame regular AMVP list. b. For example, it can be in the inter-frame AMVP-Merge list. c. For example, it can be in the inter-frame affine AMVP list. 3) For example, inheritance can be based on the type of motion candidate of the current block. a. For example, OBMC parameters (eg, flags) may be inherited from spatial neighbors adjacent to the current block. b. For example, OBMC parameters (eg, flags) may be inherited from spatial neighbors that are not adjacent to the current block. i. Alternatively, whether and how to inherit may depend on the distance between the current block and the non-neighboring blocks. c. For example, OBMC parameters (eg, flags) may be inherited from the motion candidate from the HMVP table. i. Alternatively, whether and how to inherit may depend on the distance between the current block and the HMVP candidate. d. For example, for temporal motion vectors, OBMC parameters (eg, flags) may be set to default values. i. Alternatively, OBMC parameters can be inherited from the temporal motion vectors. e. For example, for paired motion vectors, OBMC parameters (eg, flags) may be set to default values. i. Alternatively, the OBMC parameter can be inherited from one direction of the motion vectors that are configured as pairs of motion vectors. f. For example, for a zero motion vector, the OBMC parameter (e.g., flag) can be set to a default value. g. For example, the default value can indicate that OBMC is used for the current block. h. For example, the default value can indicate that OBMC is not used for the current block. 4) For example, the inheritance can be based on whether the motion candidate is LIC decoded. a. For example, if the motion candidate is LIC decoded, the inherited OBMC parameter (e.g., flag) can be equal to a value indicating that OBMC is not applied to the current block. b. For example, if the motion candidate is not LIC decoded, the inherited OBMC parameter (e.g., flag) can be equal to a value indicating that OBMC is applied to the current block. 5) For example, the inheritance can be based on the block dimension of the current block (assuming W and H represent the width and height of the current block). a. For example, when at least one of the following conditions is satisfied, the inherited OBMC parameter (e.g., flag) can be equal to a value indicating that OBMC is not applied to the current block, i. W*H < T1 or W*H <= T1, ii. W < T2 or W <= T2, iii. H < T3 or H <= T3, iv. W / H < T4 or W / H <= T4, v. H / W < T5 or H / W <= T5, vi. W*H > T6 or W*H >= T6, vii. W > T7 or W >= T7, viii. H > T8 or H >= T8, ix. W / H > T9 or W / H >= T9, x. H / W > T10 or H / W >= T10. b. For example, the thresholds T1…T10 in item a can be constant values. i. For example, T1 = 32 or 64 or 16. ii. For example, T2 = 32 or 16 or 8. iii. For example, T3 = 32 or 16 or 8. iv. For example, T4 = 4 or 8 or 16. v. For example, T5 = 4 or 8 or 16. vi. For example, T6 = 32 or 64 or 128. vi. For example, T6 = 32 or 64 or 128. vii. For example, T7=32 or 64 or 128. viii. For example, T8=32 or 64 or 128. ix. For example, T9=8 or 16 or 32. x. For example, T10=8 or 16 or 32. 6) For example, inheritance can be based on the prediction mode of the current block. a. For example, if the current block is encoded or decoded in the following mode, the OBMC parameters (eg, flags) inherited for the current block may be inherited. i. Inter-frame Merge, ii. Inter-frame AMVP, iii. Conventional inter-frame Merge, iv.MHP, v. GEO (and / or its variants such as GPM™, GPM MMVD, GPM Inter-Intra), vi. CIIP (and / or its variants, such as CIIP™), vii. Inter-frame MMVD, viii. Affine MMVD, ix. Affine Merge, x.Inter-frame TM, xi. Inter-frame BM (such as adaptive DMVR), xii.AMVP-MERGE, xiii.sbTMVP. b. For example, if the current block is encoded and decoded in the following mode, the OBMC parameters (eg, flags) inherited for the current block may not be inherited. i. Inter-frame AMVP. ii.AMVP-MERGE. iii.IBC Merge. iv.IBC AMVP. 7) For example, the inheritance of an MHP-coded block may depend on the prediction mode of the base assumption. a. For example, if the basic assumption of the MHP-coded block is inter-frame Merge coding. i. For example, the OBMC parameters (eg, flags) of the Merge candidate of the base hypothesis can be inherited to the MHP-encoded block. ii. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is always used for MHP-coded blocks. iii. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is never used for MHP-coded blocks. b. For example, if the base assumption of the block coded by MHP is inter-frame AMVP, i. For example, an OBMC parameter (eg, a flag) may be signaled in the bitstream to indicate whether OBMC is used for MHP-coded blocks. ii. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is always used for MHP-coded blocks. iii. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is never used for MHP-coded blocks. 8) For example, the inheritance of GEO / GPM (and / or its variants such as GPM™, GPM MMVD, GPM Inter-Intra, etc.) coded blocks may depend on the OBMC parameters (eg, flags) of the motion candidate. a. For example, the OBMC parameters (eg, flags) of a Merge candidate in a regular Merge list may be copied to the corresponding GEO candidate in a GEOMerge list and may be inherited to GEO / GPM coded blocks. b. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is always used for GEO / GPM coded blocks. c. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is never used for GEO / GPM coded blocks. 9) For example, the inheritance of a block coded by Affine Merge (and / or its variants, such as Affine MMVD, Affine DMVR, etc.) may depend on the OBMC parameters (eg, flags) of the affine candidate. a. For example, the OBMC parameters (eg, flags) of the Affine Merge candidate can be inherited to the Affine Merge-encoded block. b. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is always used for Affine Merge coded blocks. c. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is never used for Affine Merge coded blocks. 10) For example, the inheritance of a block coded by sbTMVP Merge (and / or its variants, such as sbTMVP™, sbTMVP DMVR, etc.) may depend on the OBMC parameters (eg, flags) of the motion displacement candidate of the sbTMVP block. a. For example, the OBMC parameters (eg, flags) of the motion displacement candidate can be inherited to the sbTMVP coded blocks. b. Alternatively, the OBMC parameters (eg, flags) of the sub-block in the corresponding CU in the reference frame (eg, where the location of the corresponding CU is identified by the motion displacement candidate) may be inherited to the sbTMVP-coded block. i. For example, a sub-block may be located at the center of the corresponding CU. ii. For example, the sub-block may be located at the upper left of the corresponding CU. c. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is always used for sbTMVP coded blocks. d. Alternatively, an OBMC parameter (eg, a flag) may be set to a value indicating that OBMC is never used for sbTMVP coded blocks. 11) For example, inheritance may depend on more than one of the conditions listed above. a. For example, assuming that OBMC parameters (e.g., flags) are inherited from a motion vector candidate, it may first check whether the current block is inter-merge coded, non-LIC coded, and has a block dimension less than a certain number. Only if all these conditions are true, the inherited OBMC parameters may be set to a value indicating that OBMC is applied to the current block. 4.2. In one example, the OBMC parameters (eg, flags) of a block may be stored in a cache. 1) For example, it can be stored at a granularity of MxM (eg, M=4 or 8) sub-blocks. 1) For example, the stored OBMC parameters can be used for encoding and decoding of a subsequent block (eg, OBMC parameter inheritance, etc.). 4.3. Whether and / or how to apply the above-disclosed method can be transmitted by a signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 4.4. Whether and / or how the above disclosed methods can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of areas containing more than one sample or pixel. 4.5. Whether and / or how to apply the above disclosed methods may depend on codec information such as block size, color format, single / dual tree partitioning, color component, slice / picture type.

[0099] The term "video unit" or "codec unit" or "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, or a TB. In the present disclosure, with respect to a "block coded in mode N", "mode N" may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a codec technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, Affine, CIIP, GPM, Spatial GPM, SGPM, GPM Inter-Inter, GPM Intra-Intra, GPM Inter-Intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, Intra TMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants thereof, etc.).

[0100] Figure 39 Flowchart of a method 3900 for video processing according to an embodiment of the present disclosure is shown. The method 3900 is implemented during conversion between a target video block of a video and a bitstream of the video.

[0101] At block 3910, for conversion between a video unit of a video and a bitstream of the video unit, whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of the video unit is determined based on inheritance from a motion vector candidate. In other words, whether OBMC is applied to the current block is inherited from the motion vector candidate.

[0102] At block 3920, conversion is performed based on the determination. In some embodiments, conversion may include encoding the video unit from the bitstream. Alternatively or additionally, conversion may include decoding the video unit from the bitstream. In this manner, block-level adaptive OBMC that inherits OBMC parameters (e.g., OBMC on / off flags) from neighboring blocks can result in higher codec gain and improved codec efficiency.

[0103] In some embodiments, the OBMC parameters of the video unit are stored in a cache. For example, the OBMC parameters may be a flag of OBMC.

[0104] In some embodiments, the OBMC parameters are stored in a granularity of MxM sub-blocks. In this case, M is an integer. For example, M can be equal to 4 or 8.

[0105] In some embodiments, the OBMC parameters are used for the codec of the following block. (eg, OBMC parameter inheritance, etc.) For example, the OBMC parameters are inherited to the codec of the following block.

[0106] In some embodiments, the motion vector candidate is a candidate in the inter-frame Merge list. In some embodiments, the motion vector candidate is in the inter-frame affine Merge list. Alternatively, the motion vector candidate is in the inter-frame regular Merge list. In some other embodiments, the motion vector candidate is in the inter-frame template matching (TM) Merge list. Alternatively, the motion vector candidate is in the inter-frame block matching (BM) Merge list. In some embodiments, the motion vector candidate is in the inter-frame geometry (GEO) Merge list. In some other embodiments, the motion vector candidate is in the inter-frame intra-frame joint prediction (CIIP) Merge list. Alternatively, the motion vector candidate is in the inter-frame Merge Mode with Motion Vector Difference (MMVD) Merge list. In some embodiments, the motion vector candidate is in the inter-frame sub-block-based temporal motion vector prediction (SbTMVP) Merge list.

[0107] In some embodiments, inheritance is based on the type of motion vector candidate of the current block.

[0108] In some embodiments, the OBMC parameters are inherited from spatially neighboring blocks adjacent to the current block.For example, the OBMC parameters may be flags.

[0109] In some embodiments, OBMC parameters are inherited from spatially neighboring blocks that are not adjacent to the current block. Alternatively, whether and / or how OBMC parameters are inherited depends on the distance between the current block and the non-neighboring blocks.

[0110] In some embodiments, OBMC parameters are inherited from motion candidates from a history-based motion vector predictor (HMVP) table. Alternatively, whether and / or how OBMC parameters are inherited depends on the distance between the current block and a candidate in the HMVP table.

[0111] In some embodiments, the OBMC parameters are set to default values for the temporal motion vector. Alternatively, the OBMC parameters are inherited from the temporal motion vector.

[0112] In some embodiments, the OBMC parameters are set to default values for the paired motion vectors.Alternatively, the OBMC parameters are inherited from one direction of the motion vectors used to construct the paired motion vectors.

[0113] In some embodiments, for a zero motion vector, the OBMC parameter is set to a default value. In some embodiments, the default value indicates that OBMC is used for the current block. In some other embodiments, the default value indicates that OBMC is not used for the current block.

[0114] In some embodiments, inheritance is based on whether the motion vector candidate is locally illumination compensated (LIC) encoded / decoded. For example, if the motion vector candidate is LIC encoded / decoded, the inherited OBMC parameter is equal to a value indicating that OBMC is not applied to the current block. As another example, if the motion vector candidate is not LIC encoded / decoded, the inherited OBMC parameter is equal to a value indicating that OBMC is applied to the current block.

[0115] In some embodiments, inheritance is based on the block dimension of the current block. For example, if at least one of the following conditions is satisfied, the inherited OBMC parameter is equal to a value indicating that OBMC is not applied to the current block: W*H < T1 or W*H <= T1, W < T2 or W <= T2, H < T3 or H <= T3, W / H < T4 or W / H <= T4, H / W < T5 or H / W <= T5, W*H > T6 or W*H >= T6, W > T7 or W >= T7, H > T8 or H >= T8, W / H > T9 or W / H >= T9, or H / W > T10 or H / W >= T10, where H represents the height of the current block, W represents the width of the current block, and T1, T2, T3, T4, T5, T6, T7, T8, T9, and T10 are integers respectively.

[0116] In some embodiments, T1, T2, T3, T4, T5, T6, T7, T8, T9, and T10 are constant values respectively. For example, T1 = 32 or 64 or 16. In some embodiments, T2 = 32 or 16 or 8. In some other embodiments, T3 = 32 or 16 or 8. Alternatively, T4 = 4 or 8 or 16. In some embodiments, T5 = 4 or 8 or 16. In some other embodiments, T6 = 32 or 64 or 128. In some embodiments, T7 = 32 or 64 or 128. In some other embodiments, T8 = 32 or 64 or 128. In some embodiments, T9 = 8 or 16 or 32. In some embodiments, T10 = 8 or 16 or 32.

[0117] In some embodiments, inheritance is based on the prediction mode of the current block. For example, if the current block is encoded / decoded in a target mode, the OBMC parameter is inherited for the current block. As an example, the target mode includes at least one of the following: affine Merge, inter-frame Merge, inter-frame advanced motion vector prediction (AMVP), regular inter-frame Merge, multi-hypothesis prediction (MHP), GEO, variants of GEO, (and / or its variants, such as GPM TM, GPM MMVD, GPM intra-intra) CIIP, variants of CIIP, MMVD, affine MMVD, inter-frame TM, inter-frame BM, AMVP-MERGE or sbTMVP.

[0118] In some embodiments, if the current block is encoded and decoded in target mode, the OBMC parameters for the current block are not inherited. For example, the target mode includes at least one of the following: inter-frame AMVP, AMVP-MERGE, intra-frame block copy (IBC) Merge, or IBC AMVP.

[0119] In some embodiments, the inheritance of a block depends on the OBMC of an affine candidate, which is a block coded by affine Merge and / or a variant of affine Merge. For example, the variant of affine Merge may include one or more of affine MMVD or affine DMVR.

[0120] In some embodiments, the OBMC parameters of an Affine Merge candidate are inherited to blocks that have been coded using the Affine Merge codec and / or a variant of the Affine Merge codec. In some other embodiments, the OBMC parameters are set to a value indicating that OBMC is used for blocks that have been coded using the Affine Merge codec and / or a variant of the Affine Merge codec. Alternatively, the OBMC parameters are set to a value indicating that OBMC is not used for blocks that have been coded using the Affine Merge codec and / or a variant of the Affine Merge codec.

[0121] In some embodiments, the motion vector candidate is a candidate in the inter-frame AMVP list. For example, the motion vector candidate is in the inter-frame regular AMVP list. In some embodiments, the motion vector candidate is in the inter-frame AMVP-Merge list. In some other embodiments, the motion vector candidate is in the inter-frame affine AMVP list.

[0122] In some embodiments, the inheritance of an MHP-coded block depends on the prediction mode of the base hypothesis. In some embodiments, the base hypothesis of an MHP-coded block is inter-merge coded. For example, if the base hypothesis of an MHP-coded block is inter-merge coded, the OBMC parameters of the Merge candidate of the base hypothesis are inherited to the MHP-coded block. Alternatively, if the base hypothesis of an MHP-coded block is inter-merge coded, the OBMC parameters are set to a value indicating that OBMC is used for the MHP-coded block. In some other embodiments, if the base hypothesis of an MHP-coded block is inter-merge coded, the OBMC parameters are set to a value indicating that OBMC is not used for the MHP-coded block.

[0123] In some embodiments, the base assumption for the MHP coded block is inter-AMVP coded. For example, if the base assumption for the MHP coded block is inter-AMVP coded, an OBMC parameter is indicated in the bitstream to indicate whether OBMC is used for the MHP coded block. In some embodiments, if the base assumption for the MHP coded block is inter-AMVP coded, the OBMC parameter is set to a value indicating that OBMC is used for the MHP coded block. Alternatively, if the base assumption for the MHP coded block is inter-AMVP coded, the OBMC parameter is set to a value indicating that OBMC is not used for the MHP coded block.

[0124] In some embodiments, the inheritance of blocks coded with Geometric Partitioning Mode (GPM) and / or variants of GPM depends on the OBMC parameters of the motion vector candidate. For example, variants of GPM include at least one of: GPM™, GPM MMVD, or GPM Inter-Intra.

[0125] In some embodiments, the OBMC parameters of the merge candidates in the regular merge list are copied to the corresponding GEO candidates in the GEO merge list, and the OBMC parameters of the merge candidates in the regular merge list are inherited by blocks coded with GPM and / or GPM variants. In some other embodiments, the OBMC parameters are set to a value indicating that OBMC is used for blocks coded with GPM and / or GPM variants. Alternatively, the OBMC parameters are set to a value indicating that OBMC is not used for blocks coded with GPM and / or GPM variants.

[0126] In some embodiments, the inheritance of OBMC parameters of motion displacement candidates for blocks encoded using sbTMVP Merge and / or variants of sbTMVP depends on the OBMC parameters of motion displacement candidates for blocks encoded using sbTMVP Merge and / or variants of sbTMVP. For example, variants of sbTMVP include at least one of the following: sbTMVP TM or sbTMVP DMVR. In some embodiments, the OBMC parameters of motion displacement candidates are inherited to blocks encoded using sbTMVP Merge and / or variants of sbTMVP.

[0127] In some embodiments, the OBMC parameters of a sub-block in a corresponding codec unit (CU) in a reference frame are inherited to a block encoded using sbTMVP Merge and / or a variant of sbTMVP. For example, the location of the corresponding CU is identified using a motion displacement candidate. In some embodiments, the sub-block is located at the center of the corresponding CU. Alternatively, the sub-block is located at the upper left of the corresponding CU.

[0128] In some embodiments, the OBMC parameter is set to a value indicating that OBMC is used for blocks coded via sbTMVP Merge and / or a variant of sbTMVP. In some other embodiments, the OBMC parameter is set to a value indicating that OBMC is not used for blocks coded via sbTMVP Merge and / or a variant of sbTMVP.

[0129] In some embodiments, inheritance depends on multiple conditions. For example, inheritance can depend on more than one of the conditions listed above.

[0130] In some embodiments, the method 3900 may further include, if the OBMC parameter is inherited from the motion vector candidate, determining whether at least one of the following conditions is true: the current block is inter-merge coded, the current block is non-LIC coded, or the block dimension of the current block is less than a threshold. For example, if all conditions are true, the inherited OBMC parameter is set to a value indicating that OBMC is applied to the current block.

[0131] In some embodiments, the indication of whether and / or how to determine whether OBMC is applied to the current block is indicated at one of the following: sequence level, group of picture level, picture level, slice level, or slice group level. In some embodiments, the indication of whether and / or how to determine whether OBMC is applied to the current block is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header.

[0132] In some embodiments, the indication of whether and / or how to determine whether OBMC is applied to the current block is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or a region containing more than one sample or pixel. In some embodiments, the method 3900 further includes: determining whether and / or how to determine whether OBMC is applied to the current block based on codec information of the video unit, the codec information including at least one of the following: block size, color format, single and / or dual tree partitioning, color component, slice type, or picture type.

[0133] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; and generating a bitstream based on the determination.

[0134] According to yet other embodiments of the present disclosure, a method for storing a bitstream of a video is provided, the method comprising: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0135] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.

[0136] Item 1. A method of video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of the video unit based on inheritance from motion vector candidates; and performing the conversion based on the determination.

[0137] Clause 2. The method of clause 1, wherein the OBMC parameters of the video unit are stored in a cache.

[0138] Clause 3. The method of clause 2, wherein the OBMC parameter comprises an OBMC flag.

[0139] Clause 4. The method of clause 2, wherein the OBMC parameters are stored at a granularity of MxM sub-blocks, where M is an integer.

[0140] Item 5. The method of Item 4, wherein M is equal to 4 or 8.

[0141] Clause 6. The method of clause 2, wherein the OBMC parameters are used for encoding and decoding of a subsequent block.

[0142] Clause 7. The method of clause 6, wherein the OBMC parameters are inherited for encoding and decoding of a subsequent block.

[0143] Clause 8. The method of any one of clauses 1 to 7, wherein whether the OBMC is applied to the current block is inherited from the motion vector candidate.

[0144] Item 9. A method according to any one of Items 1 to 7, wherein the motion vector candidate is a candidate in an inter-frame Merge list.

[0145] Item 10. A method according to Item 9, wherein the motion vector candidate is in an inter-frame affine Merge list, or wherein the motion vector candidate is in an inter-frame regular Merge list, or wherein the motion vector candidate is in an inter-frame template matching (TM) Merge list, or wherein the motion vector candidate is in an inter-frame block matching (BM) Merge list, or wherein the motion vector candidate is in an inter-frame geometric (GEO) Merge list, or wherein the motion vector candidate is in an inter-frame intra-frame joint prediction (CIIP) Merge list, or wherein the motion vector candidate is in an inter-frame Merge mode with motion vector difference (MMVD) Merge list, or wherein the motion vector candidate is in an inter-frame sub-block-based temporal motion vector prediction (SbTMVP) Merge list.

[0146] Clause 11. The method of any one of clauses 1 to 7, wherein the inheritance is based on a type of motion vector candidate for the current block.

[0147] Item 12. The method of Item 11, wherein the OBMC parameters are inherited from spatial neighboring blocks adjacent to the current block.

[0148] Clause 13. The method of clause 11, wherein the OBMC parameters are inherited from spatially neighboring blocks that are not adjacent to the current block.

[0149] Item 14. The method of Item 11, wherein whether and / or how the OBMC parameters are inherited depends on a distance between the current block and a non-neighboring block.

[0150] Item 15. The method of Item 11, wherein the OBMC parameters are inherited from a motion candidate from a history-based motion vector predictor (HMVP) table.

[0151] Item 16. The method of Item 11, wherein whether and / or how the OBMC parameters are inherited depends on a distance between the current block and a candidate in the HMVP table.

[0152] Item 17. The method of Item 11, wherein for temporal motion vectors, the OBMC parameters are set to default values.

[0153] Clause 18. The method of clause 11, wherein the OBMC parameters are inherited from temporal motion vectors.

[0154] Item 19. The method according to item 11, wherein for a pair of motion vectors, the OBMC parameter is set to a default value.

[0155] Item 20. The method according to item 11, wherein the OBMC parameter is inherited from one direction of the motion vectors that form the pair of motion vectors.

[0156] Item 21. The method according to item 11, wherein for a zero motion vector, the OBMC parameter is set to a default value.

[0157] Item 22. The method according to item 11, wherein the default value indicates that the OBMC is used for the current block.

[0158] Item 23. The method according to item 11, wherein the default value indicates that the OBMC is not used for the current block.

[0159] Item 24. The method according to any one of items 1 to 7, wherein the inheritance is based on whether the motion vector candidate is locally illuminated compensated (LIC) encoded / decoded.

[0160] Item 25. The method according to item 24, wherein if the motion vector candidate is LIC encoded / decoded, the inherited OBMC parameter is equal to a value indicating that the OBMC is not applied to the current block.

[0161] Item 26. The method according to item 24, wherein the motion vector candidate is not LIC encoded / decoded, and the inherited OBMC parameter is equal to a value indicating that the OBMC is applied to the current block.

[0162] Item 27. The method according to any one of items 1 to 7, wherein the inheritance is based on the block dimension of the current block.

[0163] Item 28. The method according to item 27, wherein if at least one of the following conditions is satisfied, the inherited OBMC parameter is equal to a value indicating that the OBMC is not applied to the current block: W*H<T1 or W*H<=T1, W<T2 or W<=T2, H<T3 or H<=T3, W / H<T4 or W / H<=T4, H / W<T5 or H / W<=T5, W*H>T6 or W*H>=T6, W>T7 or W>=T7, H>T8 or H>=T8, W / H>T9 or W / H>=T9, or H / W>T10 or H / W>=T10, and wherein H represents the height of the current block, W represents the width of the current block, and T1, T2, T3, T4, T5, T6, T7, T8, T9, and T10 are integers respectively.

[0164] Item 29. The method of Item 28, wherein T1, T2, T3, T4, T5, T6, T7, T8, T9, and T10 are each constant values.

[0165] Item 30. A method according to item 29, wherein T1 = 32 or 64 or 16, and / or wherein T2 = 32 or 16 or 8, and / or wherein T3 = 32 or 16 or 8, and / or wherein T4 = 4 or 8 or 16, and / or wherein T5 = 4 or 8 or 16, and / or wherein T6 = 32 or 64 or 128, and / or wherein T7 = 32 or 64 or 128, and / or wherein T8 = 32 or 64 or 128, and / or wherein T9 = 8 or 16 or 32, and / or wherein T10 = 8 or 16 or 32.

[0166] Item 31. A method according to any one of Items 1 to 7, wherein the inheritance is based on a prediction mode of the current block.

[0167] Item 32. The method of Item 31, wherein the OBMC parameters are inherited for the current block if the current block is encoded in target mode.

[0168] Item 33. A method according to Item 32, wherein the target mode includes at least one of the following: affine merge, inter merge, inter advanced motion vector prediction (AMVP), regular inter merge, multi-hypothesis prediction (MHP), GEO, a variant of GEO, (and / or its variants, such as GPM TM, GPM MMVD, GPM inter-intra) CIIP, a variant of CIIP, inter MMVD, affine MMVD, inter TM, inter BM, AMVP-MERGE, or sbTMVP.

[0169] Item 34. The method of Item 31, wherein the OBMC parameters are not inherited for the current block if the current block is encoded in target mode.

[0170] Item 35. The method of Item 34, wherein the target mode comprises at least one of: Inter-AMVP, AMVP-MERGE, Intra-Block Copy (IBC) Merge, or IBC AMVP.

[0171] Item 36. A method according to any one of Items 1 to 7, wherein the inheritance of a block depends on the OBMC of an affine candidate, the block being a block coded by Affine Merge and / or a variant of Affine Merge.

[0172] Item 37. The method of Item 36, wherein the OBMC parameters of the Affine Merge candidate are inherited to the block coded by the Affine Merge and / or Affine Merge variant.

[0173] Item 38. The method of Item 36, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the block, the block being coded using Affine Merge and / or a variant of Affine Merge.

[0174] Item 39. The method of Item 36, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the block, the block being coded using Affine Merge and / or a variant of Affine Merge.

[0175] Clause 40. The method of any one of clauses 1 to 7, wherein the motion vector candidate is a candidate in an inter-AMVP list.

[0176] Item 41. The method of Item 40, wherein the motion vector candidate is in an inter-conventional AMVP list, or wherein the motion vector candidate is in an inter-AMVP-Merge list, or wherein the motion vector candidate is in an inter-affine AMVP list.

[0177] Item 42. The method of any one of Items 1 to 7, wherein the inheritance of the MHP-encoded block depends on a prediction mode of a base hypothesis.

[0178] Item 43. The method of Item 42, wherein the base hypothesis of the MHP-coded block is Inter-Merge coded.

[0179] Item 44. The method of Item 43, wherein if the base hypothesis of the MHP-coded block is inter-Merge coded, the OBMC parameters of the Merge candidate of the base hypothesis are inherited to the MHP-coded block.

[0180] Item 45. The method of Item 43, wherein if the base hypothesis of the MHP-coded block is Inter-Merge coded, the OBMC parameter is set to a value indicating that OBMC is used for the MHP-coded block.

[0181] Item 46. The method of Item 43, wherein if the base hypothesis of the MHP-coded block is Inter-Merge coded, the OBMC parameter is set to a value indicating that OBMC is not used for the MHP-coded block.

[0182] Item 47. The method of Item 42, wherein the base hypothesis of the MHP-coded block is inter-frame AMVP coded.

[0183] Item 48. The method of Item 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP coded, the OBMC parameter is indicated in the bitstream to indicate whether the OBMC is used for the MHP-coded block.

[0184] Item 49. The method of Item 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP coded, the OBMC parameter is set to a value indicating that OBMC is used for the MHP-coded block.

[0185] Item 50. The method of Item 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP coded, the OBMC parameter is set to a value indicating that OBMC is not used for the MHP-coded block.

[0186] Item 51. A method according to any one of items 1 to 7, wherein the inheritance of blocks encoded with Geometric Partitioning Mode (GPM) and / or a variant of GPM depends on OBMC parameters of the motion vector candidate.

[0187] Item 52. The method of Item 51, wherein the variant of GPM comprises at least one of: GPM TM, GPM MMVD, or GPM Inter-Intra.

[0188] Item 53. The method according to Item 51, wherein the OBMC parameters of the Merge candidates in the regular Merge list are copied to the corresponding GEO candidates in the GEO Merge list, and the OBMC parameters of the Merge candidates in the regular Merge list are inherited to the blocks encoded and / or decoded by the GPM and / or GPM variant.

[0189] Item 54. The method of Item 51, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the GPM and / or GPM variant coded blocks.

[0190] Item 55. The method of Item 51, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the GPM and / or GPM variant coded blocks.

[0191] Item 56. A method according to any one of items 1 to 7, wherein the inheritance of a block coded via sbTMVP Merge and / or a variant of sbTMVP depends on the OBMC parameters of the motion displacement candidate of the block coded via sbTMVP Merge and / or a variant of sbTMVP.

[0192] Item 57. The method according to Item 56, wherein the variant of sbTMVP comprises at least one of the following: sbTMVP™, or sbTMVP DMVR.

[0193] Item 58. The method of Item 56, wherein the OBMC parameters of the motion displacement candidate are inherited to the block coded by sbTMVP Merge and / or sbTMVP variant.

[0194] Item 59. The method of Item 56, wherein the OBMC parameters of the sub-block in the corresponding codec unit (CU) in the reference frame are inherited to the block coded by sbTMVP Merge and / or sbTMVP variant.

[0195] Item 60. The method of Item 59, wherein the sub-block is located at the center of the corresponding CU, or wherein the sub-block is located above and to the left of the corresponding CU.

[0196] Item 61. The method of Item 56, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the block coded by sbTMVP Merge and / or sbTMVP variant.

[0197] Item 62. The method of Item 56, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the blocks coded by sbTMVP Merge and / or sbTMVP variants.

[0198] Item 63. A method according to any one of Items 1 to 62, wherein the inheritance is dependent on a plurality of conditions.

[0199] Item 64. The method according to Item 63 further includes: if the OBMC parameters are inherited from the motion vector candidate, determining whether at least one of the conditions is true: the current block is inter-frame Merge encoded and decoded, the current block is non-LIC encoded and decoded, or the block dimension of the current block is less than a threshold.

[0200] Item 65. The method of Item 64, wherein if all of the conditions are true, the inherited OBMC parameter is set to a value indicating that the OBMC is applied to the current block.

[0201] Item 66. A method according to any one of Items 1 to 65, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at one of: sequence level, picture group level, picture level, slice level, or slice group level.

[0202] Item 67. A method according to any one of Items 1 to 65, wherein an indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.

[0203] Item 68. A method according to any one of Items 1 to 65, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or a region containing more than one sample or pixel.

[0204] Item 69. The method according to any one of Items 1 to 65 further includes: determining whether and / or how to determine whether the OBMC is applied to the current block based on codec information of the video unit, the codec information including at least one of the following: block size, color format, single and / or double tree partitioning, color component, slice type, or picture type.

[0205] Item 70. The method of any one of Items 1 to 69, wherein the converting comprises encoding the video unit into the bitstream.

[0206] Item 71. A method according to any one of Items 1 to 69, wherein the converting comprises decoding the video unit from the bitstream.

[0207] Item 72. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of items 1 to 71.

[0208] Item 73. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of Items 1 to 71.

[0209] Item 74. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; and generating the bitstream based on the determination.

[0210] Item 75. A method for storing a bitstream of a video, comprising: determining, based on inheritance from motion vector candidates, whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium. Example device

[0211] Figure 40 A block diagram of a computing device 4000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0212] It should be understood that Figure 40 The computing device 4000 shown in FIG. 4 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.

[0213] like Figure 40 As shown, computing device 4000 comprises a general computing device 4000. Computing device 4000 may include at least one or more processors or processing units 4010, memory 4020, storage unit 4030, one or more communication units 4040, one or more input devices 4050, and one or more output devices 4060.

[0214] In some embodiments, the computing device 4000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4000 can support any type of interface to the user (such as a "wearable" circuit device, etc.).

[0215] The processing unit 4010 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 4020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 4000. The processing unit 4010 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0216] The computing device 4000 typically includes various computer storage media. Such media can be any media accessible by the computing device 4000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4020 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 4030 can be any removable or non-removable medium and can include machine-readable media, such as memory, flash drive, disk or other media that can be used to store information and / or data and can be accessed in the computing device 4000.

[0217] The computing device 4000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 40 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.

[0218] The communication unit 4040 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 4000 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 4000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0219] Input device 4050 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 4060 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 4040, computing device 4000 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 4000, or, if desired, any device that enables computing device 4000 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).

[0220] In some embodiments, some or all components of the computing device 4000 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers in a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.

[0221] In an embodiment of the present disclosure, the computing device 4000 may be used to implement video encoding / decoding. The memory 4020 may include one or more video encoding / decoding modules 4025 having one or more program instructions. These modules are accessible and executable by the processing unit 4010 to perform the functions of the various embodiments described herein.

[0222] In an example embodiment performing video encoding, an input device 4050 may receive video data as input to be encoded 4070. The video data may be processed, for example, by a video codec module 4025 to generate an encoded bitstream. The encoded bitstream may be provided as output 4080 via an output device 4060.

[0223] In an example embodiment performing video decoding, an input device 4050 may receive an encoded bitstream as input 4070. The encoded bitstream may be processed, for example, by a video codec module 4025 to generate decoded video data. The decoded video data may be provided as output 4080 via an output device 4060.

[0224] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A video processing method, comprising: determining, for conversion between a video unit of a video and a bitstream of the video unit, whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of the video unit based on inheritance from motion vector candidates; as well as The converting is performed based on the determination. 2 . The method of claim 1 , wherein the OBMC parameters of the video unit are stored in a buffer. The method according to claim 2 , wherein the OBMC parameter comprises an OBMC flag. The method of claim 2 , wherein the OBMC parameters are stored at a granularity of M×M sub-blocks, where M is an integer. The method according to claim 4 , wherein M is equal to 4 or 8. The method according to claim 2 , wherein the OBMC parameters are used for encoding and decoding of a subsequent block. The method according to claim 6 , wherein the OBMC parameters are inherited for encoding and decoding of a subsequent block. 8 . The method according to claim 1 , wherein whether the OBMC is applied to the current block is inherited from the motion vector candidate.

9. The method according to any one of claims 1 to 7, wherein the motion vector candidate is a candidate in an inter-frame Merge list.

10. The method according to claim 9, wherein the motion vector candidate is in an inter-frame affine Merge list, or Where the motion vector candidate is in the inter-frame regular Merge list, or Wherein the motion vector candidate is in the Inter-Template Matching (TM) Merge list, or Wherein the motion vector candidate is in the inter-frame block matching (BM) Merge list, or Wherein the motion vector candidate is in the inter-frame geometry (GEO) Merge list, or Wherein the motion vector candidate is in the inter-frame intra-frame joint prediction (CIIP) Merge list, or wherein the motion vector candidate is in a Merge Mode (MMVD) Merge List with a motion vector difference between frames, or The motion vector candidate is in an inter-frame sub-block-based temporal motion vector prediction (SbTMVP) Merge list.

11. The method according to any one of claims 1 to 7, wherein the inheritance is based on a type of a motion vector candidate of the current block. 12 . The method of claim 11 , wherein the OBMC parameters are inherited from spatial neighboring blocks adjacent to the current block. 13 . The method of claim 11 , wherein the OBMC parameters are inherited from spatial neighboring blocks that are not adjacent to the current block. The method according to claim 11 , wherein whether and / or how the OBMC parameters are inherited depends on the distance between the current block and non-neighboring blocks.

15. The method of claim 11, wherein the OBMC parameters are inherited from a motion candidate from a history-based motion vector predictor (HMVP) table.

16. The method according to claim 11, wherein whether and / or how to inherit the OBMC parameter depends on the distance between the current block and the candidates in the HMVP table.

17. The method according to claim 11, wherein for the time-domain motion vector, the OBMC parameter is set to a default value.

18. The method according to claim 11, wherein the OBMC parameter is inherited from the time-domain motion vector.

19. The method according to claim 11, wherein for the paired motion vector, the OBMC parameter is set to a default value.

20. The method according to claim 11, wherein the OBMC parameter is inherited from one direction of the motion vectors that constitute the paired motion vector.

21. The method according to claim 11, wherein for the zero motion vector, the OBMC parameter is set to a default value.

22. The method according to claim 11, wherein the default value indicates that the OBMC is used for the current block.

23. The method according to claim 11, wherein the default value indicates that the OBMC is not used for the current block.

24. The method according to any one of claims 1 to 7, wherein the inheritance is based on whether the motion vector candidate is locally illumination compensated (LIC) encoded / decoded.

25. The method according to claim 24, wherein if the motion vector candidate is LIC encoded / decoded, the inherited OBMC parameter is equal to the value indicating that the OBMC is not applied to the current block.

26. The method according to claim 24, wherein the motion vector candidate is not LIC encoded / decoded, and the inherited OBMC parameter is equal to the value indicating that the OBMC is applied to the current block.

27. The method according to any one of claims 1 to 7, wherein the inheritance is based on the block dimension of the current block.

28. The method according to claim 27, wherein if at least one of the following conditions is satisfied, the inherited OBMC parameter is equal to the value indicating that the OBMC is not applied to the current block: W*H < T1 or W*H <= T1, W < T2 or W <= T2, H < T3 or H <= T3, W / H < T4 or W / H <= T4, H / W < T5 or H / W <= T5, W*H > T6 or W*H >= T6, W > T7 or W >= T7, H > T8 or H >= T8, W / H > T9 or W / H >= T9, or H / W > T10 or H / W >= T10, and where H represents the height of the current block, W represents the width of the current block, and T1, T2, T3, T4, T5, T6, T7, T8, T9, and T10 are integers respectively.

29. The method according to claim 28, wherein T1, T2, T3, T4, T5, T6, T7, T8, T9, and Tl0 are constant values respectively.

30. The method according to claim 29, wherein T1 = 32 or 64 or 16, and / or where T2 = 32 or 16 or 8, and / or where T3 = 32 or 16 or 8, and / or Where T4=4 or 8 or 16, and / or Where T5 = 4 or 8 or 16, and / or Where T6 = 32 or 64 or 128, and / or Where T7 = 32 or 64 or 128, and / or Where T8 = 32 or 64 or 128, and / or Where T9=8 or 16 or 32, and / or Where T10=8 or 16 or 32.

31. The method of any one of claims 1 to 7, wherein the inheritance is based on a prediction mode of the current block.

32. The method of claim 31, wherein if the current block is encoded in target mode, the OBMC parameters are inherited for the current block.

33. The method of claim 32, wherein the target mode comprises at least one of: Affine Merge, Inter-frame Merge, Inter-frame Advanced Motion Vector Prediction (AMVP), Conventional inter-frame Merge, Multiple Hypothesis Prediction (MHP), GEO, Variants of GEO, (and / or its variants such as GPM™, GPM MMVD, GPM Inter-Intra) CIIP, Variants of CIIP, Inter-frame MMVD, Affine MMVD, Inter-frame TM, Inter-frame BM, AMVP-MERGE, or sbTMVP.

34. The method of claim 31, wherein if the current block is encoded in target mode, the OBMC parameters are not inherited for the current block.

35. The method of claim 34, wherein the target mode comprises at least one of: Inter-frame AMVP, AMVP-MERGE, Intra-block copy (IBC) Merge, or IBC AMVP.

36. The method according to any one of claims 1 to 7, wherein the inheritance of a block depends on the OBMC of an affine candidate, the block being a block coded by Affine Merge and / or a variant of Affine Merge.

37. The method according to claim 36, wherein the OBMC parameters of the affine merge candidate are inherited to the block coded by the affine merge and / or a variant of the affine merge.

38. The method of claim 36, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the block, the block being coded using Affine Merge and / or a variant of Affine Merge.

39. The method of claim 36, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the block, the block being coded by Affine Merge and / or a variant of Affine Merge.

40. The method of any one of claims 1 to 7, wherein the motion vector candidate is a candidate in an inter-AMVP list.

41. The method of claim 40, wherein the motion vector candidate is in an inter-frame regular AMVP list, or Wherein the motion vector candidate is in the inter-frame AMVP-Merge list, or The motion vector candidate is in the inter-frame affine AMVP list.

42. The method according to any one of claims 1 to 7, wherein the inheritance of an MHP-coded block depends on a prediction mode of a base hypothesis.

43. The method of claim 42, wherein the base hypothesis of the MHP-coded block is Inter-Merge coded.

44. The method of claim 43, wherein if the base hypothesis of the MHP-coded block is inter-Merge coded, the OBMC parameters of the Merge candidate of the base hypothesis are inherited to the MHP-coded block.

45. The method of claim 43, wherein if the base hypothesis of the MHP-coded block is Inter Merge coded, the OBMC parameter is set to a value indicating that the OBMC is used for the MHP-coded block.

46. The method of claim 43, wherein if the base hypothesis of the MHP-coded block is Inter Merge coded, the OBMC parameter is set to a value indicating that the OBMC is not used for the MHP-coded block.

47. The method of claim 42, wherein the base hypothesis of the MHP-coded block is inter-frame AMVP codec.

48. The method of claim 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP coded, the OBMC parameter is indicated in the bitstream to indicate whether the OBMC is used for the MHP-coded block.

49. The method of claim 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP-coded, the OBMC parameter is set to a value indicating that OBMC is used for the MHP-coded block.

50. The method of claim 47, wherein if the base hypothesis of the MHP-coded block is inter-AMVP coded, the OBMC parameter is set to a value indicating that OBMC is not used for the MHP-coded block.

51. The method according to any one of claims 1 to 7, wherein the inheritance of blocks coded with Geometric Partitioning Mode (GPM) and / or variants of GPM depends on OBMC parameters of the motion vector candidate.

52. The method of claim 51, wherein the variant of the GPM comprises at least one of: GPM™, GPM MMVD, or GPM Inter-Intra.

53. The method according to claim 51, wherein the OBMC parameters of the Merge candidates in the regular Merge list are copied to the corresponding GEO candidates in the GEO Merge list, and the OBMC parameters of the Merge candidates in the regular Merge list are inherited to the blocks encoded by GPM and / or GPM variant.

54. The method of claim 51, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the GPM and / or GPM variant coded blocks.

55. The method of claim 51, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the GPM and / or GPM variant coded blocks.

56. The method according to any one of claims 1 to 7, wherein the inheritance of a block coded via sbTMVP Merge and / or a variant of sbTMVP depends on the OBMC parameters of the motion displacement candidates of the block coded via sbTMVP Merge and / or a variant of sbTMVP.

57. The method of claim 56, wherein the variant of sbTMVP comprises at least one of: sbTMVP™, or sbTMVP DMVR.

58. The method according to claim 56, wherein the OBMC parameters of the motion displacement candidate are inherited to the block coded by sbTMVP Merge and / or a variant of sbTMVP.

59. The method of claim 56, wherein the OBMC parameters of a sub-block in a corresponding codec unit (CU) in a reference frame are inherited to the block coded by sbTMVP Merge and / or a variant of sbTMVP.

60. The method of claim 59, wherein the sub-block is located at the center of the corresponding CU, or The sub-block is located at the upper left of the corresponding CU.

61. The method of claim 56, wherein the OBMC parameter is set to a value indicating that the OBMC is used for the block coded by sbTMVP Merge and / or sbTMVP variant.

62. The method of claim 56, wherein the OBMC parameter is set to a value indicating that the OBMC is not used for the blocks coded by sbTMVP Merge and / or a variant of sbTMVP.

63. A method according to any one of claims 1 to 62, wherein the inheritance depends on a plurality of conditions.

64. The method of claim 63, further comprising: If the OBMC parameters are inherited from the motion vector candidate, determining whether at least one of the conditions is true: The current block is inter-frame Merge coded and decoded. The current block is non-LIC coded, or The block dimension of the current block is smaller than a threshold.

65. The method of claim 64, wherein if all of the conditions are true, the inherited OBMC parameter is set to a value indicating that the OBMC is applied to the current block.

66. The method of any one of claims 1 to 65, wherein an indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.

67. The method according to any one of claims 1 to 65, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated in one of the following: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.

68. A method according to any one of claims 1 to 65, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is included in one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Virtual Pipeline Data Unit (VPDU), Codec Tree Unit (CTU), CTU line, strips, piece, sub-image, or An area containing more than one sample or pixel.

69. The method according to any one of claims 1 to 65, further comprising: Determine whether and / or how to determine whether the OBMC is applied to the current block based on codec information of the video unit, the codec information including at least one of the following: Block size, Color format, Single and / or double tree partitioning, Color component, Strip type, or Image type.

70. The method of any one of claims 1 to 69, wherein the converting comprises encoding the video unit into the bitstream.

71. The method of any one of claims 1 to 69, wherein the converting comprises decoding the video unit from the bitstream.

72. An apparatus for video processing, comprising a processor and non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1 to 71.

73. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of claims 1 to 71.

74. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; as well as The bitstream is generated based on the determination.

75. A method for storing a bitstream of a video, comprising: determining whether overlapped sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on inheritance from motion vector candidates; generating the bitstream based on the determination; as well as The bitstream is stored in a non-transitory computer-readable medium.