Method and device for video processing and medium

By inheriting the prediction patterns of motion vector candidates between video units and adopting the OBMC flag inheritance method, the problem of improving encoding and decoding efficiency in existing video encoding and decoding technologies is solved, and higher encoding and decoding gain is achieved.

CN120883618APending Publication Date: 2025-10-31DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480018909.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-14
Filing Date
2024-03-13
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in efficiency, especially in video compression technology, where the efficiency of video encoding/decoding needs to be improved.

Method used

An Overlapping Sub-Block Motion Compensation (OBMC) flag inheritance method is adopted, which adaptively applies OBMC parameters to improve encoding and decoding efficiency by inheriting the prediction modes or prediction patterns of motion vector candidates during the transition between video units.

Benefits of technology

By applying adaptive OBMC parameters, the encoding and decoding gain of video codecs is improved, thus enhancing encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120883618A_ABST
    Figure CN120883618A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method includes determining, for a conversion between a video unit of a video and a bitstream of the video unit, whether overlapping sub-block based motion compensation (OBMC) is applied to a current block of the video unit based on inheritance from a motion vector candidate; and performing the conversion based on the determination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to overlapping sub-block motion compensation (OBMC) flag inheritance. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. This method includes: conversion between video units and bitstreams of video units; determining whether Overlapping Subblock Motion Compensation (OBMC) is applied to the video unit based on inheritance from motion vector candidates, wherein the inheritance is based on one of: the prediction mode of the video unit or the prediction mode of the motion vector candidates; and performing the conversion based on the determination. In this manner, block-level adaptive OBMC that inherits OBMC parameters (e.g., OBMC on / off flags) from neighboring blocks can lead to higher encoding / decoding gain and improved encoding / decoding efficiency.

[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to video units of the video based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; and generating a bitstream based on the determination.

[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining whether overlapping sub-block-based motion compensation (OBMC) is applied to video units of the video based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0009] The present invention is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. The present invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0010] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown;

[0012] Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown;

[0013] Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown;

[0014] Figure 4 Intra-frame prediction mode is shown;

[0015] Figure 5 Reference samples for wide-angle intra-frame prediction are shown;

[0016] Figure 6 The discontinuity problem is shown when the orientation exceeds 45°;

[0017] Figure 7A A schematic diagram is shown showing the definition of the sample points used by the PDPC in the diagonal top-right mode applied to the diagonal and adjacent angle intra-frame modes;

[0018] Figure 7B A schematic diagram is shown showing the definition of the sample points used by the PDPC in the diagonal lower left mode applied to the diagonal and adjacent angle intra-frame modes;

[0019] Figure 7CA schematic diagram is shown illustrating the definition of the sample points used by the PDPC in the adjacent diagonal top-right mode applied to the diagonal and adjacent angle intra-frame modes;

[0020] Figure 7D A schematic diagram is shown illustrating the definition of the sample points used by the PDPC in the adjacent diagonal lower left mode applied to the diagonal and adjacent angle intra-frame modes;

[0021] Figure 8 An example of four reference rows adjacent to the prediction block is shown;

[0022] Figure 9A A schematic diagram of the sub-segmentation process depending on the block size is shown;

[0023] Figure 9B A schematic diagram of the sub-segmentation process depending on the block size is shown;

[0024] Figure 10 The matrix-weighted intra-frame prediction process is illustrated.

[0025] Figure 11 The spatial GPM candidates are shown;

[0026] Figure 12 The GPM template is shown;

[0027] Figure 13 GPM mixing is shown;

[0028] Figure 14 The locations of the spatial merge candidates are shown;

[0029] Figure 15 The candidate pairs considered for redundancy checks of spatial merge candidates are shown.

[0030] Figure 16 A schematic diagram of motion vector scaling for temporal Merge candidates is shown;

[0031] Figure 17 The candidate positions for time-domain Merge candidates C0 and C1 are shown;

[0032] Figure 18 The MMVD search point is shown;

[0033] Figure 19 The extended CU region used in BDOF is shown;

[0034] Figure 20 A schematic diagram for the symmetric MVD mode is shown;

[0035] Figure 21 This shows the refinement of motion vectors on the decoding side;

[0036] Figure 22 The top and left neighbor blocks used in the CIIP weight derivation are shown;

[0037] Figure 23 An example of GPM partitioning grouped at the same angle is shown;

[0038] Figure 24 The unidirectional prediction MV selection for geometric segmentation patterns is shown;

[0039] Figure 25 An exemplary generation of the bending weight w0 using a geometric segmentation pattern is shown;

[0040] Figure 26 The current CTU processing order and its available reference points in the current CTU and the left CTU are shown;

[0041] Figure 27 The residual encoding / decoding passes for the transform skip block are shown;

[0042] Figure 28 An example of a block encoded and decoded in palette mode is shown;

[0043] Figure 29 This shows a sub-block-based index map scan of the palette, with the left side used for horizontal scanning and the right side used for vertical scanning;

[0044] Figure 30 A flowchart of the decoding process using ACT is shown;

[0045] Figure 31 The intra-frame template matching search area used is shown;

[0046] Figure 32 Five locations in the reconstructed brightness sample points are shown;

[0047] Figure 33 The prediction process for the DBV pattern is shown;

[0048] Figure 34 The low-frequency inseparable transform (LFNST) process is illustrated.

[0049] Figure 35 The location, type, and transformation type of the SBT are shown;

[0050] Figure 36 The ROI for LFNST16 is shown;

[0051] Figure 37 The ROI for LFNST8 is shown;

[0052] Figure 38 Discontinuous measurements are shown;

[0053] Figure 39 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and

[0054] Figure 40 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0055] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation

[0056] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0057] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0058] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0059] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0060] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example Environment

[0061] Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0062] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0063] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0064] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0065] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0066] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0067] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0068] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0069] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0070] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2The examples are shown separately.

[0071] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0072] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0073] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0074] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0075] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0076] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0077] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0078] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0079] In another example, motion estimation unit 204 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0080] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0081] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0082] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0083] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0084] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0085] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0086] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block, which is stored in the buffer 213.

[0087] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0088] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0089] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0090] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0091] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0092] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge pattern. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge pattern" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0093] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0094] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0095] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0096] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0097] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding predicted block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0098] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Brief Overview This disclosure relates to video codec technology. Specifically, it concerns overlapping sub-block-based motion compensation (OBMC) and related techniques in image / video codecs. It can be applied to existing video codec standards such as HEVC, VVC, etc. It can also be applied to future video codec standards or video codecs. 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was established in 2015 by VCEG and MPEG. JVET meetings are held quarterly. The new video codec standard was officially named Multifunctional Video Codec (VVC) at the April 2018 JVET meeting, and the first version of the VVC Test Model (VTM) was also released at that time. The VVC working draft and the VTM test model are updated after each meeting. The VVC project achieved technical completion (FDIS) at a meeting in July 2020. 2.1. Existing encoding / decoding tools 2.1.1. Intra-frame prediction 2.1.1.1. Intra-mode encoding and decoding with 67 intra-prediction modes In order to capture arbitrary edge directions presented in natural video, the number of directional intra-frame modes in VVC has been expanded from 33 used in HEVC to 65. Figure 4 The red dashed arrows in the image depict new oriented modes not present in HEVC, while the planar and DC modes remain unchanged. These denser oriented intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction. In VVC, for non-square blocks, several traditional angle intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. In HEVC, each intra-codec block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is needed to generate intra-prediction values ​​using DC mode. In VVC, blocks can have rectangular shapes, which generally requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks. 2.1.1.2. Intra-frame mode encoding and decoding To maintain low complexity in generating the Most Probable Mode (MPM) list, an intra-mode encoding / decoding approach with 6 MPMs is used by considering two available neighboring intra-modes. The following three aspects are considered when constructing the MPM list: -Default intra-frame mode; -Nearest intra-frame mode; –Derived intra-frame mode. A unified 6-MPM list is used for intra-blocks, regardless of whether MRL and ISP codec tools are applied. The MPM list is constructed based on the intra-mode of the left and upper neighboring blocks. Assuming the left-hand mode is represented as Left and the upper block's mode as Above, the unified MPM list is constructed as follows: – When neighboring blocks are unavailable, the intra-frame mode is set to planar mode by default. –If both Left and Above modes are non-angular modes: –MPM list → {Plane mode, DC, V, H, V-4, V+4}. –If one of the modes Left and Above is an angular mode, and the other is a non-angular mode: – Set the mode Max to the larger of Left and Above modes. –MPM list→{Plane mode,Max,DC,Max-1,Max+1,Max-2}. –If Left and Above are both angle modes and they are different: – Set the mode Max to the larger of Left and Above modes. -If the difference between the patterns Left and Above is in the range of 2 to 62 (inclusive). –MPM list→{Plane mode,Left,Above,DC,Max-1,Max+1} -otherwise –MPM list→{Plane mode,Left,Above,DC,Max-2,Max+2}. –If Left and Above are both angle modes and they are the same: –MPM list → {Plane mode,Left,Left-1,Left+1,DC,Left-2}. Furthermore, the first binary bit of the MPM index codeword is encoded and decoded using a CABAC context. A total of three contexts are used, corresponding to whether the current intra-block is MRL-enabled, ISP-enabled, or a normal intra-block. During the 6MPM list generation process, deduplication is used to remove duplicate patterns, ensuring that only unique patterns can be included in the MPM list. For entropy encoding and decoding of the 61 non-MPM patterns, Truncation Binary Codes (TBC) are used. 2.1.1.3. Wide-angle intra-frame prediction for non-square blocks Traditional angular intra-prediction directions are defined clockwise from 45 degrees to -135 degrees. In VVC, for non-square blocks, multiple traditional angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are transmitted via signaling using the original mode indices, which are then remapped to the wide-angle mode indices after resolution. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding / decoding method also remains unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as follows: Figure 5 As shown. The number of modes replaced in the wide-angle directional mode depends on the block aspect ratio. The replaced intra-prediction modes are shown in Table 1. Table 1 – Intra-prediction modes replaced by wide-angle mode like Figure 6 As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap Δp. α The negative impact of wide-angle mode. If the wide-angle mode represents a non-fractional offset. There are 8 wide-angle modes that satisfy this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted through these modes, the samples in the reference cache are directly copied without applying any interpolation. This modification reduces the number of samples that need to be smoothed. Furthermore, it aligns the design of non-fractional modes in traditional prediction modes with that of wide-angle modes. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, with the number of entries expanded from 35 to 67 to align with the expansion of intra-prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra-prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​in the mapping table entries to more accurately convert the prediction angle for chroma blocks. 2.1.1.4. Mode-dependent intra-frame smoothing (MDIS) A four-tap intra-interpolation filter is used to improve the accuracy of oriented intra-prediction. In HEVC, a two-tap linear interpolation filter has been used to generate intra-prediction blocks in oriented prediction modes (i.e., excluding plane and DC predictions). In VVC, a simplified 6-bit 4-tap Gaussian interpolation filter is used only for oriented intra-prediction modes. The non-oriented intra-prediction process remains unchanged. The 4-tap filter is selected based on the MDIS condition for oriented intra-prediction modes that provide non-fractional shifts, i.e., all oriented modes except the following: 2, HOR_IDX, DIA_IDX, VER_IDX, 66. Based on the intra-frame prediction mode, perform the following reference sample processing: – Directional intra-prediction modes are classified into one of the following groups: – Vertical or horizontal mode (HOR_IDX, VER_IDX) – Represents diagonal patterns for angles that are multiples of 45 degrees (2, DIA_IDX, VDIA_IDX), - Remaining directional patterns; – If the directional intra-prediction mode is classified as belonging to group A, no filter is applied to the reference sample to generate the prediction sample. – Otherwise, if the pattern falls into group B, the [1,2,1] reference sample filter can be applied to the reference samples (according to the MDIS condition) to further copy these filtered values ​​into the intra-frame prediction values ​​according to the selected direction, but no interpolation filter is applied. Otherwise, if the mode is classified as belonging to group C, only the intra-frame reference sample interpolation filter is applied to the reference samples to generate predicted samples that fall between the reference samples at fractional or integer positions, depending on the selected direction (no reference sample filtering is performed). 2.1.1.5. Location-dependent intra-frame prediction combination In VVC, the intra-prediction results for DC, planar, and multiple angle modes are further modified using the Position-Related Intra-Prediction Combination (PDPC) method. PDPC is an intra-prediction method that calls a combination of unfiltered boundary reference samples and HEVC-style intra-prediction with filtered boundary reference samples. PDPC is applied to the following intra-prediction modes without signal transmission: planar, DC, horizontal, vertical, lower-left angle mode and its eight adjacent angle modes, and upper-right angle mode and its eight adjacent angle modes. The predicted sample point pred(x',y') is predicted using the intra-frame prediction mode (DC, plane, angle) and a linear combination of reference samples according to the following equation 3-8: pred(x',y')=(wL×R -1,y’+wT×R x’,-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x',y')+32)>>6 (2-1) Where R x,-1 R -1,y Let R represent the reference points located at the top and left boundaries of the current sample point (x, y), respectively, and R... -1,-1 This represents the reference sample point located at the top left corner of the current block. If PDPC is applied to DC, planar, horizontal, and vertical intra-frame modes, no additional boundary filters are required, as in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is the same for DC and planar modes, avoiding clipping operations. For angled modes, the PDPC scaling factor is adjusted so that range checking is not required, and the angle condition for enabling PDPC is removed (using scale>=0). Additionally, in all angled mode cases, the PDPC weights are based on 32. The PDPC weights depend on the prediction mode, as shown in Table 2. PDPC is applied to blocks with both width and height greater than or equal to 4. Figures 7A-7D Reference samples (R) of PDPC applied to various prediction modes are shown. x,-1 R -1,y and R -1,-1 The definition of ). The predicted sample point pred(x',y') is located at (x',y') within the prediction block. For example, for diagonal mode, the reference sample point R x,-1 The coordinates x are given by the following formula: x = x' + y' + 1, and the reference point R is... -1,y The coordinates y are similarly given by the following formula: y = x' + y' + 1. For other ring patterns, refer to sample point R. x,-1 and R -1,y It can be located at a fractional sample point position. In this case, the sample value at the nearest integer sample point position is used. Table 2 - Examples of PDPC weights based on prediction models 2.1.1.6. Multi-reference line (MRL) intra-frame prediction Multi-reference line (MRL) intra-prediction uses more reference lines for intra-prediction. Figure 8The example depicts four reference lines, where the samples for segments A and F are not obtained from reconstructed neighboring samples, but are instead filled with the nearest samples from segments B and E, respectively. HEVC intra-frame image prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines are used (reference line 1 and reference line 3). The index (mrl_idx) of the selected reference line is signaled and used to generate the intra-prediction value. For reference line idx greater than 0, only the additional reference line mode is included in the MPM list, and only the MPM index is signaled, not the remaining modes. The reference line index is signaled before the intra-prediction modes, and in the case of signaling a non-zero reference line index, the planar mode is excluded from the intra-prediction modes. MRL is disabled for the first row block within the CTU to prevent the use of extended reference samples outside the current CTU row. Additionally, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in the intra-frame prediction mode for a non-zero reference row index is aligned with the derivation for reference row index 0. MRL requires storing three neighboring luma reference rows as well as the CTU to generate predictions. The Cross-Component Linear Model (CCLM) tool also requires three neighboring luma reference rows for its downsampling filter. The definition of MRL using the same three rows is aligned with CCLM to reduce storage requirements for the decoder. 2.1.1.7. Intra-Frame Sub-Segmentation (ISP) Intra-frame sub-segmentation (ISP) divides the luma intra-prediction block vertically or horizontally into 2 or 4 sub-segments based on the block size. For example, the minimum block size for ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into 4 sub-segments. It has been noted that M×128 (with M≤64) and 128×N (with N≤64) ISP blocks may pose potential problems with a 64×64 VDPU. For example, an M×128 CU in a single-tree case has M×128 luma TBs and two corresponding... Chroma TB. If the CU uses an ISP, the luminance TB will be divided into four M×32 TBs (horizontal division is possible only), each TB smaller than a 64×64 block. However, in the current ISP design, the chroma block is not divided. Therefore, both chroma components will have a size larger than a 32×32 block. Similarly, a 128×N CU using an ISP may cause a similar situation. Therefore, both cases are problems for a 64×64 decoder pipeline. Therefore, the CU size that can use an ISP is limited to a maximum of 64×64. Figure 9A and Figure 9B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples. In ISP, 1xN / 2xN sub-block prediction is not allowed to rely on the reconstructed values ​​of previously decoded 1xN / 2xN sub-blocks of the codec block, making the minimum prediction width of the sub-block four samples. For example, an 8xN (N>4) codec block encoded and decoded using an ISP with vertical partitioning is divided into two prediction regions of size 4xN and four transforms of size 2xN. Furthermore, a 4xN codec block encoded and decoded using an ISP with vertical partitioning is predicted using the complete 4xN block; it uses four transforms of size 1xN. Although 1xN and 2xN transform sizes are allowed, it is asserted that the transforms of these blocks within the 4xN region can be performed in parallel. For example, when a 4xN prediction region contains four 1xN transforms, there are no transforms in the horizontal direction; the transform in the vertical direction can be performed as a single 4xN transform in the vertical direction. Similarly, when a 4xN prediction region contains two 2xN transform blocks, the transform operations of the two 2xN blocks in each direction (horizontal and vertical) can be performed in parallel. Therefore, no latency is added when processing these smaller blocks compared to processing intra-blocks of regular 4x4 encoding and decoding. Table 3 – Entropy Encoding / Decoding Coefficient Group Dimensions Block size Coefficient group size 1×N, N≥16 1×16 N×1, N≥16 16×1 2×N, N≥8 2×8 N×2, N≥8 8×2 All other possible M×N cases 4×4 For each subsegment, reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated through processes such as entropy decoding, inverse quantization, and inverse transform. Therefore, the reconstructed sample values ​​of each subsegment can be used to generate the prediction for the next subsegment, and each subsegment is processed repeatedly. Furthermore, the first subsegment to be processed is the one containing the upper left sample of the CU, and then processing continues downwards (horizontal division) or to the right (vertical division). As a result, the reference samples used to generate the subsegment prediction signal are only located to the left and top of the row. All subsegments share the same intra-frame mode. The following is a summary of the interactions between the ISP and other codec tools. - Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP encoding / decoding mode will be presumed to be 0, and therefore the ISP mode information will not be sent to the decoder. – Entropy Encoding Coefficient Group Size: The size of the entropy encoding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 3. Note that the new size only affects blocks generated by the ISP where one of their dimensions is less than 4 samples. In all other cases, the coefficient group remains 4×4. –CBF encoding / decoding: Assume that at least one subsegment has a non-zero CBF. Therefore, if n is the number of subsegments and the first n-1 subsegments produce zero CBF, then the CBF of the nth subsegment is presumed to be 1. –MPM usage: The MPM flag will be presumed to be 1 in blocks encoded and decoded by ISP mode, and the MPM list will be modified to exclude DC mode and give priority to horizontal intra-frame modes for ISP horizontal division and vertical intra-frame modes for vertical division. – Transformation size limitation: All ISP transforms with a length greater than 16 points use DCT-II. –PDPC: When the CU uses ISP encoding / decoding mode, the PDPC filter will not be applied to the resulting sub-segments. –MTS Flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0, and it will not be sent to the decoder. Therefore, the encoder will not perform RD tests for each different available transform in the resulting sub-segment. Instead, the transform selection for ISP mode will be fixed and selected based on the intra-frame mode utilized, the processing order, and the block size. Therefore, no signal transmission is required. For example, making t H and t V Let w be the horizontal transformation and h be the vertical transformation selected for the w×h sub-segment, respectively, where w is the width and h is the height. The transformation is selected according to the following rules: – If w = 1 or h = 1, then there is no horizontal or vertical transformation, respectively. – If w = 2 or w > 32, then t H =DCT-II. – If h = 2 or h > 32, then t V =DCT-II. Otherwise, the transformations are selected as shown in Table 4. Table 4 – Transform selection depends on intra-frame mode In ISP mode, all 67 intra-frame modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. Furthermore, the selection criteria for intra-frame interpolation filters no longer exist, and a cubic (DCT-IF) filter is always applied to fractional position interpolation in ISP mode. 2.1.1.8. Matrix-Weighted Intra-Prediction (MIP) Matrix-weighted intra-prediction (MIP) is a new intra-prediction technique added to VVC. To predict samples of a rectangular block of width W and height H, MIP takes H reconstructed neighbor boundary samples from the left row of the block and W reconstructed neighbor boundary samples from the top row of the block as input. If reconstructed samples are unavailable, they are generated in the same manner as traditional intra-prediction. The generation of the predicted signal is based on three steps: averaging, matrix-vector multiplication, and linear interpolation, as follows: Figure 10 As shown. • Calculate the average of neighboring samples In the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging neighboring boundary samples according to predefined rules depending on the block size. top and bdry left Shrink to a smaller boundary and Then, the two narrowing boundaries and spliced ​​to the reduced boundary vector bdry red Therefore, for a block with a shape of 4×4, its size is 4, while for all other shapes, its size is 8. If `mode` refers to the MIP mode, then the splicing is defined as follows: Matrix multiplication Using averaged samples as input, matrix-vector multiplication is performed, followed by the addition of an offset. The result is a scaled-down predicted signal on a downsampled set of samples from the original block. This is derived from the scaled input vector bdry. red Generate a reduced prediction signal pred red, Its width is W red And the height is H red The signal on the downsampled block. Here, W red and H red Defined as: Reduced prediction signal pred red It is calculated by calculating the matrix-vector product and adding an offset: pred red =A·bdry red +b. Here, if W = H = 4, then A has W. red ·H red A is a matrix with 4 rows and 4 columns, and in all other cases, A is a matrix with 8 columns. b is a matrix of size W. red ·H red The vector. Matrix A and offset vector b are from sets S0, S1, S... 2. The index is obtained from one of the following: idx = idx(W, H) is defined as follows: Here, each coefficient of matrix A is represented with 8 bits of precision. Set S0 consists of 16 matrices. and 16 offset vectors The set consists of 8 matrices, each with 16 rows and 4 columns, and each offset vector with a size of 16. The matrices and offset vectors of this set are used in blocks of size 4×4. and 8 offset vectors The set S2 consists of 6 matrices, each with 16 rows and 8 columns, and each offset vector has a size of 16. and 6 offset vectors of size 64 The matrix is ​​composed of 64 rows and 8 columns. Interpolation The predicted signals at the remaining locations are generated from the predicted signals on the downsampled set through linear interpolation, which is a single-step linear interpolation in each direction. Regardless of the block shape or size, the interpolation is first performed in the horizontal direction and then in the vertical direction. • Signaling in MIP mode and coordination with other codec tools For each codec unit (CU) in intra-frame mode, a flag indicating whether MIP mode should be applied is transmitted. If MIP mode is to be applied, the MIP mode (predModeIntra) is transmitted via signaling. For MIP mode, the transpose flag (isTransposed) determining whether the mode is transposed, and the MIP mode Id (modeId) determining which matrix to use for a given MIP mode, are derived as follows: isTransposed=predModeIntra&1 modeId=predModeIntra>>1(2-6). The MIP codec mode coordinates with other codec tools by taking into account the following aspects: – Enable LFNST for MIPs on large blocks. Here, the LFNST transform for planar mode is used. – The reference sample derivation for MIP is performed in exactly the same way as the reference sample derivation for conventional intra-prediction modes. – For the upsampling step used in MIP prediction, the original reference sample is used instead of the downsampling reference sample. - The limiting is performed before upsampling, not after upsampling. – Regardless of the maximum transform size, MIPs are allowed to reach 64x64. – For sizeId=0, the number of MIP patterns is 32; for sizeId=1, the number of MIP patterns is 16; and for sizeId=2, the number of MIP patterns is 12. 2.1.1.9. Airspace GPM (SGPM) In the spatial GPM, a candidate list is constructed, comprising segmentation and two intra-prediction modes. Up to 11 intra-prediction modes are used in the MPM to form a combination, and the length of the candidate list is set to 16. The selected candidate indices are transmitted via signaling. use Figure 11 The template shown is used to reorder the list. The GPM blending process is not used in the template, and the SAD between the template's prediction and reconstruction is used for sorting. Figure 12 The GPM template is shown. Figure 13 The GPM split boundary is shown. The SGPM mode is applied to blocks whose width and height meet the same constraints as in inter-frame GPM. Consider the following items: ●Airspace GPM segmentation mode: 26 predefined patterns An adaptive derivation algorithm based on the ratio of horizontal gradient to vertical gradient. ● Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation pattern, an IPM list is derived for each segment using an intra-to-inter-frame GPM list. The IPM list size is 3. In the list, a TIMD-derived pattern is replaced by two derived patterns with horizontal and vertical orientations (using top or left templates), or the TIMD-derived pattern is excluded. MPM list: A uniform MPM list (up to 11 elements) is used for all splitting patterns. ● Template size (left and top): 1 or 4. ● Expanded block size: The spatial domain GPM is extended to be further applied to 4x8, 8x4, 4x16 and 16x4 blocks, which can be described as 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32. ●Adaptive Hybridization: For the adaptive mixing of spatial GPM tests, the mixing depth τ is derived as follows: ■If min(width, height) = 4, then 1 / 2τ is selected. ■ Otherwise, if min(width, height) = 8, then τ is selected. ■ Otherwise, if min(width, height) = 16, then 2τ is selected. ■ Otherwise, if min(width, height) = 32, then 4τ is selected. ■Otherwise, 8τ is selected. 2.1.2. Inter-frame prediction For each inter-frame prediction CU, motion parameters include motion vectors, reference picture indices, reference picture list usage indices, and additional information required for inter-frame prediction sample generation using new encoding / decoding features of the VVC. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded motion vector increments, or reference picture indices. A Merge mode is specified, whereby motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as additional scheduling introduced in the VVC. The Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other required information are explicitly transmitted via signaling for each CU. In addition to the inter-frame coding and decoding features in HEVC, VVC also includes many new and improved inter-frame predictive coding and decoding tools, as listed below: – Extended Merge Forecast – Merge pattern with MVD (MMVD) –Symmetrical MVD (SMVD) signal transmission -Affine Motion Compensation Prediction – Sub-block-based temporal motion vector prediction (SbTMVP) -Adaptive Motion Vector Resolution (AMVR) –Sports field storage: 1 / 16 luminance sample MV storage and 8x8 sports field compression – Bidirectional prediction (BCW) with CU-level weights – Bidirectional optical flow (BDOF) –Decoder-side motion vector refinement (DMVR) – Geometric Partitioning (GPM) – Intra-frame and inter-frame joint prediction (CIIP). The following section provides details of these inter-frame prediction methods specified in VVC. 2.1.2.1. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: 1) Airspace MVP from adjacent CUs 2) Temporal MVP from the same CU 3) Historical MVPs from FIFO tables 4) Paired average MVP 5) Zero MV. The size of the Merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU encoded in Merge mode, the index of the best Merge candidate is encoded using rounding univariate binarization (TU). The first bit of the Merge index is encoded using the context, and bypass encoding is used for the remaining bits. This section provides the derivation process for each category of merge candidates. Similar to HEVC, VVC also supports parallel derivation of the merge candidate list for all CUs within a region of a specific size. 2.1.2.1.1. Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Figure 14 From the candidates at the indicated positions, a maximum of four merge candidates can be selected. The derivation order is B. 0, A 0, B 1, A1 and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or when it is intra-frame encoded / decoded. After adding candidates at position A1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only... Figure 15 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information. 2.1.2.1.2. Derivation of Time-Domain Candidates In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU is explicitly transmitted via signal transmission in the strip header. Figure 16 As shown by the dashed lines, the scaled motion vectors of the temporal merge candidate are obtained by scaling the motion vectors of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index of the temporal merge candidate is set to 0. like Figure 17 As shown, the position of the temporal candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded or decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. 2.1.2.1.3. Derivation of Merge Candidates Based on History Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, indicating that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained First-In-First-Out (FIFO) rule is used, where a redundancy check is first applied to find if a duplicate HMVP already exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked sequentially and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied to HMVP candidates for spatial or temporal Merge candidates. To reduce the number of redundant check operations, the following simplifications are introduced: 1. Is the number of HMPV candidates used for the Merge list generation set to (N<=4)? M:(8-N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list building process starting from HMVP is terminated. 2.1.2.1.4. Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid. When the Merge list is not full after adding pairwise average Merge candidates, zero MVP is inserted at the end until the maximum number of Merge candidates is reached. 2.1.2.2. Merge Estimation Region The Merge Estimation Region (MER) allows for the independent derivation of Merge candidate lists for CUs within the same Merge Estimation Region (MER). Candidate blocks located within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. Furthermore, the update process for the historical motion vector prediction candidate list is only updated if (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel), where (xCb, yCb) is the top-left brightness sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and transmitted via signaling as log2_parallel_merge_level_minus2 in the sequence parameter set. 2.1.2.3. Merge Pattern with MVD (MMVD) In addition to the Merge mode (where implicitly derived motion information is directly used for generating prediction samples for the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is transmitted immediately after the skip flag and the Merge flag are sent via signaling to indicate whether the MMVD mode is used for the CU. In MMVD, after selecting a Merge candidate, the Merge candidate is further refined using MVD information transmitted via signals. This further information includes a Merge candidate flag, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, one of the top two candidates in the Merge list is selected as the MV basis. The Merge candidate flag is transmitted via signals to specify which one to use. The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 18 As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 5. Table 5 – Relationship between Distance Index and Predefined Offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 6. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is a unidirectional or bidirectional predicted MV, and both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbols in Table 6 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional predicted MV, and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), the symbols in Table 6 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the signs for list 1 MVs have the opposite values. Table 6 – Signs of MV Offsets Defined by Direction Index Directional IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.4. Bidirectional prediction with CU-level weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-7) Weighted average bidirectional prediction allows five weights, w∈{-2,3,4,5,10}. For each bidirectionally predicted CU, the weights w are determined in one of two ways: 1) for non-merge CUs, the weight index is transmitted via signal after the motion vector difference; 2) for merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency images, all five weights are used. For non-low-latency images, only three weights (w∈{3,4,5}) are used. – At the encoder, a fast search algorithm is applied to find the weight indices without significantly increasing encoder complexity. These algorithms are summarized below. When combined with AMVR, if the current image is a low-latency image, unequal weights are conditionally checked only for 1-pixel and 4-pixel motion vector precision. – When combined with an affine pattern, an affine ME is performed on unequal weights if and only if the affine pattern is selected as the current best pattern. – When the two reference images in bidirectional prediction are the same, only conditional checks are performed on unequal weights. –Unequal weights are not searched when certain conditions are met, depending on the POC distance between the current image and its reference image, the encoding / decoding QP, and the temporal level. The BCW weight index is encoded using a context-coded bit and a subsequent bypass-coded bit. The first context-coded bit indicates whether equal weights are used; if unequal weights are used, the bypass-coded bit is used to signal additional bits to indicate which unequal weights are used. Weighted Prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with fading. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be transmitted via signaling for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the corresponding weights and offsets of the reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not transmitted via signaling, and w is presumed to be 4 (i.e., equal weights are applied). For MergeCU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to normal Merge patterns and inherited affine Merge patterns. For constructed affine Merge patterns, affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of the CU using the constructed affine Merge pattern is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be used together for a CU. When a CU is encoded or decoded using CIIP mode, the BCW index of the current CU is set to 2, for example, with equal weights. 2.1.2.5. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version, requiring far fewer computations, especially in terms of the number of multiplications and multiplier size. BDOF is used to refine the bidirectional prediction signal of the CU at the 4x4 sub-block level. BDOF is applied to the CU if all of the following conditions are met: –CU is encoded and decoded using a “true” bidirectional prediction mode, meaning that one of the two reference images is displayed before the current image in the order of display, and the other is displayed after the current image in the order of display. – The distances (i.e., the difference in point of view) from the two reference images to the current image are the same; Both reference images are short-term reference images; –CU is not encoded or decoded using affine mode or ATMVP Merge mode; –The CU has more than 64 luminance samples; - Both the CU height and CU width are greater than or equal to 8 luminance samples; – The BCW weight index indicates equal weights; –WP is not enabled for the current CU; –CIIP mode is not used in the current CU. BDOF is applied only to the luminance component. As the name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of an object is smooth. For each 4x4 sub-block, motion refinement (v) is calculated by minimizing the difference between the L0 and L1 predicted samples. x ,v y Then, motion refinement is used to adjust the bidirectional prediction sample values ​​in the 4x4 sub-blocks. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two predicted signals. and It is calculated by directly calculating the difference between two neighboring sample points, that is, Where I (k) (i,j) is the sample value at coordinate (i,j) of the predicted signal in list k (k=0,1), and shift1 is calculated based on the luminance bit depth bitDepth as shift1=max(6,bitDepth-6). Then, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5, and S6 are calculated as follows: in θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ) Where Ω is the 6x6 window surrounding the 4x4 sub-block, and n a and n b The values ​​are set to min(1, bitDepth-11) and min(4, bitDepth-8), respectively. Then, using the following formula, the motion is refined (v x ,v y It is derived using cross-correlation and autocorrelation terms: in th′ BIO =2 max(5,BD-7) . It is a floor function, and Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4x4 sub-block: Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset >> shift (2-13). These values ​​were chosen such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. To derive the gradient value, some predicted samples I in the list k (k = 0, 1) outside the current CU boundary are used. (k) (i,j) needs to be generated. For example... Figure 19 As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values ​​are used only for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled from their nearest neighbors (i.e., repeated). When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 and L1 prediction samples is less than a threshold. The threshold is set to equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 and L1 prediction samples calculated in the DVMR process is reused here. Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning either of the two reference images has a luma_weight_lx_flag of 1. BDOF is also disabled when the CU is encoded and decoded using symmetric MVD mode or CIIP mode. 2.1.2.6. Symmetric MVD Encoding and Decoding In VVC, in addition to the normal one-way and two-way prediction MVD signal transmission, a symmetrical MVD mode for two-way prediction MVD signal transmission is also applied. In the symmetrical MVD mode, the reference image indices of both List 0 and List 1, as well as the motion information of the MVD in List 1, are not transmitted via signal transmission but are derived. The decoding process of the symmetric MVD mode is as follows: 1) At the strip level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to 0. Otherwise, if the most recent reference image in list 0 and the most recent reference image in list 1 form a forward and backward reference pair or a backward and forward reference pair, then BiDirPredFlag is set to 1, and both the reference images in list 0 and list 1 are short-term reference images. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectional predictive codec and BiDirPredFlag is equal to 1, then the symmetric mode flag indicating whether to use symmetric mode is explicitly indicated by the signal transmission. Figure 20 This is a schematic diagram of the symmetric MVD mode. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly transmitted via signals. The reference indices of list-0 and list-1 are each set to a pair of reference images. MVD1 is set to equal to (-MVD0). The final motion vector is shown in the following formula. In the encoder, symmetric MVD motion estimation begins with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from a one-way prediction search, MVs obtained from a two-way prediction search, and MVs from an AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.2.7. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the motion vector refinement (MV) in the Merge mode, a decoder-side motion vector refinement based on bilateral matching is applied in the VVC. In the bidirectional prediction operation, a refined MV is searched around the initial MV in reference image lists L0 and L1. The BM method computes the distortion between two candidate blocks in reference image lists L0 and L1. Figure 21 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs encoded and decoded using the following modes and features: - CU-level Merge pattern with bidirectional prediction of MV – Relative to the current image, one reference image is from the past and the other is from the future. – The distances (i.e., the difference in point of view) from the two reference images to the current image are the same. Both reference images are short-term reference images. –CU has more than 64 luminance samples. - Both the CU height and CU width are greater than or equal to 8 luminance samples. –BCW weight index indicates equal weights – Do not enable WP for the current block –CIIP mode is not used in the current block. The refined motion vector (MV) derived through the DMVR process was used to generate inter-frame prediction samples and also for temporal motion vector prediction in future image encoding and decoding. The original MV was used in the deblocking process and also for spatial motion vector prediction in future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-entries. 2.1.2.7.1. Search Scheme In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirror rule. In other words, any point examined by DMVR, represented by the candidate MV pair (MV0, MV1), obeys the following two equations: MV0 ′ =MV0 + MV_offset (2-15) MV1 ′ =MV1 - MV_offset (2-16) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference images. The refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainties, a bias towards the original MV is proposed during the DMVR process. The SAD value between the reference blocks of the initial MV candidate reference is reduced by 1 / 4. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the surface equation of the parameter error, rather than through an additional search with SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center. In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form. E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (2-17) Where (x) min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points, (x) min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-18) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-19). x min and y min The value is automatically constrained between -8 and 8 because all cost values ​​are positive, and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score (x min ,y min An integer distance is added to the thin MV to obtain a subpixel accurate thin increment MV. 2.1.2.7.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 of a lumen sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search point surrounds the initial fractional pixel MV with an integer sample offset; therefore, for the DMVR search process, samples at those fractional positions need to be interpolated. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, utilizing a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not needed by the interpolation process based on the original MV but are needed by the interpolation process based on the refined MV are filled from these available samples. 2.1.2.7.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luminance samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luminance samples. The maximum cell size for the DMVR search process is limited to 16x16. 2.1.2.8. Intra-Frame and Inter-Frame Joint Prediction (CIIP) In VVC, when a CU is encoded or decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether the inter-frame / intra-frame joint prediction (CIIP) mode is applied to the current CU. Figure 22 The top and left neighbor blocks used in the CIIP weight derivation are shown. As the name suggests, CIIP prediction combines the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal P in CIIP mode... inter The inter-frame prediction process is derived using the same procedure as in the regular Merge mode; and the intra-frame prediction signal P intra The conventional intra-frame prediction process with a planar pattern is derived. Then, a weighted average is used to combine the intra-frame and inter-frame prediction signals, where weights are calculated based on the encoding / decoding modes of the top and left neighboring blocks, as follows: - If the top nearest neighbor is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. - If the left nearest neighbor is available and is intra-coded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. – If (isIntraLeft+isIntraTop) equals 2, then wt is set to 3; Otherwise, if (isIntraLeft+isIntraTop) equals 1, then wt is set to 2; Otherwise, set wt to 1. The CIIP predictions are formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2(2-20). 2.1.2.9. Multiple Hypothesis Prediction (MHP) In inter-frame AMVP mode, regular Merge mode, and MMVD mode, up to two additional prediction values ​​are transmitted via signaling. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α) n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the table below. add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-frame AMVP mode, MHP is applied only when unequal weights are selected in BCW in bidirectional prediction mode. 2.1.2.10. Overlapping Sub-block Motion Compensation (OBMC) When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information from neighboring blocks with weighted prediction. The following conditions should not be used for OBMC: • When OBMC is disabled at the SPS level • When the current block has intra-frame mode or IBC mode • When the current block applies a LIC • When the area of ​​the current luminance block is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom, and right sub-block boundary pixels using motion information from neighboring sub-blocks. Enable this for sub-block-based codec tools: • Affine AMVP mode; • Affine Merge pattern and sub-block-based temporal motion vector prediction (SbTMVP); • Bilateral matching based on sub-blocks. 2.1.2.11. Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique used to model the local illumination variation between the current block and its predicted block as a function of the local illumination variation between the current block template and the reference block template. The parameters of this function can be represented by scaling α and offset β, which form a linear equation, α*p[x]+β, to compensate for the illumination variation, where p[x] is the reference sample pointed to by the MV at position x on the reference image. When surround motion compensation is enabled, the MV must account for the surround offset being clipped. Since α and β can be derived based on the current block template and the reference block template, they require no signaling overhead except for signaling the LIC flag for AMVP mode to indicate its use. The local illumination compensation proposed in JVET-O0066 was used for unidirectional prediction of inter-frame CU, and the following modifications were made. • Intra-frame neighbor samples can be used in LIC parameter derivation; • Disable LIC for blocks with fewer than 32 luminance samples; • For both non-sub-block mode and affine mode, the LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than based on the partial template block samples corresponding to the first upper left 16x16 cell; • The samples of the reference block template are generated by using a MC with block MV without rounding it to integer pixel precision. 2.1.2.12. Geometric Partitioning (GPM) In VVC, geometric partitioning modes are supported for inter-frame prediction. Geometric partitioning modes are transmitted via signaling as a merge mode using CU-level flags. Other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. A total of 64 partitions are supported for each possible CU size w×h=2. m ×2 n , where m,n∈{3…6} does not include 8x64 and 64x8. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line. Figure 23 The position of the dividing line is mathematically derived from the angle and offset parameters of a specific segment. Each part of the geometric segment in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, meaning each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, each CU requires only two motion-compensated predictions. If a geometric segmentation pattern is used for the current CU, the geometric segmentation pattern (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate sizes is explicitly transmitted in SPS, and syntax binarization is specified for the GPM merge indices. After predicting each part of the geometric segmentation, a blending process with adaptive weights is used to adjust the sample values ​​along the geometric segmentation edges. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction patterns. Finally, the motion field of the CU predicted using the geometric segmentation pattern is stored. 2.1.2.12.1. Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for the geometric segmentation pattern. These motion vectors in... Figure 24 The value is marked with "x". If the corresponding LX motion vector for the nth extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional predicted motion vector for the geometric segmentation pattern. 2.1.2.12.2. Blending along geometric segmentation edges Figure 25 An exemplary generation of the bending weight w0 using the geometric segmentation pattern is shown. After predicting each part of the geometric segmentation using its own motion, a blend is applied to the two predicted signals to derive samples around the geometric segmentation edge. The blend weight for each location of the CU is derived based on the distance between the individual location and the segmentation edge. The distance from position (x, y) to the segmentation edge is derived as follows: Where i,j are indices for the angle and offset used for geometric segmentation, which depend on the geometric segmentation index transmitted via signal. ρ x,j and ρ y,j The sign depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-25) w1(x,y)=1-w0(x,y)(2-27). partIdx depends on the angle index i. An example of weight w0 is shown below. 2.1.2.12.3. Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and Mv, a combination of Mv1 and Mv2, are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern. The type of motion vector stored for each individual location in the sports field is determined as follows: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-28) where motionIdx equals d(4x+2,4y+2). partIdx depends on the angle index i. If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the combination of Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector. 2) Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored. 2.1.2.12.4. GPM with inter-frame and intra-frame prediction (GPM inter-frame - intra-frame) Utilizing GPM's inter- and intra-frame approach, in addition to selecting merge candidates for each non-rectangular partitioned region in the CU where GPM is applied, a predefined intra-prediction mode for the geometric partition line can also be selected. In the proposed method, for each GPM-separated region, a flag from the encoder is used to determine whether it is an intra-prediction mode or an inter-prediction mode. When it is an inter-prediction mode, the unidirectional prediction signal is generated from the MV from the merge candidate list. On the other hand, when it is an intra-prediction mode, for the intra-prediction mode specified by the index from the encoder, the unidirectional prediction signal is generated from neighboring pixels. The possible variations in intra-prediction modes are limited by the geometry. Finally, the two unidirectional prediction signals are mixed in the same manner as with ordinary GPM. 2.1.3. Screen Content Encoding / Decoding Tools 2.1.3.1. Intra-Block Copying (IBC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer-precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less. On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-Merge mode, block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed. In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4x4 sub-blocks. For larger current blocks, a hash key match with a reference block is determined when all hash keys of all 4x4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected. In block matching search, the search scope is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is transmitted via a flag, and it can be transmitted via a flag as either IBC AMVP mode or IBC Skip / Merge mode, as follows: –IBC Skip / Merge Mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and paired candidates. –IBC AMVP Mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left nearest neighbor and one from the top nearest neighbor (if IBC encoded and decoded). When either nearest neighbor is unavailable, the default block vector is used as the prediction value. A flag is transmitted via signaling to indicate the index of the block vector prediction value. 2.1.3.1.1. IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined region, which includes the current CTU region and some regions of the left CTU. Figure 26 The reference area of ​​the IBC mode is shown, where each block represents a 64x64 lumen sample unit. Based on the current location of the encoding / decoding CU within the current CTU, the following applies: – If the current block falls within the top-left 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, reference samples in the bottom-right 64x64 block of the left CTU can be used in CPR mode. The current block can also use CPR mode to reference reference samples in the bottom-left 64x64 block of the left CTU and reference samples in the top-right 64x64 block of the left CTU. – If the current block falls within the upper right 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, if the brightness position (0,64) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference the reference samples in the lower left and lower right 64x64 blocks of the left CTU; otherwise, the current block can also reference the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls within the lower left 64x64 block of the current CTU, then in addition to the samples already reconstructed in the current CTU, if the brightness position (64,0) relative to the current CTU has not yet been reconstructed, the current block can also use CPR mode to reference reference samples in the upper right and lower right 64x64 blocks of the left CTU. Otherwise, the current block can also use CPR mode to reference reference samples in the lower right 64x64 block of the left CTU. – If the current block falls within the bottom right 64x64 block of the current CTU, then only CPR mode can be used to reference samples that have already been reconstructed in the current CTU. This restriction allows for the use of local on-chip memory to implement IBC mode for hardware implementations. 2.1.3.1.2. Interaction between IBC and other codec tools The interaction between the IBC mode in VVC and other inter-frame coding and decoding tools (such as Paired Merge Candidate, History-Based Motion Vector Prediction (HMVP), Intra / Inter Joint Prediction Mode (CIIP), Merge Mode with Motion Vector Difference (MMVD), and Geometric Partitioning Mode (GPM)) is as follows: IBC can be used with pairwise merge candidates and HMVP. New pairwise IBC merge candidates can be generated by averaging two IBC merge candidates. For HMVP, IBC movements are inserted into the history cache for future reference. –IBC cannot be used in combination with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM. – When using DUAL_TREE segmentation, IBC is not allowed for chroma codec blocks. Unlike HEVC screen content encoding / decoding extensions, the current image is no longer included as one of the reference images in reference image list 0 for IBC prediction. The derivation process of motion vectors for the IBC mode excludes all neighboring blocks in inter-frame modes and vice versa. The following IBC design aspects are applied: –IBC shares the same process as regular MV Merge, including pairwise Merge candidates and history-based motion predictions, but does not allow TMVP and zero vectors because they are invalid for IBC mode. – Separate HMVP caches (5 candidates each) are used for regular MV and IBC. – Block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that no invalid vectors exist in the bitstream, and if a Merge candidate is invalid (out of range or 0), then the Merge must not be used. As described below, such bitstream consistency constraints are expressed based on virtual buffers. – For deblocking, IBC is processed as an inter-frame mode. – If the current block is encoded or decoded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is transmitted via signaling to indicate only whether the MV is an inter-frame pixel or 4 integer pixels. The number of IBC Merge candidates can be transmitted separately in the strip header from the number of regular Merge candidates, sub-block Merge candidates, and geometric Merge candidates. The concept of a virtual cache is used to describe the permissible reference region and valid block vector for IBC prediction modes. Representing the CTU size as ctbSize, the virtual cache ibcBuf has a width wIbcBuf = 128x128 / ctbSize and a height hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of the VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64). The virtual IBC cache ibcBuf is maintained as follows. - At the beginning of decoding each CTU line, refresh the entire ibcBuf with an invalid value of -1. – At the beginning of decoding the VPDU(xVPDU, yVPDU) relative to the top left corner of the image, set ibcBuf[x][y] = -1, where x = xVPDU%wIbcBuf, ..., xVPDU%wIbcBuf + W v -1;y=yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1. – After decoding the CU containing (x, y) relative to the top-left corner of the image, set ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]. For a block covering coordinates (x, y), it is valid if the following are true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] must not be equal to -1. 2.1.3.2. Block Differential Pulse Codec Modulation (BDPCM) VVC supports Block Differential Pulse Codec Modulation (BDPCM) for screen content encoding and decoding. At the sequence level, the BDPCM enable flag is signaled in SPS; this flag is only signaled when transform skip mode is enabled in SPS (described in the next section). When BDPCM is enabled, if the CU size is less than or equal to MaxTsSize multiplied by MaxTsSize in terms of luma samples, and if the CU is intra-coded, a flag is transmitted at the CU level, where MaxTsSize is the maximum block size allowed for transform skip mode. This flag indicates whether regular intra-coding or BDPCM is used. If BDPCM is used, a BDPCM prediction direction flag is transmitted to indicate whether the prediction is horizontal or vertical. The block is then predicted using a regular horizontal or vertical intra-prediction process with unfiltered reference samples. The residuals are quantized, and each quantized residual is encoded and decoded as the difference between its predicted value and its previously encoded residual (i.e., the residual at a neighboring horizontal or vertical location, depending on the BDPCM prediction direction). For a block with dimensions M (height) × N (width), let ri,j Let Q(r) be the prediction residual, where 0 ≤ i ≤ M-1 and 0 ≤ j ≤ N-1. i,j ), 0≤i≤M-1, 0≤j≤N-1 represent residuals r i,j The quantized version. BDPCM is applied to the quantized residuals to produce a value with element-wise... Modified M×N array in It is predicted from the quantized residual values ​​of its neighbors. For the vertical BDPCM prediction mode, for 0≤j≤(N-1), the following is used for derivation. For the horizontal BDPCM prediction mode, for 0 ≤ i ≤ (M-1), the following is used for derivation. On the decoder side, the above process is reversed to compute Q(r). i,j ), 0≤i≤M-1, 0≤j≤N-1, as follows: If using vertical BDPCM (2-31) If using horizontal BDPCM(2-32) The residual Q of dequantization -1 (Q(r i,j The values ​​are added to the intra-block prediction values ​​to produce reconstructed sample values. Predicted quantized residual values The residual encoding / decoding process, identical to that used in transform-skip mode residual encoding / decoding, is sent to the decoder. For lossless encoding / decoding, if `slice_ts_residual_coding_disabled_flag` is set to 1, the quantized residual values ​​are sent to the decoder using regular transform residual encoding / decoding. For MPM modes encoded / decoded for future intra-frame modes, if the BDPCM prediction directions are horizontal or vertical, the horizontal or vertical prediction mode is stored for the CU encoded / decoded using BDPCM. For deblocking, if both blocks on either side of a block boundary are encoded / decoded using BDPCM, that particular block boundary is not deblocked. 2.1.3.3. Residual Encoding and Decoding for Transform Skip Mode VVC allows transform skip mode to be used for luma blocks with a maximum size of MaxTsSize multiplied by MaxTsSize, where the value of MaxTsSize is transmitted through the signal in the PPS and can be up to 32. When the CU is encoded and decoded in transform skip mode, a transform skip residual encoding and decoding process is used to quantize and encode its prediction residuals. This process is modified from the transform coefficient encoding and decoding process. In transform skip mode, the residuals of the TU are also encoded and decoded in units of non-overlapping sub-blocks of size 4x4. For better encoding and decoding efficiency, some modifications were made to customize the residual encoding and decoding process for the characteristics of the residual signal. The following summarizes the differences between transform skip residual encoding and decoding and conventional transform residual encoding and decoding: – The forward scan sequence is applied to scan sub-blocks within the transform block and positions within sub-blocks; - No signaling for the final (x, y) position; – When all previous flags are equal to 0, coded_sub_block_flag is encoded and decoded for each sub-block except for the last sub-block; The context modeling of sig_coeff_flag uses a reduced template, and the context model of sig_coeff_flag depends on the top neighbor and the left neighbor. The context model of the –abs_level_gt1 flag also depends on the left-hand sig_coeff_flag value and the top-hand sig_coeff_flag value; –par_level_flag uses only one context model; – Additional flags greater than 3, 5, 7, and 9 are transmitted via signaling to indicate coefficient levels, with each flag representing a context; - Binarization of the remainder value is derived using a fixed order parameter of 1 (rice). The context model of the symbol flag is determined based on the left neighbor and the top neighbor, and the symbol flag is parsed after sig_coeff_flag to keep all the context-encoded bits together. For each subblock, if coded_subblock_flag equals 1 (i.e., there is at least one non-zero quantized residual in the subblock), then the quantized residual-level encoding and decoding is performed in three scan passes (see...). Figure 27 ): – First scan pass: The significance flag (sig_coeff_flag), the sign flag (coeff_sign_flag), the flag indicating an absolute level greater than 1 (abs_level_gtx_flag[0]), and the parity flag (par_level_flag) are encoded and decoded. For a given scan position, if sig_coeff_flag equals 1, then coeff_sign_flag is encoded and decoded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[0] equals 1, then par_level_flag is encoded and decoded separately to specify the parity of the absolute level. - Scan cycles greater than x: For each scan position where the absolute level is greater than 1, up to four abs_level_gtx_flag[i] (i = 1...4) are encoded or decoded to indicate whether the absolute level at the given position is greater than 3, 5, 7 or 9 respectively. – Remainder scan iterations: The remainders of the absolute level abs_remainder are encoded and decoded in bypass mode. The remainders of the absolute level are binarized using a fixed Rice parameter value of 1. The bits in scan cycles #1 and #2 (the first scan cycle and scan cycles greater than x) are context-coded until the maximum number of context-coded bits in the TU is exhausted. The maximum number of context-coded bits in the residual block is limited to 1.75 * block_width * block_height, or equivalently, an average of 1.75 context-coded bits per sample location. The bits in the last scan cycle (remainder scan cycle) are bypass-coded. The variable RemCcbs is initially set to the maximum number of context-coded bits in the block and decrements by 1 each time context-coded bits are encoded. When RemCcbs is greater than or equal to 4, syntax elements (including sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag) in the first encoding cycle are encoded using context-coded bits. If RemCcbs becomes less than 4 during the first pass of encoding / decoding, then the remaining coefficients that have not yet been encoded / decoded in the first pass are encoded / decoded in the remainder scan pass (pass #3). After the first pass of encoding / decoding is completed, if RemCcbs is greater than or equal to 4, the syntax elements (including abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag) in the second pass of encoding / decoding are encoded / decoded using the binary bits encoded / decoded in the context. If RemCcbs becomes less than 4 when encoding / decoding the second pass, the remaining coefficients that have not yet been encoded / decoded in the second pass are encoded / decoded in the remainder scan pass (pass #3). Figure 27 This illustrates the transformation skipping the residual encoding / decoding process. An asterisk marks the position where the binary bits encoded / decoded by the context are exhausted; at this point, bypass encoding / decoding is used to encode / decode all remaining binary bits. Furthermore, for blocks not encoded in BDPCM mode, a magnitude mapping mechanism is applied to skip residual encoding until the maximum number of bits encoded in the context is reached. Magnitude mapping uses the magnitudes of the top and left neighboring coefficients to predict the current coefficient magnitude, thus reducing signal transmission costs. For a given residual position, `absCoeff` is represented as the absolute coefficient magnitude before mapping, and `absCoeffMod` is represented as the coefficient magnitude after mapping. `X0` represents the absolute coefficient magnitude at the left neighboring position, and `X1` represents the absolute coefficient magnitude at the top neighboring position. Magnitude mapping is performed as follows: pred = max(X0,X1); if (absCoeff == pred) absCoeffMod = 1; else absCoeffMod = (absCoeff <pred)?absCoeff+1:absCoeff. Then, the absCoeffMod value is encoded and decoded as described above. After all context-encoded bits have been exhausted, amplitude mapping is disabled for all remaining scan positions in the current block. 2.1.3.4. Palette Mode In VVC, palette mode is used for screen content encoding and decoding in all chroma formats supported in the 4:4:4 gradation (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome). When palette mode is enabled, if the CU size is less than or equal to 64x64 and the number of samples in the CU is greater than 16, a flag is transmitted at the CU level to indicate whether palette mode is used. Considering that applying palette mode on small CUs introduces negligible encoding / decoding gain and adds extra complexity to small blocks, palette mode is disabled for CUs with 16 or fewer samples. A palette-encoded codec unit (CU) is treated as a prediction mode distinct from intra-prediction, inter-prediction, and intra-block copy (IBC) modes. If a palette mode is used, the sample values ​​in the CU are represented by a set of representative color values. This set is called the palette. For sample values ​​that are close to the palette colors, the palette index is transmitted via signaling. Samples outside the palette can also be specified by transmitting an escape symbol via signaling. For samples encoded and decoded within the CU using the escape symbol, their component values ​​are transmitted directly via signaling using (possibly) quantized component values. This is in... Figure 28 The jump symbol is shown in the figure. The quantized jump symbol is binarized using the fifth-order Exp-Golomb binarization process (EG5). To encode and decode the palette, palette predictions are maintained. For non-wavefront cases, palette predictions are initialized to 0 at the beginning of each stripe. For WPP cases, the palette prediction at the beginning of each CTU line is initialized to a prediction derived from the first CTU in the previous CTU line, ensuring consistency between the initialization scheme of the palette predictions and CABAC synchronization. For each entry in the palette prediction, a reuse flag is signaled to indicate whether it is part of the current palette in the CU. The reuse flag is transmitted using zero run length encoding and decoding. Subsequently, the number of new palette entries and the component values ​​of the new palette entries are signaled. After encoding the palette-encoded CU, the palette predictions are updated using the current palette, and entries from previous palette predictions not reused in the current palette are added to the end of the new palette predictions until the maximum allowed size is reached. A skip flag is signaled for each CU to indicate the presence of a skip symbol in the current CU. If a jump symbol exists, the palette table is incremented by one, and the last index is assigned as the jump symbol. Similar to the coefficient groups (CGs) used in transform coefficient encoding and decoding, the CU encoded and decoded in palette mode is divided into multiple row-based coefficient groups. Each row-based coefficient group consists of m samples (i.e., m = 16), where the index run, palette index value, and quantized color for the skip mode are encoded / parsed sequentially for each CG. As in HEVC, horizontal or vertical traversal scans can be applied to scan the samples, such as... Figure 29 As shown. The encoding order of the palette run-length encoding / decoding in each segment is as follows: For each sample location, one context-encoded bit `run_copy_flag = 0` is signaled to indicate whether the pixel has the same pattern as the previous sample location; that is, whether both the previously scanned sample and the current sample are run-length type `COPY_ABOVE`, or whether both the previously scanned sample and the current sample are run-length type `INDEX` and have the same index value. Otherwise, `run_copy_flag = 1` is signaled. If the current sample and the previous sample have different patterns, one context-encoded bit `copy_above_palette_indices_flag` is signaled to indicate the run-length type of the current sample, i.e., `INDEX` or `COPY_ABOVE`. Here, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not need to resolve the run-length type because the default is to use the `INDEX` mode. Similarly, if the previously resolved run type is COPY_ABOVE, the decoder does not need to resolve the run type. After palette run encoding and decoding of samples in one encoding / decoding pass, the index values ​​(for INDEX mode) and quantized jump colors are grouped and encoded / decoded in another encoding / decoding pass using CABAC bypass encoding / decoding. This separation of context-encoded and bypass-encoded bits improves throughput within each line of CG. For stripes with dual luma / chroma trees, the palette is applied separately to the luma (Y component) and chroma (Cb and Cr components), where luma palette entries contain only Y values ​​and chroma palette entries contain both Cb and Cr values. For single-tree stripes, the palette is applied jointly to the Y, Cb, and Cr components, meaning each entry in the palette contains Y, Cb, and Cr values, unless the CU is encoded using a local dual-tree, in which case luma and chroma encoding / decoding are handled separately. In this case, if the corresponding luma or chroma blocks are encoded using a palette mode, their palettes are applied in a manner similar to the dual-tree case (this relates to non-4:4:4 encoding / decoding and will be further explained in 2.1.3.4.1). For stripes encoded with dual-trees, the maximum palette prediction size is 63, and the maximum palette table size used for encoding and decoding the current CU is 31. For stripes encoded with dual-trees, the maximum prediction size and maximum palette table size for each of the luma and chroma palettes are halved, i.e., the maximum prediction size is 31 and the maximum table size is 15. For deblocking, palette-encoded blocks on either side of the block boundary are not deblocked. 2.1.3.4.1. Palette mode for non-4:4:4 content The palette mode in VVC supports all chroma formats in a similar way to the palette mode in HEVC SCC. For non-4:4:4 content, the following customizations apply: 1. When transmitting a jump-out value for a given sample location via signal transmission, if the sample location has only a luminance component and no chrominance component due to chrominance downsampling, then only the luminance jump-out value is transmitted via signal transmission. This is the same as in HEVC SCC. 2. For local double-tree blocks, the palette pattern is applied to the block in the same way as the palette pattern is applied to single-tree blocks, with two exceptions: a. The palette prediction update process is slightly modified as follows. Since the local dual-tree block contains only the luma (or chroma) component, the prediction update process uses the value of the luma (or chroma) component transmitted through the signal and fills in the "missing" chroma (or luma) component by setting it to the default value (1 << (component bit depth - 1)). b. The maximum palette prediction size remains at 63 (because stripes are encoded and decoded using a single tree), but the maximum palette table size for luma / chroma blocks remains at 15 (because blocks are encoded and decoded using a separate palette). 3. For monochrome palette mode, the number of color components in the palette-encoded block is set to 1 instead of 3. 2.1.3.4.2. Encoder Algorithm for Palette Mode On the encoder side, the following steps are used to generate the current CU's palette table. 1. First, to derive the initial entries in the current CU's palette table, a simplified K-means clustering approach is applied. The current CU's palette table is initialized as an empty table. For each sample location in the CU, the SAD (Sum of Averages) between that sample and each palette table entry is calculated, and the minimum SAD among all palette table entries is obtained. If the minimum SAD is less than a predefined error limit (errorLimit), the current sample is clustered with the palette table entry with the minimum SAD. Otherwise, a new palette table entry is created. The threshold (errorLimit) is QP-related and is obtained from a lookup table containing 57 elements covering the entire QP range. After all samples in the current CU have been processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entries after the 31st entry are discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroids of each cluster from step 1 or using one of the palette colors from the palette predictions. The option with lower rate-distortion cost is selected as the final color of the palette table. If a cluster has only a single sample point and the corresponding palette entry is not in the palette predictions, the corresponding sample point is converted to a jump sign in the next step. 3. The resulting palette table contains some new entries from the centroids of the clusters in step 1 and some entries from the palette predictions. Therefore, the table is reordered again so that all the new entries (i.e., centroids) are placed at the beginning of the table, followed by the entries from the palette predictions. Given the current palette table of the CU, the encoder selects the palette index for each sample location in the CU. For each sample location, the encoder checks all index values ​​corresponding to the palette table entry and the RD cost of the index representing the jump symbol, and selects the index with the minimum RD cost using the following equation: RD cost = Distortion × (isChroma? 0.8 : 1) + lambda × bits after bypass encoding / decoding (2-33) After the index map of the current CU is determined, each entry in the palette table is checked to see if it is used by at least one sample location in the CU. Any unused palette entries are deleted. After the index map of the current CU is determined, mesh RD optimization is applied to find the optimal values ​​for `run_copy_flag` and run type for each sample location by comparing the RD costs of three options: same as the previously scanned location, run type `COPY_ABOVE`, or run type `INDEX`. When calculating the SAD value, the sample value is reduced to 8 bits unless the CU is encoded / decoded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. Furthermore, in the case of lossless encoding / decoding, only the rate is used in the rate-distortion optimization step described above (because lossless encoding / decoding does not produce distortion). 2.1.3.5. Adaptive Color Transformation In the HEVC SCC extension, Adaptive Color Transform (ACT) is applied to reduce redundancy among the three color components in the 444 chroma format. ACT is also adopted in the VVC standard to improve the encoding and decoding efficiency of the 444 chroma format. Similar to HEVC SCC, ACT performs loop color space conversion in the prediction residual domain by adaptively converting the residual from the input color space to the YCgCo space. Figure 30 The decoding flowchart for applying ACT is shown. Two color spaces are adaptively selected by transmitting an ACT flag at the CU level. When the flag is equal to 1, the CU residual is encoded / decoded in the YCgCo color space; otherwise, the CU residual is encoded / decoded in the original color space. Furthermore, similar to the HEVC ACT design, for inter-frame and IBC CUs, ACT is enabled only if there is at least one non-zero coefficient in the CU. For intra-frame CUs, ACT is enabled only if the chroma component selects the same intra-frame prediction mode (i.e., DM mode) as the luma component. 2.1.3.5.1. ACT Mode In the HEVC SCC extension, ACT supports both lossless and lossy encoding / decoding based on a lossless flag (i.e., cu_transquant_bypass_flag). However, no flag is transmitted via signaling in the bitstream to indicate whether lossy or lossless encoding / decoding is applied. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown below. Because the YCgCo-R transform is not normalized, QP adjustments (-5, 1, 3) are applied to the transform residuals of the Y, Cg, and Co components, respectively, to compensate for the change in dynamic range of the residual signals before and after the color transform. The adjusted quantization parameters only affect the quantization and dequantization of the residuals in the CU. For other encoding / decoding processes (e.g., deblocking), the original QP is still applied. Additionally, because forward and inverse color transforms require access to the residuals of all three components, ACT mode is always disabled for ISP modes with individual tree splits and different prediction block sizes for different color components. When ACT is applied, Transform Skip (TS) extended to the codec chroma residuals and Block Differential Pulse Codec Modulation (BDPCM) are also enabled. 2.1.3.5.2. ACT Fast Encoding Algorithm To avoid brute-force RD search in both the original and converted color spaces, the following fast encoding algorithm is applied to the VTM reference software to reduce encoder complexity when ACT is enabled. – The order in which RD checks are enabled / disabled in ACT depends on the original color space of the input video. For RGB video, the RD cost in ACT mode is checked first; for YCbCr video, the RD cost in non-ACT mode is checked first. The RD cost in the second color space is checked only if there is at least one non-zero coefficient in the first color space. – When a CU is acquired via a different split path, the same ACT enable / disable decision is reused. Specifically, when a CU is first encoded or decoded, the selected color space used to encode and decode the residual of a CU is stored. Then, when the same CU is acquired via another split path, instead of checking the RD costs of the two spaces, the stored color space decision is directly reused. The RD cost of the parent CU is used to determine whether to check the RD cost for the second color space of the current CU. For example, if the RD cost for the first color space of the parent CU is less than the RD cost for the second color space, then the second color space is not checked for the current CU. – To reduce the number of codec modes tested, the selected codec modes are shared between the two color spaces. Specifically, for intra-frame modes, pre-selected intra-frame mode candidates based on SATD-based intra-frame mode selection are shared between the two color spaces. For inter-frame and IBC modes, block vector search or motion estimation is performed only once. Block vectors and motion vectors are shared between the two color spaces. 2.1.3.6. Intra-Template Matching (IntraTMP) Intra-template matching prediction (IntraTM) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signaling, and the same prediction operation is performed on the decoder side. The prediction signal is obtained by comparing the L-shaped causal nearest neighbors of the current block with... Figure 31 This is generated by matching another block within a predefined search region, which includes: R1: Current CTU R2: Top Left CTU R3: Above CTU R4: Left CTU. SAD was used as the cost function. Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to ensure a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW SearchRange_h = a * BlkH Where "a" is a constant that controls the gain / complexity tradeoff. In practice, "a" equals 5. Enable intra-frame template matching for CUs with width and height dimensions less than or equal to 64. The maximum CU size for intra-frame template matching is configurable. When DIMD is not used in the current CU, the intra-template matching prediction mode is transmitted at the CU level via a dedicated flag. 2.1.3.6.1. Using block vectors derived from IntraTMP for IBC The block vector (BV) derived from IntraTMP (Intra-Temporal Match Prediction) is used for IntraBlock Copy (IBC). The IntraTMP BV and IBC BV of the stored neighboring blocks are used as spatial BV candidates in the construction of the IBC candidate list. 2.1.3.6.2. Direct Block Vector (DBV) Mode for Chromaticity Prediction For the chroma component, when the chroma dual tree is activated in an intra-strip, if one of the luma blocks (five locations) is encoded / decoded using MODE_IBC, its block vector bvL is used and scaled to derive the chroma block vector bvC. The scaling factor depends on the chroma format sampling structure. Then, by using the position (xCb, yCb) of the current chroma block and its bvC, the corresponding offset position (xCb+bvC[0], yCb+bvC[1]) is determined, and block copy prediction is performed. Figure 32 Five locations in the reconstructed brightness sample are shown. Figure 33 The prediction process for the DBV pattern is shown. The CU level flag is transmitted via signal to indicate whether the proposed DBV mode is applied, as shown in Table 7. Table 7 shows the binarization process of intra_chroma_pred_mode in the proposed method. intra_chroma_pred_mode binary string Chroma intra-frame mode 0 11100 List [0] 1 11101 List [1] 2 11110 List [2] 3 11111 List [3] 4 110 DIMD chromaticity 5 10 DM 6 0 DBV 2.1.4. Transformation and Quantization 2.1.4.1. Large-size transformation with high-frequency zeroing In VVC, a large block size transform of up to 64×64 is enabled, primarily for higher resolution video such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both) equal to 64, high-frequency transform coefficients are zeroed out, leaving only low-frequency coefficients. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of transform coefficients are retained. Similarly, when N equals 64, only the first 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without zeroing out any values. Additionally, transform shifts are removed in transform skip mode. VTM also supports a configurable maximum transform size in SPS, allowing the encoder to flexibly select a transform size of up to 32 or 64 lengths depending on the specific implementation requirements. 2.1.4.2. Multiple Transform Selection (MTS) for Kernel Transformation In addition to DCT-II, which is used in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding and decoding of both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. Newly introduced transform matrices are DST-VII and DCT-VIII. Table 8 shows the basis functions of the selected DST / DCT. Table 8 - Transform basis functions for DCT-II / VIII and DSTVII for N-point inputs To maintain the orthogonality of the transformation matrices, the transformation matrices are quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformation coefficients within a 16-bit range, all coefficients must have 10 bits after both the horizontal and vertical transformations. To control the MTS scheme, separate enable flags are specified at the SPS level for intra-frame and inter-frame operations. When MTS is enabled at SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is applied only to luminance. MTS signaling is skipped when one of the following conditions is met. – The position of the last significant coefficient for luminance TB is less than 1 (i.e., DC only). – The last significant coefficient of luminance TB is located within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are transmitted via signaling to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 9. A unified transform selection for ISP and implicit MTS is used by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame explicit MTS and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform kernels. An 8-bit master transform kernel is used when transform matrix precision is involved. Therefore, all transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, and 32-point DST-7 and DCT-8) use an 8-bit master transform kernel. Table 9 - Transformation and Signaling Mapping Table To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed out for DST-7 and DCT-8 blocks with a size (width or height, or both) equal to 32. Only the coefficients in the 16x16 low-frequency region are retained. Similar to HEVC, block residuals can be encoded and decoded using transform skip mode. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Implicit MTS can also be enabled when MTS is enabled for inter-frame encoded blocks. 2.1.4.3. Low-Frequency Inseparable Transform (LFNST) In VVC, LFNST is applied between the forward master transform and quantization (at the encoder) and between dequantization and the inverse master transform (at the decoder). Within LFNST, either a 4x4 or 8x8 inseparable transform is applied depending on the block size. For example, a 4x4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and an 8x8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). Figure 34 The Low Frequency Inseparable Transform (LFNST) process is illustrated. Using the input as an example, the application of the inseparable transform used in LFNST is described below. To apply 4x4 LFNST, a 4x4 input block X... is first represented as a vector The non-separable transform is computed as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. Subsequently, using the scan order (horizontal, vertical or diagonal) for the block, the 16x1 coefficient vector is reorganized into 4x4 blocks. Coefficients with smaller indices will be placed at positions with smaller scan indices in the 4x4 coefficient block. 2.1.4.3.1. Reduced Non-Separable Transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, the dimension of the non-separable transform matrix needs to be reduced to minimize the computational complexity and the memory space for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in the LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (for 8x8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix is not an NxN matrix but becomes an R×N matrix as follows: The R rows of the transform form R bases in N-dimensional space. The inverse transform matrix of RT is the transpose of its forward transform. For 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix (i.e., the size of the regular 8x8 inseparable transform matrix) is reduced to a 16x48 direct matrix. Therefore, the 48×16 inverse RST matrix is ​​used on the decoder side to generate the core (primary) transform coefficients in the upper left region of the 8×8. When the 16x48 matrix is ​​applied instead of the 16x64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4x4 blocks in the upper left 8x8 block, excluding the lower right 4x4 block. Thanks to the reduced dimension, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB with a reasonable performance degradation. To reduce complexity, LFNST is only allowed to be applied if all coefficients outside the first coefficient subgroup are insignificant. Therefore, when LFNST is applied, all primary transform coefficients must be zero. This allows LFNST indexing to be conditional on the last salient position, thus avoiding the extra coefficient scan required in the current LFNST design, which is only used to check salient coefficients at specific locations. The worst-case handling of LFNST (in terms of multiplication per pixel) restricts the indivisible transforms of 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In these cases, the last salient scan position must be less than 8 when LFNST is applied, and less than 16 for other sizes. For blocks of shapes 4xN and Nx4 where N>8, the proposed restrictions mean that LFNST is now applied only once, and only to the top-left 4x4 region. Since all principal coefficients are zero when LFNST is applied, the number of operands required for the principal transform is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when testing the LFNST transform. Rate-distortion optimized quantization must be performed on at most the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.1.4.3.2. LFNST Transform Selection There are a total of 4 transform sets in LFNST, and each transform set uses 2 inseparable transform matrices (kernels). The mapping from intra-prediction modes to transform sets is predefined, as shown in Table 10. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable quadratic transform candidate is further specified by an LFNST index explicitly transmitted via signaling. This index is transmitted via signaling once in the bitstream for each intra-CU after the transform coefficients. Table 10 - Transformation Selection Table IntraPredMode Transform set index IntraPredMode<0 1 0<=IntraPredMode<=1 0 2<=IntraPredMode<=12 1 13<=IntraPredMode<=23 2 24<=IntraPredMode<=44 3 45<=IntraPredMode<=55 2 56<=IntraPredMode<=80 1 81<=IntraPredMode<=83 0 2.1.4.3.3. LFNST Index Signaling and Interaction with Other Tools Since LFNST is only allowed to be applied if all coefficients outside the first coefficient subgroup are insignificant, LFNST index encoding / decoding depends on the position of the last significant coefficient. Furthermore, the LFNST index is context-coded but not dependent on the intra-prediction mode, and only the first bit is context-coded. Additionally, LFNST is applied to intra-CUs in both intra-slice and inter-slice scenarios, and to both luma and chroma. If dual-tree is enabled, the luma and chroma LFNST indices are transmitted separately via signaling. For inter-slice (dual-tree is disabled), a single LFNST index is transmitted via signaling and used for both luma and chroma. Considering that large CUs larger than 64x64 are implicitly partitioned (TU fragmentation) due to the existing maximum transform size limit (64x64), LFNST index search can quadruple the data buffer for a certain number of decoding pipeline stages. Therefore, the maximum allowed size for LFNST is limited to 64x64. Note that LFNST is only enabled using DCT2. LFNST index signaling is placed before MTS index signaling. The use of scaling matrices for perceptual quantization is not obvious, as scaling matrices specified for the master matrix might be useful for LFNST coefficients. Therefore, scaling matrices are not permitted for LFNST coefficients. Chromatic LFNST is not applied for single-tree splitting mode. 2.1.4.4. Subblock Transformation (SBT) In VTM, a sub-block transform is introduced for inter-frame prediction CUs. In this transform mode, only a sub-part of the residual block is encoded or decoded for the CU. When the inter-frame prediction CU has cu_cbf equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded or decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is encoded or decoded using a presumed adaptive transform, and the other portion of the residual block is zeroed out. When SBTs are used in inter-frame encoding / decoding CUs, the SBT type and SBT position information are transmitted via signals in the bitstream. For example... Figure 35As shown, there are two SBT types and two SBT locations. For SBT-V (or SBT-H), the TU width (or height) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. A 2:2 partition is like a binary tree (BT) partition, while a 1:3 / 3:1 partition is like an asymmetric binary tree (ABT) partition. In an ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 in the luma sample, a 1:3 / 3:1 partition along that dimension is prohibited. There can be a maximum of 8 SBT modes for a CU. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). Two positions in SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are... Figure 35 The subblock transformation is specified. For example, the horizontal and vertical transformations for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transformations in both dimensions are set to DCT-2. Therefore, the subblock transformation jointly specifies the TU slicing, cbf, and horizontal and vertical kernel transformation types of the residual block. SBT should not be used in CUs encoded and decoded using a combination of inter-frame and intra-frame modes. 2.1.4.5. Resetting the maximum transformation size and transformation coefficients to zero. The CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the largest intra-codec block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform, no canonical zeroing operation is applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are canonically zeroed. 2.1.4.6. Enhanced MTS for Intra-Frame Coding and Decoding In the current VVC design, only the DST7 and DCT8 transform cores used for intra-frame and inter-frame encoding / decoding are utilized for MTS. An additional master transform, including DCT5, DST4, DST1, and the identity transform (IDT), is employed. Furthermore, the MTS set depends on the TU size and intra-frame mode information. Sixteen different TU sizes are considered, and for each TU size, five different categories are considered based on the intra-frame mode information. For each category, four different transform pairs are considered, the same as the transform pairs for VVC. Note that although a total of 80 different categories are considered, some of these different categories typically share the exact same transform set. Therefore, there are 58 unique entries in the resulting LUT (less than 80). For angular modes, the joint symmetry of TU shape and intra-frame prediction is considered. Therefore, mode i (i>34) with TU shape AxB will be mapped to the same category as mode j = (68-i) with TU shape BxA. However, for each transform pair, the order of the horizontal and vertical transform kernels is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical and horizontal transform kernels are swapped. For wide-angle modes, the most recent regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 through 80. The MTS index [0,3] is transmitted via signaling using 2-bit fixed-length encoding and decoding. 2.1.4.7. Quadratic Transformation: LFNST Extension with Large Kernel The LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is extended to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra-frame mode (predModeIntra) is derived according to the following formula: For predModeIntra < 2, lfnstTrSetIdx equals 2. For predModeIntra in [0,34], lfnstTrSetIdx = predModeIntra For predModeIntra in [35, 66], lfnstTrSetIdx = 68 – predModeIntra • Define three distinct kernels, LFNST4, LFNST8, and LFNST16, to indicate the LFNST kernel set, which are applied to 4xN / Nx4 (N≥4), 8xN / Nx8 (N≥8), and M x N (M, N≥16), respectively. The kernel dimension is specified by the following formula: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96). Forward LFNST is applied to the upper left low-frequency region, known as the region of interest (ROI). When LFNST is applied, the principal transform coefficients in regions other than the ROI are zeroed out, which is unchanged from the VVC standard. The ROI of LFNST16 is Figure 36 The diagram is depicted in the image. It consists of six 4x4 sub-blocks arranged sequentially in scan order. Since the number of input samples is 96, the transformation matrix of the forward LFNST16 can be Rx96. In this contribution, R is chosen to be 32, and 32 coefficients (two 4x4 sub-blocks) are generated accordingly from the forward LFNST16, placed in the order of coefficient scan. The ROI of LFNST8 is Figure 37 The results are shown in the diagram. The positive LFNST8 matrix can be Rx64, and R is chosen to be 32. The generated coefficients are positioned in the same manner as LFNST16. The mapping from intra-prediction modes to these sets is shown in the table below. Table 11. Mapping of Intra-Prediction Modes to LFNST Set Indices Intra-prediction mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST collection index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra-prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST collection index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra-prediction mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST collection index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.1.4.8. Indivisible Main Transform (NSPT) for Intra-Frame Encoding and Decoding For block sizes of 4x4, 4x8, 8x4, and 8x8, DCT-II+LFNST is replaced by NSPT. NSPT follows the design of LFNST, with 3 candidates and 35 sets, selected based on intra-frame mode. The kernel size is as follows: • NSPT4x4: 16x16; • NSPT4x8 / NSPT8x4: 32x20; ·NSPT8x8: 64x32. Therefore, 12 and 32 coefficients for NSPT4x8 / NSPT8x4 and NSPT8x8, respectively, were cleared to zero. 2.1.4.9. Symbolic Prediction The basic idea of ​​coefficient sign prediction methods (JVET-D0031 and JVET-J0021) is to calculate the reconstruction residual for both positive and negative sign combinations for applicable transformation coefficients and select the assumption that minimizes the cost function. To derive the optimal sign, the cost function is defined as a measure of discontinuity across block boundaries, such as... Figure 38 As shown. The cost function is measured for all hypotheses, and the hypothesis with the minimum cost is selected as the predicted value for the sign of the coefficients. The cost function is defined as the sum of the absolute second derivatives over the residual domain of the upper row and the left column, as shown below: Where R is the reconstructed nearest neighbor, P is the prediction for the current block, and r is the residual hypothesis. The term (-R) -1 +2R0-P1) Each block can only be computed once, and only the residual assumption is subtracted. 3. Problem In ECM-7.0, OBMC can be applied to inter-frame encoded blocks, regardless of whether they are inter-frame AMVP encoded or inter-frame Merge encoded. For inter-frame AMVP encoded blocks, the application of OBMC can be indicated at the block level via signaling syntax flags. For inter-frame Merge encoded blocks, OBMC is implicitly presumed to be applied, regardless of block characteristics and the encoding / decoding information of neighboring blocks. However, there may be cases where inter-frame Merge encoded blocks do not prefer OBMC modes. For example, blocks containing sharp edges, few gradients, or few colors may not prefer OBMC modes. Block-level adaptive OBMC that inherits the OBMC on / off flags from its nearest neighbor can result in higher encoding / decoding gains. 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way. The terms “video unit” or “code-decoder unit” or “block” can refer to code-decoder tree block (CTB), code-decoder tree unit (CTU), code-decoder block (CB), CU, PU, ​​TU, PB, TB. In this disclosure, with respect to "blocks encoded or decoded using mode N", "mode N" can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a encoding / decoding technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.). It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to encoding / decoding tools also apply. 4.1. In one example, whether OBMC is applied to the current block can be inherited from the motion vector candidate. 1) For example, motion vector candidates can be candidates in the inter-frame merge list. a. For example, it can be in the inter-frame regular Merge list. b. For example, it can be in the inter-frame TM Merge list. c. For example, it can be in the inter-frame BM Merge list. d. For example, it can be in the inter-frame GEO Merge list. e. For example, it can be in the inter-frame CIIP Merge list. f. For example, it can be in the inter-frame MMVD Merge list. g. For example, it can be in the inter-frame affine Merge list. h. For example, it can be in the inter-frame SbTMVP Merge list. 2) For example, motion vector candidates can be candidates in the inter-frame AMVP list. a. For example, it can be in the list of regular AMVPs for inter-frame use. b. For example, it can be in the inter-frame AMVP-Merge list. c. For example, it can be in the list of inter-frame affine AMVPs. 3) For example, inheritance can be based on the type of motion candidate in the current block. a. For example, OBMC parameters (e.g., flags) can be inherited from the spatial nearest neighbor of the current block. b. For example, OBMC parameters (e.g., flags) can be inherited from spatial neighbors that are not adjacent to the current block. i. Alternative locations, whether and how they are inherited, can depend on the distance between the current block and non-adjacent blocks. c. For example, OBMC parameters (e.g., flags) can be inherited from motion candidates from the HMVP table. i. Alternative sites, whether and how they are inherited, can depend on the distance between the current block and the HMVP candidate. d. For example, for time-domain motion vectors, OBMC parameters (e.g., flags) can be set to default values. i. Alternatively, OBMC parameters can be inherited from the time-domain motion vector. e. For example, for paired motion vectors, OBMC parameters (e.g., flags) can be set to default values. i. Alternatively, OBMC parameters can be inherited from one direction of the motion vector that forms a pair of motion vectors. f. For example, for a zero motion vector, the OBMC parameter (e.g., a flag) can be set to a default value. g. For example, the default value can indicate that OBMC is used for the current block. h. For example, the default value can indicate that OBMC is not used for the current block. 4) For example, inheritance can be based on whether the motion candidate is LIC - decoded. a. For example, if the motion candidate is LIC - decoded, the inherited OBMC parameter (e.g., a flag) can be equal to a value indicating that OBMC is not applied to the current block. b. For example, if the motion candidate is not LIC - decoded, the inherited OBMC parameter (e.g., a flag) can be equal to a value indicating that OBMC is applied to the current block. 5) For example, inheritance can be based on the block dimension of the current block (assuming W and H represent the width and height of the current block). a. For example, in the case of satisfying at least one of the following conditions, the inherited OBMC parameter (e.g., a flag) can be equal to a value indicating that OBMC is not applied to the current block, i. W*H < T1 or W*H <= T1, ii. W < T2 or W <= T2, iii. H < T3 or H <= T3, iv. W / H < T4 or W / H <= T4, v. H / W < T5 or H / W <= T5, vi. W*H > T6 or W*H >= T6, vii. W > T7 or W >= T7, viii. H > T8 or H >= T8, ix. W / H > T9 or W / H >= T9, x. H / W > T10 or H / W >= T10. b. For example, the thresholds T1…T10 in item a can be constant values. i. For example, T1 = 32 or 64 or 16. ii. For example, T2 = 32 or 16 or 8. iii. For example, T3 = 32 or 16 or 8. iv. For example, T4 = 4 or 8 or 16. v. For example, T5 = 4 or 8 or 16. vi. For example, T6 = 32 or 64 or 128. vii. For example, T7 = 32 or 64 or 128. viii. For example, T8 = 32 or 64 or 128. ix. For example, T9 = 8, 16, or 32. x. For example, T10 = 8, 16, or 32. 6) For example, inheritance can be based on the prediction pattern of the current block. a. For example, if the current block is encoded or decoded using the following modes, then the OBMC parameters (e.g., flags) inherited for the current block can be inherited. i. Inter-frame Merge ii. Inter-frame AMVP iii. Regular inter-frame merge iv. MHP (e.g., MHP Merge and / or MHP AMVP) v.GEO (and / or its variants, such as GPM™, GPM MMVD, GPM inter-intra) vi. CIIP (and / or its variants, such as CIIP™) vii. Inter-frame MMVD viii. Sub-block encoding and decoding (e.g., affine Merge, affine AMVP, sbTMVP, etc.) ix. Affine MMVD x. Affine Merge xi. Inter-frame™ xii. Inter-frame BM (such as adaptive DMVR) xiii.AMVP-MERGE xiv.sbTMVP. b. For example, if the current block is encoded or decoded using the following modes, then the OBMC parameters (e.g., flags) inherited for the current block may not be inherited. i. Inter-frame AMVP ii.AMVP-MERGE iii. IBC Merge iv.IBC AMVP. 7) For example, inheritance can be based on the predicted patterns of motion candidates (e.g., neighboring blocks). a. For example, if a motion candidate is encoded or decoded using one of the following modes, then OBMC parameters (e.g., flags) for such a motion candidate for the current block can be inherited. i. Inter-frame Merge ii. Inter-frame AMVP iii. Regular inter-frame merge iv. MHP (e.g., MHP Merge and / or MHP AMVP) v.GEO (and / or its variants, such as GPM™, GPM MMVD, GPM inter-intra) vi. CIIP (and / or its variants, such as CIIP™) vii. Inter-frame MMVD viii. Sub-block encoding and decoding (e.g., affine Merge, affine AMVP, sbTMVP, etc.) ix. Affine MMVD x. Affine Merge xi. Affine AMVP xii. Inter-Frame™ xiii. Inter-frame BM (such as adaptive DMVR) xiv.AMVP-MERGE xv.sbTMVP xvi.IBC Merge xvii.IBC AMVP xviii.LIC. b. For example, if a motion candidate is encoded or decoded using one of the following modes, then OBMC parameters (e.g., flags) for such a motion candidate for the current block may not be inherited. i.LIC ii. Inter-frame AMVP iii.MHP AMVP iv. Affine AMVP v.AMVP-MERGE vi. IBC Merge vii.IBC AMVP. c. For example, if OBMC parameters (e.g., flags) are not inherited from motion candidates, then OBMC flags can be set to specific values. i. For example, a specific value may depend on the LIC flag of the motion candidate. ii. For example, a particular value may depend on the sub-block pattern of the motion candidate (affine and / or sbTMVP). iii. For example, a particular value may depend on the AMVP-MERGE pattern of the motion candidate. iv. For example, a particular value may depend on the block width and / or height of the current video unit. v. For example, a particular value can be fixed (e.g., 0 or 1). 8) For example, the inheritance of blocks encoded and decoded by MHP can depend on the prediction mode of the underlying assumptions. a. For example, if the fundamental assumption of a block encoded / decoded via MHP is that it is encoded / decoded via inter-frame Merge, i. For example, the OBMC parameters (e.g., flags) of the Merge candidates based on the basic assumption can be inherited into blocks encoded and decoded by MHP. ii. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is always used for blocks encoded or decoded by MHP. iii. Alternatively, OBMC parameters (e.g., flags) may be set to values ​​that indicate that OBMC is never used for blocks encoded or decoded by MHP. b. For example, if the fundamental assumption of a block encoded / decoded via MHP is that it is encoded / decoded via inter-frame AMVP, i. For example, OBMC parameters (e.g., flags) can be transmitted in the bitstream via signals, indicating whether the OBMC is used for blocks encoded and decoded by MHP. ii. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is always used for blocks encoded or decoded by MHP. iii. Alternatively, OBMC parameters (e.g., flags) may be set to values ​​that indicate that OBMC is never used for blocks encoded or decoded by MHP. 9) For example, the OBMC parameters (e.g., flags) of a block encoded or decoded by MHP can be set according to the prediction mode of the underlying assumptions. a. For example, it can depend on whether the basic assumptions of the blocks encoded and decoded by MHP are encoded and decoded in a specific sub-block mode (e.g., affine AMVP, affine Merge, sbTMVP, etc.). b. For example, based on the fundamental assumption that the code is encoded or decoded in a sub-block mode (e.g., sbTMVP and / or affine AMVP and / or affine Merge), the OBMC flag can be set to a value. c. For example, if the underlying assumption is that the code is encoded and decoded via sbTMVP and / or affine AMVP and / or affine Merge, the OBMC flag can be set to a fixed value (e.g., 1 or 0). d. For example, it can depend on whether the basic assumption of the block encoded / decoded by MHP is that it is encoded / decoded by LIC. i. For example, it can depend on whether the basic assumption of the block encoded / decoded by MHP is that it is encoded / decoded in AMVP-MERGE mode. 10) For example, the inheritance of blocks encoded and decoded by GEO / GPM (and / or its variants, such as GPM TM, GPM MMVD, GPM inter-intra, etc.) may depend on the motion candidate's OBMC parameters (e.g., flags). a. For example, the OBMC parameters (e.g., flags) of a Merge candidate in the regular Merge list can be copied to the corresponding GEO candidate in the GEO Merge list and can be inherited into the GEO / GPM encoded / decoded block. b. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is always used for blocks encoded / decoded via GEO / GPM. c. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is never used for blocks encoded or decoded by GEO / GPM. 11) For example, the inheritance of blocks encoded and decoded by affine Merge (and / or its variants, such as affine MMVD, affine DMVR, etc.) may depend on the OBMC parameters (e.g., flags) of the affine candidate. a. For example, OBMC parameters (e.g., flags) of affine Merge candidates can be inherited into blocks encoded and decoded by affine Merge. b. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is always used for blocks encoded and decoded via affine Merge. c. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is never used in blocks encoded or decoded via affine Merge. 12) For example, the inheritance of blocks encoded and decoded by sbTMVP Merge (and / or its variants, such as sbTMVP TM, sbTMVP DMVR, etc.) may depend on the OBMC parameters (e.g., flags) of the motion displacement candidates of the sbTMVP blocks. a. For example, OBMC parameters (e.g., flags) of motion displacement candidates can be inherited into blocks encoded and decoded by sbTMVP. b. Alternatively, the OBMC parameters (e.g., flags) of a sub-block in the corresponding CU in the reference frame (e.g., where the position of the corresponding CU is identified by motion displacement candidates) can be inherited into the block encoded and decoded by sbTMVP. i. For example, a sub-block can be located at the center of the corresponding CU. ii. For example, a sub-block can be located to the upper left of the corresponding CU. c. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is always used for blocks encoded and decoded by sbTMVP. d. Alternatively, OBMC parameters (e.g., flags) can be set to values ​​that indicate that OBMC is never used in blocks encoded or decoded by sbTMVP. 13) For example, inheritance can depend on more than one of the conditions listed above. a. For example, suppose the OBMC parameter (e.g., a flag) is inherited from the motion vector candidate. It can first check whether the current block is inter-merge encoded, whether it is not LIC encoded, and whether the block dimension is less than a certain number. Only when all these conditions are true can the inherited OBMC parameter be set to a value that indicates the OBMC to be applied to the current block. 4.2. In one example, the block's OBMC parameters (e.g., flags) can be stored in a cache. 1) For example, it can be stored in M×M (e.g., M=4 or 8) sub-blocks. 1) For example, the stored OBMC parameters can be used for encoding and decoding of subsequent blocks (e.g., OBMC parameter inheritance). 4.3. For example, if the OBMC flag is transmitted via signal for a video unit, the OBMC flag can be context-coded. 1) In one example, it can be encoded and decoded by at least two context models. 2) In one example, which context model is used may depend on the prediction mode / method of the current video unit. a. In one example, it can be based on whether the current block is encoded or decoded using IBC. b. In one example, it can be based on whether the current block is encoded or decoded in LIC mode. c. In one example, it can be based on whether the current block is encoded or decoded in AMVP-MERGE mode. d. In one example, it can be based on whether the current block is encoded or decoded using sub-block mode. e. In one example, it can be based on whether the current block is encoded or decoded in affine mode. f. In one example, it can be based on whether the current block is encoded or decoded using MHP (e.g., with the additional assumption that the size is greater than 0). 4.4. Whether and / or how the methods disclosed above can be applied to transmit signals at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header. 4.5. Whether and / or how the methods disclosed above can be applied to transmit signals at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU lines / strips / films / sub-images / other types of areas containing more than one sample point or pixel. 4.6. Whether and / or how the methods disclosed above are applied may depend on encoding / decoding information such as block size, color format, single / dual tree segmentation, color components, and stripe / picture type.

[0099] The terms "video unit," "code-decoder unit," or "block" can refer to a code-decoder tree block (CTB), a code-decoder tree unit (CTU), a code-decoder block (CB), a CU, a PU, a TU, a PB, or a TB. In this disclosure, with respect to "a block encoded and decoded using mode N," "mode N" can be a predictive mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a code-decoder technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame-inter, GPM intra-frame-intra, GPM inter-frame-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.).

[0100] Figure 39 A flowchart of a method 3900 for video processing according to an embodiment of the present disclosure is shown. Method 3900 is implemented during the conversion between a target video block and a video bitstream. It should be noted that embodiments of the present disclosure can be implemented individually or combined in any suitable manner.

[0101] At box 3910, for the conversion between video units and video unit bitstreams, whether Overlapping Sub-Block Motion Compensation (OBMC) is applied to the video unit is determined based on inheritance from the motion vector candidate. In other words, whether OBMC is applied to the current block is inherited from the motion vector candidate. Inheritance is based on one of the following: the prediction mode of the video unit or the prediction mode of the motion vector candidate.

[0102] At box 3920, the conversion is performed based on determination. In some embodiments, the conversion may include encoding video units from a bitstream. Alternatively or additionally, the conversion may include decoding video units from a bitstream. In this way, block-level adaptive OBMC that inherits OBMC parameters (e.g., OBMC on / off flags) from neighboring blocks can result in higher encoding / decoding gain and improved encoding / decoding efficiency. In some embodiments, OBMC parameters may be OBMC flags.

[0103] In some embodiments, if a video unit is encoded and decoded in a target mode, the OBMC parameters for that video unit are inherited. For example, the target mode includes at least one of the following: a multiple hypothesis prediction (MHP) mode or a sub-block encoding / decoding mode. For example, the MHP mode includes at least one of the following: an MHP Merge mode or an MHP inter-frame advanced motion vector prediction (AMVP) mode. In some other examples, the sub-block encoding / decoding mode includes at least one of the following: an affine Merge mode, an affine AMVP mode, or a sub-block-based temporal motion vector prediction (SbTMVP).

[0104] In some embodiments, if motion vector candidates are encoded and decoded through a target mode, the OBMC parameters for the motion vector candidates of the video unit are inherited. For example, the target mode includes at least one of the following: inter-frame merge mode, inter-frame AMVP mode, regular inter-frame merge mode, MHP mode, geometric (GEO) mode, a variant of GEO, inter-frame intra-frame joint prediction (CIIP) mode, a variant of CIIP, inter-frame merge mode with motion vector difference (MMVD) mode, sub-block encoding / decoding mode, affine MMVD, affine merge mode, inter-frame template matching (TM) mode, inter-frame block matching (BM) mode, AMVP-MERGE mode, sbTMVP mode, intra-frame block copy (IBC) merge mode, IBC AMVP mode, or local illumination compensation (LIC).

[0105] In some embodiments, the MHP mode includes at least one of the following: MHP Merge mode or MHP AMVP mode. Alternatively or additionally, the sub-block encoding / decoding mode includes at least one of the following: affine Merge mode, affine AMVP mode, or sbTMVP.

[0106] In some embodiments, if motion vector candidates are encoded and decoded through a target mode, the OBMC parameters for the motion vector candidates of video units are not inherited. For example, the target mode includes at least one of the following: LIC mode, inter-frame AMVP mode, MHP AMVP mode, affine AMVP mode, AMVP-MERGE mode, IBC Merge mode, or IBC AMVP mode.

[0107] In some embodiments, if the OBMC parameters are not inherited from the motion vector candidate, the OBMC flag is set to a target value. For example, the target value depends on the LIC flag of the motion vector candidate. As another example, the target value depends on at least one of the following of the motion vector candidate: subblock mode, affine mode, or sbTMVP. For example, the target value depends on the AMVP-MERGE mode of the motion vector candidate. In another example, the target value depends on at least one of the following of the video cell: block width or block height.

[0108] In some embodiments, the target value is fixed. For example, the target value is 0 or 1.

[0109] In some embodiments, the OBMC parameters of a block encoded / decoded by MHP are set according to the prediction mode of the underlying assumption. For example, the setting of the OBMC parameters depends on whether the underlying assumption of the block encoded / decoded by MHP is that it is encoded / decoded in a sub-block mode. For example, the sub-block mode includes at least one of the following: sbTMVP, affine AMVP, or affine Merge.

[0110] In some embodiments, if the underlying assumptions are encoded or decoded in at least one of the following ways: sbTMVP, affine AMVP, or affine Merge, the OBMC parameter is set to a fixed value. For example, the fixed value is 0 or 1.

[0111] In some embodiments, the OBMC parameter settings depend on whether the fundamental assumption of the MHP-encoded block is encoded and decoded in LIC mode. In some other embodiments, the OBMC parameter settings depend on whether the fundamental assumption of the MHP-encoded block is encoded and decoded in AMVP-MERGE mode.

[0112] In some embodiments, if OBMC parameters are indicated for a video unit, the OBMC parameters are encoded and decoded using context. For example, OBMC parameters are encoded and decoded using at least two context models.

[0113] In some embodiments, which context model is used for OBMC parameters depends on the prediction mode or prediction method of the video unit. In some other embodiments, which context model is used is based on whether the video unit is encoded / decoded using IBC. In some further embodiments, which context model is used is based on whether the video unit is encoded / decoded using LIC mode. Alternatively or additionally, which context model is used is based on whether the video unit is encoded / decoded using AMVP-MERGE mode. In some embodiments, which context model is used is based on whether the video unit is encoded / decoded using subblock mode. In some other embodiments, which context model is used is based on whether the video unit is encoded / decoded using affine mode.

[0114] In some embodiments, which context model is used is based on whether the video unit is MHP encoded or decoded. For example, which context model is used is based on whether the size of the additional assumption is greater than 0.

[0115] In some embodiments, the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header. In some embodiments, the indication of whether and / or how to determine whether the OBMC is applied to the current block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample or pixel. In some embodiments, method 3900 further includes: determining, based on the encoding and decoding information of the video unit, whether to determine whether OBMC is applied to the current block and / or how to determine whether OBMC is applied to the current block, wherein the encoding and decoding information includes at least one of the following: block size, color format, single and / or dual tree segmentation, color components, stripe type, or picture type.

[0116] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores video data generated as a bitstream by a method performed by means of a video processing apparatus. The method includes: determining whether overlapping sub-block-based motion compensation (OBMC) is applied to video units of the video based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; and generating a bitstream based on the determination.

[0117] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: determining whether overlapping sub-block-based motion compensation (OBMC) is applied to video units of the video based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.

[0118] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0119] Item 1. A video processing method comprising: a conversion between a video unit and a bitstream of the video unit for a video; determining whether overlapping sub-block-based motion compensation (OBMC) is applied to the video unit based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; and performing the conversion based on the determination.

[0120] Item 2. The method according to Item 1, wherein if the video unit is encoded or decoded in target mode, the OBMC parameters for the video unit are inherited.

[0121] Item 3. The method according to Item 2, wherein the target mode includes at least one of the following: a multiple hypothesis prediction (MHP) mode or a sub-block encoding / decoding mode.

[0122] Item 4. The method according to Item 3, wherein the MHP mode includes at least one of the following: MHP Merge mode or MHP Inter-Frame Advanced Motion Vector Prediction (AMVP) mode.

[0123] Item 5. The method according to Item 3, wherein the sub-block encoding / decoding mode includes at least one of the following: affine Merge mode, affine AMVP mode, or sub-block-based temporal motion vector prediction (sbTMVP).

[0124] Item 6. The method according to Item 1, wherein if the motion vector candidate is encoded or decoded through the target mode, the OBMC parameters of the motion vector candidate for the video unit are inherited.

[0125] Item 7. The method according to Item 6, wherein the target mode includes at least one of the following: inter-frame merge mode, inter-frame AMVP mode, regular inter-frame merge mode, MHP mode, geometry (GEO) mode, a variant of GEO, inter-frame intra-frame joint prediction (CIIP) mode, a variant of CIIP, inter-frame merge mode with motion vector difference (MMVD) mode, sub-block coding / decoding mode, affine MMVD, affine merge mode, inter-frame template matching (TM) mode, inter-frame block matching (BM) mode, AMVP-MERGE mode, sbTMVP mode, intra-frame block copy (IBC) merge mode, IBC AMVP mode, or local illumination compensation (LIC).

[0126] Item 8. The method according to Item 7, wherein the MHP mode includes at least one of the following: MHP Merge mode or MHP AMVP mode; and / or wherein the sub-block encoding / decoding mode includes at least one of the following: affine Merge mode, affine AMVP mode or sbTMVP.

[0127] Item 9. The method according to Item 1, wherein if the motion vector candidate is encoded or decoded via a target mode, the OBMC parameters of the motion vector candidate for the video unit are not inherited.

[0128] Item 10. The method according to Item 9, wherein the target mode includes at least one of the following: LIC mode, inter-frame AMVP mode, MHP AMVP mode, affine AMVP mode, AMVP-MERGE mode, IBC Merge mode, or IBC AMVP mode.

[0129] Item 11. The method according to Item 1, wherein if the OBMC parameter is not inherited from the motion vector candidate, the OBMC flag is set to the target value.

[0130] Item 12. The method according to Item 11, wherein the target value depends on the LIC flag of the motion vector candidate.

[0131] Item 13. The method according to Item 11, wherein the target value depends on at least one of the following motion vector candidates: sub-block mode, affine mode, or sbTMVP.

[0132] Item 14. The method according to Item 11, wherein the target value depends on the AMVP-MERGE pattern of the motion vector candidate.

[0133] Item 15. The method according to Item 11, wherein the target value depends on at least one of the following of the video unit: block width or block height.

[0134] Item 16. The method according to Item 11, wherein the target value is fixed.

[0135] Item 17. The method according to Item 16, wherein the target value is 0 or 1.

[0136] Item 18. The method according to any one of items 1-17, wherein the OBMC parameters of the MHP-encoded block are set according to a prediction mode based on the fundamental assumption.

[0137] Item 19. The method according to Item 18, wherein the setting of the OBMC parameter depends on whether the basic assumption of the MHP-encoded block is encoded and decoded in sub-block mode.

[0138] Item 20. The method according to Item 19, wherein the sub-block pattern includes at least one of the following: sbTMVP, affine AMVP, or affine Merge.

[0139] Item 21. The method according to Item 18, wherein the OBMC parameter is set to a fixed value if the basic assumption is encoded or decoded in at least one of the following: sbTMVP, affine AMVP, or affine Merge.

[0140] Item 22. The method according to Item 21, wherein the fixed value is 0 or 1.

[0141] Item 23. The method according to Item 18, wherein the setting of the OBMC parameter depends on whether the basic assumption of the MHP-encoded block is encoded or decoded in LIC.

[0142] Item 24. The method according to Item 18, wherein the setting of the OBMC parameter depends on whether the fundamental assumption of the MHP-encoded block is encoded and decoded in AMVP-MERGE mode.

[0143] Item 25. The method according to any one of items 1-24, wherein if the OBMC parameter is indicated for the video unit, the OBMC parameter is context-coded.

[0144] Item 26. The method according to Item 25, wherein the OBMC parameters are encoded and decoded by at least two context models.

[0145] Item 27. The method described in Item 25, wherein which context model is used for the OBMC parameters depends on the prediction mode or prediction method of the video unit.

[0146] Item 28. The method described in Item 25, wherein which context model is used is based on whether the video unit is IBC encoded or decoded.

[0147] Item 29. The method described in Item 25, wherein which context model is used is based on whether the video unit is encoded or decoded in LIC mode.

[0148] Item 30. The method described in Item 25, wherein which context model is used is based on whether the video unit is encoded and decoded in AMVP-MERGE mode.

[0149] Item 31. The method described in Item 25, wherein which context model is used is based on whether the video unit is encoded and decoded in sub-block mode.

[0150] Item 32. The method described in Item 25, wherein which context model is used is based on whether the video unit is encoded or decoded in an affine mode.

[0151] Item 33. The method described in Item 25, wherein which context model is used is based on whether the video unit is encoded or decoded using MHP.

[0152] Item 34. The method described in Item 33, wherein which context model is used is based on whether the size of the additional assumption is greater than 0.

[0153] Item 35. The method according to any one of items 1-34, wherein an indication of whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0154] Item 36. The method according to any one of items 1-35, wherein the indication of whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is indicated by one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.

[0155] Item 37. The method according to any one of items 1-35, wherein the indication of whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, strip, slice, sub-picture, or region containing more than one sample point or pixel.

[0156] Item 38. The method according to any one of items 1-35 further comprises: determining, based on the encoding and decoding information of the video unit, whether to determine whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit, wherein the encoding and decoding information includes at least one of the following: block size, color format, single and / or dual tree segmentation, color components, stripe type, or picture type.

[0157] Item 39. The method according to any one of items 1-38, wherein the conversion includes encoding the video unit into the bitstream.

[0158] Item 40. The method according to any one of items 1-38, wherein the conversion includes decoding the video unit from the bitstream.

[0159] Item 41. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-40.

[0160] Item 42. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-40.

[0161] Item 43. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining, based on inheritance from motion vector candidates, whether overlapping sub-block-based motion compensation (OBMC) is applied to video units of the video, and wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; and generating the bitstream based on the determination.

[0162] Item 44. A method for storing a bitstream of video, comprising: determining whether overlapping subblock-based motion compensation (OBMC) is applied to video units of the video based on inheritance from motion vector candidates, wherein the inheritance is based on one of: a prediction mode of the video unit or a prediction mode of the motion vector candidates; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium. Example device

[0163] Figure 40 A block diagram of a computing device 4000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4000 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0164] It should be understood that, Figure 40 The computing device 4000 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0165] like Figure 40 As shown, the computing device 4000 includes a general-purpose computing device 4000. The computing device 4000 may include at least one or more processors or processing units 4010, memory 4020, storage unit 4030, one or more communication units 4040, one or more input devices 4050, and one or more output devices 4060.

[0166] In some embodiments, the computing device 4000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4000 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0167] Processing unit 4010 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 4020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 4000. Processing unit 4010 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0168] Computing device 4000 typically includes various computer storage media. Such media can be any media accessible by computing device 4000, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 4020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 4030 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 4000.

[0169] The computing device 4000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 40Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0170] The communication unit 4040 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 4000 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 4000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0171] Input device 4050 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4060 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 4040, computing device 4000 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 4000 can also communicate with one or more devices that enable a user to interact with computing device 4000, or, if necessary, with any device that enables computing device 4000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).

[0172] In some embodiments, some or all components of computing device 4000 may be deployed in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and work together to achieve the functions described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (e.g., the Internet) using suitable protocols. For example, a cloud computing provider provides applications via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functions described herein from service providers at remote locations. Alternatively, the components and functions described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0173] In embodiments of this disclosure, computing device 4000 may be used to implement video encoding / decoding. Memory 4020 may include one or more video codec modules 4025 having one or more program instructions. These modules are accessible and executable by processing unit 4010 to perform the functions of the various embodiments described herein.

[0174] In an example embodiment of performing video encoding, input device 4050 may receive video data as input 4070 to be encoded. The video data may be processed, for example, by video codec module 4025 to generate an encoded bitstream. The encoded bitstream may be provided as output 4080 via output device 4060.

[0175] In an example embodiment performing video decoding, input device 4050 may receive an encoded bitstream as input 4070. The encoded bitstream may be processed, for example, by video codec module 4025 to generate decoded video data. The decoded video data may be provided as output 4080 via output device 4060.

[0176] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A video processing method, comprising: For the conversion between video units and their bitstreams, it is determined whether Overlapping Sub-Block Motion Compensation (OBMC) is applied to the video unit based on inheritance from motion vector candidates, wherein said inheritance is based on one of: the prediction mode of the video unit or the prediction mode of the motion vector candidates; and The conversion is performed based on the determination.

2. The method of claim 1, wherein if the video unit is encoded or decoded in target mode, the OBMC parameters for the video unit are inherited.

3. The method of claim 2, wherein the target mode includes at least one of the following: Multiple hypothesis prediction (MHP) model, or Sub-block encoding / decoding mode.

4. The method of claim 3, wherein the MHP mode includes at least one of the following: MHP Merge mode or MHP Inter-Frame Advanced Motion Vector Prediction (AMVP) mode.

5. The method of claim 3, wherein the sub-block encoding / decoding mode includes at least one of the following: Affine Merge mode, Affine AMVP mode, or Sub-block-based temporal motion vector prediction (sbTMVP).

6. The method of claim 1, wherein if the motion vector candidate is encoded or decoded via a target mode, the OBMC parameters of the motion vector candidate for the video unit are inherited.

7. The method of claim 6, wherein the target mode comprises at least one of the following: Inter-frame Merge mode Inter-frame AMVP mode, Standard inter-frame merge mode, MHP mode Geometry (GEO) pattern A variant of GEO Inter-frame and intra-frame joint prediction (CIIP) mode, Variants of CIIP Merge mode (MMVD) with motion vector difference between frames. Sub-block encoding / decoding mode, Affine MMVD, Affine Merge mode, Inter-frame template matching (TM) mode, Inter-frame block matching (BM) mode, AMVP-MERGE mode sbTMVP mode, Intra-Block Copy (IBC) Merge mode, IBC AMVP mode, or Local lighting compensation (LIC).

8. The method of claim 7, wherein the MHP mode includes at least one of the following: MHP Merge mode or MHP AMVP mode; and / or The sub-block encoding / decoding mode includes at least one of the following: affine Merge mode, affine AMVP mode, or sbTMVP.

9. The method of claim 1, wherein if the motion vector candidate is encoded or decoded via a target mode, the OBMC parameters of the motion vector candidate for the video unit are not inherited.

10. The method of claim 9, wherein the target mode comprises at least one of the following: LIC mode, Inter-frame AMVP mode, MHP AMVP mode Affine AMVP mode, AMVP-MERGE mode IBC Merge mode, or IBC AMVP mode.

11. The method of claim 1, wherein if the OBMC parameter is not inherited from the motion vector candidate, the OBMC flag is set to the target value.

12. The method of claim 11, wherein the target value depends on the LIC flag of the motion vector candidate.

13. The method of claim 11, wherein the target value depends on at least one of the following motion vector candidates: sub-block mode, affine mode, or sbTMVP.

14. The method of claim 11, wherein the target value depends on the AMVP-MERGE pattern of the motion vector candidate.

15. The method of claim 11, wherein the target value depends on at least one of the following of the video unit: block width or block height.

16. The method of claim 11, wherein the target value is fixed.

17. The method of claim 16, wherein the target value is 0 or 1.

18. The method according to any one of claims 1-17, wherein the OBMC parameters of the MHP-encoded block are set according to a prediction mode based on the fundamental assumption.

19. The method of claim 18, wherein the setting of the OBMC parameter depends on whether the basic assumption of the MHP-encoded block is encoded and decoded in sub-block mode.

20. The method of claim 19, wherein the sub-block pattern includes at least one of the following: sbTMVP, affine AMVP, or affine Merge.

21. The method of claim 18, wherein the OBMC parameter is set to a fixed value if the basic assumption is encoded or decoded in at least one of the following: sbTMVP, affine AMVP, or affine Merge.

22. The method of claim 21, wherein the fixed value is 0 or 1.

23. The method of claim 18, wherein the setting of the OBMC parameter depends on whether the basic assumption of the MHP-encoded block is encoded or decoded in LIC.

24. The method of claim 18, wherein the setting of the OBMC parameter depends on whether the fundamental assumption of the MHP-encoded block is encoded and decoded in AMVP-MERGE mode.

25. The method according to any one of claims 1-24, wherein if the OBMC parameter is indicated for the video unit, the OBMC parameter is context-coded.

26. The method of claim 25, wherein the OBMC parameters are encoded and decoded by at least two context models.

27. The method of claim 25, wherein which context model is used for the OBMC parameters depends on the prediction mode or prediction method of the video unit.

28. The method of claim 25, wherein which context model is used is based on whether the video unit is IBC encoded or decoded.

29. The method of claim 25, wherein which context model is used is based on whether the video unit is encoded or decoded in LIC mode.

30. The method of claim 25, wherein which context model is used is based on whether the video unit is encoded and decoded in AMVP-MERGE mode.

31. The method of claim 25, wherein which context model is used is based on whether the video unit is encoded and decoded in sub-block mode.

32. The method of claim 25, wherein which context model is used is based on whether the video unit is encoded or decoded in an affine pattern.

33. The method of claim 25, wherein which context model is used is based on whether the video unit is encoded or decoded using MHP.

34. The method of claim 33, wherein which context model is used is based on whether the size of the additional assumption is greater than 0.

35. The method according to any one of claims 1-34, wherein an indication of whether to determine whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is indicated in one of the following places: sequence level, Image group level, Image quality, strip level, or Film series level.

36. The method according to any one of claims 1-35, wherein the indication of whether to determine whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

37. The method according to any one of claims 1-35, wherein the indication of whether to determine whether the OBMC is applied to the video unit and / or how to determine whether the OBMC is applied to the video unit is included in one of the following: Predicted Block (PB), Transform block (TB), Code Block (CB), Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Virtual Pipeline Data Unit (VPDU) Code-decode tree unit (CTU), CTU line, strip, piece, Sub-images, or A region containing more than one sample point or pixel.

38. The method according to any one of claims 1-35, further comprising: Based on the encoding and decoding information of the video unit, it is determined whether and / or how to determine whether the OBMC is applied to the video unit, wherein the encoding and decoding information includes at least one of the following: Block size, Color format, Single and / or dual tree partitioning, Color components, Strip type, or Image type.

39. The method according to any one of claims 1-38, wherein the conversion comprises encoding the video unit into the bitstream.

40. The method according to any one of claims 1-38, wherein the conversion comprises decoding the video unit from the bitstream.

41. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-40.

42. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-40.

43. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: Whether Overlapping Subblock Motion Compensation (OBMC) is applied to a video unit of the video is determined based on inheritance from motion vector candidates, and wherein the inheritance is based on one of the following: the prediction mode of the video unit or the prediction mode of the motion vector candidates; as well as The bit stream is generated based on the determination.

44. A method for storing a bitstream of video, comprising: Whether Overlapping Subblock Motion Compensation (OBMC) is applied to a video unit of the video is determined based on inheritance from motion vector candidates, and wherein the inheritance is based on one of the following: the prediction mode of the video unit or the prediction mode of the motion vector candidates; The bit stream is generated based on the determination; as well as The bitstream is stored in a non-transitory computer-readable medium.