Method and device for video processing and medium
By using the sample point value, adjacent sample point value, template cost or motion vector accuracy in the current block to make adaptive judgments based on the motion compensation (OBMC) of the overlapping sub-blocks during the conversion process between the video unit and the bitstream, the problem of improving the encoding and decoding efficiency in the existing video encoding and decoding technology is solved, and higher encoding and decoding gain and efficiency are achieved.
Patent Information
- Application Number
- CN202380085349.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-12
- Filing Date
- 2023-12-06
- Publication Date
- 2025-07-22
AI Technical Summary
The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in block-level adaptability and motion compensation, which is difficult to further improve.
During the conversion between the video unit and the bitstream, it is determined whether to apply motion compensation (OBMC) based on overlapping sub-blocks (OBMC) to achieve block-level adaptive OBMC based on the sample point value, adjacent sample point value, template cost or motion vector accuracy.
It improves the encoding and codec gain of video encoding and codec, improves the encoding and codec efficiency, and improves the quality and efficiency of video processing.
Smart Images

Figure CN120359756A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to block-level adaptive overlapping sub-block based motion compensation (OBMC) in video codec sample value dependencies. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is generally a desire to further improve the encoding / decoding efficiency of video codec technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a video unit of a video and the bitstream of the video unit, determining whether overlapping sub-block based motion compensation (OBMC) is applied to a current block of the video unit based on at least one of the following: the sample values of the samples within the current block, the sample values of the samples adjacent to the current block, the template cost, or the motion vector precision of the current block; and performing the conversion based on the determination. In this way, block-level adaptive OBMC that takes into account block characteristics based on decoded information can bring higher encoding / decoding gain and improve encoding / decoding efficiency.
[0005] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing for a video. The method includes: determining whether overlapping sub-block based motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: the sample values of the samples within the current block, the sample values of the samples adjacent to the current block, the template cost, or the motion vector precision of the current block; and generating the bitstream based on the determination.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining whether overlapping block motion compensation (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: a sample value of a sample within the current block, a sample value of a sample adjacent to the current block, a template cost, or a motion vector precision of the current block; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.
[0009] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the example embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram showing an example video codec system according to some embodiments of the present disclosure;
[0012] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 An intra prediction mode is shown;
[0015] Figure 5 Reference samples for wide-angle intra prediction are shown;
[0016] Figure 6 The problem of discontinuity in the case of a direction exceeding 45° is shown;
[0017] Figure 7A A schematic diagram showing the definition of samples used by PDPC of the diagonal upper right mode applied to diagonal and adjacent angle intra modes;
[0018] Figure 7B A schematic diagram showing the definition of samples used by PDPC of the diagonal lower left mode applied to diagonal and adjacent angle intra modes;
[0019] Figure 7CSchematic diagram showing the definition of samples used by PDPC of the adjacent diagonal upper right mode applied to diagonal and adjacent angle intra modes;
[0020] Figure 7D Schematic diagram showing the definition of samples used by PDPC of the adjacent diagonal lower left mode applied to diagonal and adjacent angle intra modes;
[0021] Figure 8 Shows an example of four reference lines adjacent to the prediction block;
[0022] Figure 9A Schematic diagram showing the process of sub - division depending on the block size;
[0023] Figure 9B Schematic diagram showing the process of sub - division depending on the block size;
[0024] Figure 10 Shows the matrix - weighted intra - prediction process;
[0025] Figure 11 Shows the spatial GPM candidates;
[0026] Figure 12 Shows the GPM template;
[0027] Figure 13 Shows the GPM mixing;
[0028] Figure 14 Shows the positions of the spatial Merge candidates;
[0029] Figure 15 Shows the candidate pairs considered for the redundancy check of the spatial Merge candidates;
[0030] Figure 16 Shows the graph for the motion vector scaling of the temporal Merge candidates;
[0031] Figure 17 Shows the candidate positions for the temporal Merge candidates C0 and C1;
[0032] Figure 18 Shows the MMVD search points;
[0033] Figure 19 Shows the extended CU region used in BDOF;
[0034] Figure 20 Shows the graph for the symmetric MVD mode;
[0035] Figure 21 Shows the motion vector refinement on the decoding side;
[0036] Figure 22 Shows the top and left neighboring blocks used in CIIP weight derivation;
[0037] Figure 23 Shows an example of GPM partitioning grouped at the same angle;
[0038] Figure 24 Shows the unidirectional prediction MV selection for geometric partitioning mode;
[0039] Figure 25 Shows the generation of hybrid weight w0 using geometric partitioning mode;
[0040] Figure 26 Shows the current CTU processing order and its available reference sample points in the current CTU and the left CTU;
[0041] Figure 27 Shows the residual encoding / decoding passes for transform skip blocks;
[0042] Figure 28 Shows an example of a block encoded / decoded in palette mode;
[0043] Figure 29 Shows the sub-block based index map scan for the palette, with horizontal scan on the left and vertical scan on the right;
[0044] Figure 30 Shows the decoding flowchart with ACT;
[0045] Figure 31 Shows the in-frame template matching search area used;
[0046] Figure 32 Shows five positions in the reconstructed luma samples;
[0047] Figure 33 Shows the prediction process of DBV mode;
[0048] Figure 34 Shows the low-frequency non-separable transform (LFNST) process;
[0049] Figure 35 Shows the SBT position, type, and transform type;
[0050] Figure 36 Shows the ROI for LFNST16;
[0051] Figure 37 Shows the ROI for LFNST8;
[0052] Figure 38 Shows the discontinuity measurement;
[0053] Figure 39 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; and
[0054] Figure 40 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.
[0055] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0056] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.
[0057] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.
[0058] References to "an embodiment", "embodiment", "example embodiment", etc. in the present disclosure indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0059] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0060] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example environment
[0061] Figure 1 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0062] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0063] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 over a network 130A. The encoded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0064] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from a source device 110 or a storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0065] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0066] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0067] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0068] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0069] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0070] In addition, although some components such as the motion estimation unit 204 and the motion compensation unit 205 may be integrated, for purposes of explanation, these components are shown Figure 2is shown separately in the example of
[0071] The splitting unit 201 can split the picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0072] The mode selection unit 203 can select, for example, one of a plurality of coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select an intra-inter joint prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0073] To perform inter prediction on the current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the cache 213 other than the picture associated with the current video block.
[0074] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" can refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" can refer to portions of a picture composed of macroblocks that are independent of the macroblocks in the same picture.
[0075] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0076] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, where the plurality of reference indices indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0077] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0078] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, where the value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0079] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0080] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0081] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0082] The residual generation unit 207 can generate residual data for a current video block by subtracting (e.g., indicated by a minus sign) a (plurality of) predicted video blocks of the current video block from the current video block. The residual data of the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0083] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0084] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0085] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0086] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0087] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce block effect artifacts in the video block.
[0088] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0089] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0090] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3In an example, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0091] In Figure 3 an example, video decoder 300 includes entropy decoding unit 301, motion compensation unit 302, intra prediction unit 303, inverse quantization unit 304, inverse transform unit 305, and reconstruction unit 306 and buffer 307. In some examples, video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to video encoder 200.
[0092] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and in the case of a prediction region in a B slice, also includes an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0093] Motion compensation unit 302 can generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0094] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during the encoding of a video block to calculate the interpolated values for sub-integer pixels of a reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 according to the received syntax information, and motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0095] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.
[0096] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0097] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter may also be applied to filter the decoded block in order to remove blocking artifact. The decoded video block is then stored in the buffer 307, and the buffer 307 provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also produces the decoded video for presentation on a display device.
[0098] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to the multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bitrates. 1. Brief Overview The present disclosure relates to video coding and decoding techniques. Specifically, the present disclosure relates to overlapping block motion compensation (OBMC) and related techniques in picture / video coding and decoding. The present disclosure can be applied to existing video coding and decoding standards such as HEVC, VVC, etc. The present disclosure can also be applicable to future video coding and decoding standards or video codecs. 2. Introduction Video coding standards have mainly evolved from the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, and ISO / IEC developed MPEG-1 and MPEG-4 Visual. The two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly simultaneously. The new video coding standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. Subsequently, the VVC working draft and the test model VTM were updated after each meeting. The VVC project achieved Final Draft International Standard (FDIS) at the meeting in July 2020. 2.1 Existing Coding Tools 2.1.1. Intra Prediction 2.1.1.1. Intra Mode Coding with 67 Intra Prediction Modes To capture any edge direction presented in natural videos, the number of directional intra modes in VVC is extended from 33 used in HEVC to 65. The new directional modes not in HEVC are depicted as red dashed arrows in Figure 4 and the planar mode and DC mode remain the same. These denser directional intra prediction modes are applicable to all block sizes and both luminance intra prediction and chrominance intra prediction. In VVC, several conventional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes for non-square blocks. In HEVC, each intra-coded block has a square shape and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using the DC mode. In VVC, blocks can have a rectangular shape, which generally requires division operations for each block. To avoid division operations for DC prediction, only the longer side is used to calculate the average value of non-square blocks. 2.1.1.2. Intra Mode Coding To keep the complexity of the Most Probable Mode (MPM) list generation low, an intra mode coding method with 6 MPMs is used by considering two available neighboring intra modes. The MPM list is constructed considering the following three aspects: – Default Intra mode; – Neighboring Intra mode; – Derived Intra mode. A unified 6-MPM list is used for Intra blocks regardless of whether MRL and ISP codec tools are applied. The MPM list is constructed based on the Intra modes of the left and upper neighboring blocks. Assuming the mode of the left is denoted as Left and the mode of the upper block is denoted as Above, the unified MPM list is constructed as follows: – When neighboring blocks are not available, their Intra modes are defaultly set to Planar. – If both modes Left and Above are non-angular modes: - MPM list → {Planar, DC, V, H, V-4, V+4}. – If one of the modes Left and Above is an angular mode and the other is a non-angular mode: - Set the mode Max to the larger mode among Left and Above - MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2}. – If both Left and Above are angular modes and they are different: - Set the mode Max to the larger mode among Left and Above - If the difference between the modes Left and Above is in the range of 2 to 62 (including the boundary values) - MPM list → {Planar, Left, Above, DC, Max-1, Max+1} - Otherwise - MPM list → {Planar, Left, Above, DC, Max-2, Max+2}. – If both Left and Above are angular modes and they are the same: - MPM list → {Planar, Left, Left-1, Left+1, DC, Left-2}. In addition, the first binary bit of the MPM index codeword is context-coded by CABAC. A total of three contexts are used, corresponding to whether the current Intra block is MRL-enabled, ISP-enabled, or a normal Intra block. During the 6-MPM list generation process, deduplication is used to remove duplicate modes so that only unique modes can be included in the MPM list. For the entropy coding of 61 non-MPM modes, Truncated Binary Codes (TBC) are used. 2.1.1.3 Wide-angle Intra prediction for non-square blocks The traditional angular intra prediction directions are defined as 45 degrees to -135 degrees in a clockwise direction. In VVC, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The original mode index is used to signal the replaced mode, and the original mode index is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains the same, i.e., 67, and the intra mode encoding and decoding methods remain unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 5 shown. The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 1. Table 1 Intra prediction modes replaced by wide-angle modes As Figure 6 shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the negative impact brought by the increased gap Δp α If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When predicting a block through these modes, the samples in the reference buffer are directly copied without applying any interpolation. By this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fractional modes in traditional prediction modes and wide-angle modes. In VVC, 4:2:2 and 4:4:4 chrominance formats as well as 4:2:0 chrominance format are supported. The chrominance derivation mode (DM) derivation table for the 4:2:2 chrominance format was initially ported from HEVC, and the number of entries was extended from 35 to 67 to align with the extension of the intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes in the range from 2 to 5 are mapped to 2. Therefore, the chrominance DM derivation table for the 4:2:2 chrominance format is updated by replacing some values of the entries in the mapping table to more accurately convert the prediction angles of chrominance blocks. 2.1.1.4. Mode-Dependent Intra Smoothing (MDIS) The four-tap frame interpolation value filter is used to improve the accuracy of intra prediction with direction. In HEVC, a two-tap linear interpolation filter has been used to generate an intra prediction block in the prediction mode with direction (i.e., excluding the planar and DC prediction values). In VVC, a simplified 6-bit four-tap Gaussian interpolation filter is only used for the intra mode with direction. The intra prediction process without direction is not modified. The selection of the four-tap filter is performed according to the MDIS condition for the intra prediction mode with direction, and the intra prediction mode with direction provides a non-fractional displacement, that is, all direction modes below are excluded: 2, HOR_IDX, DIA_IDX, VER_IDX, 66. According to the intra prediction mode, the following reference sample processing is performed: – The intra prediction mode with direction is classified into one of the following groups: - Vertical or horizontal mode (HOR_IDX, VER_IDX), - Diagonal mode (2, DIA_IDX, VDIA_IDX) representing an angle that is a multiple of 45 degrees, - The remaining modes with aspects; – If the intra prediction mode with direction is classified as belonging to Group A, then no filter is applied to the reference samples to generate the predicted samples; – Otherwise, if the mode falls into Group B, then a [1, 2, 1] reference sample filter may be applied (depending on the MDIS condition) to the reference samples to further copy these filtered values to the intra prediction value according to the selected direction, but no interpolation filter is applied; – Otherwise, if the mode is classified as belonging to Group C, then only the intra reference sample interpolation filter is applied to the reference samples to generate predicted samples at fractional or integer positions between the reference samples according to the selected direction (no reference sample filtering is performed). 2.1.1.5. Position-Dependent Intra Prediction Combination In VVC, the result of intra prediction for DC, planar, and several angular modes is further modified by the position-dependent intra prediction combination (PDPC) method. PDPC is an intra prediction method that calls a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, bottom-left corner mode and its eight adjacent angular modes, and top-right angular mode and its eight adjacent angular modes. According to Equation 3-8, a linear combination of the intra prediction mode (DC, planar, angular) and reference samples is used to predict the predicted sample pred(x’, y’), as follows: pred(x’,y’) = (wL × R -1,y ’ + wT × Rx ’ ,-1 -wTL × R -1,-1 +(64 - wL - wT + wTL) × pred(x’, y’) + 32) >> 6 (2 - 1) where R x,-1 , R -1,y respectively represent the reference samples located at the top and left boundaries of the current sample (x, y), and R -1,-1 represents the reference sample located at the upper left corner of the current block. If PDPC is applied to DC, planar, horizontal, and vertical intra - modes, then no additional boundary filters are required, as is the case for the HEVC DC - mode boundary filter or the horizontal / vertical - mode edge filter. The PDPC processes for DC and planar modes are the same and avoid the clipping operation. For angular modes, the pdpc scaling factor is adjusted such that no range check is required and the condition on the angle enabling pdpc (scaling >= 0) is removed. Additionally, the PDPC weights are based on 32 in all angular - mode cases. The PDPC weights depend on the prediction mode and are shown in Table 2. PDPC is applied to blocks where both the width and height are greater than or equal to 4. Figures 7A to 7D Shows the reference samples (R x,-1 , R -1,y and R -1,-1 ) defined for applying PDPC over various prediction modes. The predicted sample pred(x’, y’) is located at (x’, y’) within the prediction block. For example, for the diagonal mode, the x - coordinate of the reference sample R x,-1 is given by x = x’ + y’ + 1, and the y - coordinate of the reference sample R -1,y is similarly given by y = x’ + y’ + 1. For another angular mode, the reference samples R x,-1 and R -1,y can be located at fractional - sample positions. In this case, the sample values at the nearest integer - sample positions are used. Table 2 - Examples of PDPC weights according to prediction mode 2.1.1.6. Multiple - reference - line (MRL) intra - prediction Multiple - reference - line (MRL) intra - prediction uses more reference lines for intra - prediction. In Figure 8 , an example of 4 reference lines is depicted, where the samples in segments A and F are not taken from the reconstructed neighboring samples but are filled with the closest samples from segments B and E respectively. HEVC intra - picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 3) are used. The index of the selected reference line (mrl_idx) is signaled and used to generate the intra prediction value. For reference line idx greater than 0, only additional reference line modes are included in the MPM list, and only the MPM index is signaled without the residual mode. The reference line index is signaled before the intra prediction mode, and in the case where a non-zero reference line index is signaled, the planar mode is excluded from the intra prediction modes. MRL is disabled for the first line of blocks within a CTU to prevent the use of extended reference samples outside the current CTU row. Additionally, PDPC will be disabled when additional lines are used. For the MRL mode, the derivation of the DC value in the DC intra prediction mode for a non-zero reference line index is aligned with the derivation for reference line index 0. MRL requires the use of 3 neighboring luma reference lines stored in the CTU to generate the prediction. The cross-component linear model (CCLM) tool also requires 3 neighboring luma reference lines for its downsampling filter. The definition of MRL using the same 3 lines is aligned with CCLM to reduce the storage requirements of the decoder. 2.1.1.7. Intra Sub-Partitioning (ISP) Intra Sub-Partitioning (ISP) divides the luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions according to the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is greater than 4×8 (or 8×4), the corresponding block will be divided into 4 sub-partitions. It has been noted that M×128 (M≤64) and 128×N (N≤64) ISP blocks may pose potential problems when using the 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding chroma TBs. If the CU uses ISP, then the luma TB will be divided into 4 M×32 TBs (only horizontal division is possible), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, the sizes of both chroma components will be greater than 32×32 blocks. Similarly, using ISP with a 128×N CU can create a similar situation. Therefore, these two cases are problems for the 64×64 decoder pipeline. Therefore, the CU size for which ISP can be used is limited to a maximum of 64×64. Figure 9A and Figure 9B Examples showing two possibilities are presented. All sub-partitions meet the condition of having at least 16 samples. In the ISP, it is not allowed that the 1×N / 2×N sub-block prediction depends on the reconstructed values of the previously decoded 1×N / 2×N sub-blocks of the coding / decoding block, such that the minimum prediction width of the sub-block becomes four samples. For example, an 8×N (N>4) coding / decoding block using ISP coding / decoding with vertical partitioning is divided into two prediction regions, each with a size of 4×N, and four transforms have a size of 2×N. Similarly, a 4×N coding / decoding block using ISP coding / decoding with vertical partitioning is predicted using the complete 4×N block; four transforms are used, each being 1×N. Although transform sizes of 1×N and 2×N are allowed, it can be asserted that the transforms of these blocks within the 4×N region can be executed in parallel. For example, when the 4×N prediction region contains four 1×N transforms, there is no transform in the horizontal direction; the transforms in the vertical direction can be executed as a single 4×N transform in the vertical direction. Similarly, when the 4×N prediction region contains two 2×N transform blocks, the transform operations of the two 2×N blocks can be performed in parallel in each direction (horizontal and vertical). Therefore, compared with processing the intra-block of 4×4 conventional coding / decoding, no additional latency is added when processing these smaller blocks. Table 3 - Entropy Coding / Decoding Coefficient Group Size Block size Coefficient group size 1×N, N≥16 1×16 N×1, N≥16 16×1 2×N, N≥8 2×8 N×2, N≥8 8×2 All other possible M×N cases 4×4 For each sub-partition, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated through processes such as entropy decoding, inverse quantization, and inverse transform. Therefore, the reconstructed sample values of each sub-partition can be used to generate the prediction of the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the sub-partition that contains the top-left sample of the CU and then continues downward (horizontal partitioning) or to the right (vertical partitioning). Therefore, the reference samples for generating the sub-partition prediction signal are only located on the left and above the row. All sub-partitions share the same intra-mode. The following is a summary of the interaction between the ISP and other coding / decoding tools. – Multiple Reference Lines (MRL): If the MRL index of the block is not 0, the ISP coding / decoding mode will be presumed to be 0, so the ISP mode information will not be sent to the decoder. – Entropy Coding / Decoding Coefficient Group Size: As shown in Table 3, the size of the entropy coding / decoding sub-block has been modified so that there are 16 samples in all possible cases. Note that the new size only affects the blocks generated by the ISP, where one size is less than 4 samples. In all other cases, the coefficient group remains at a 4×4 size. – CBF Coding / Decoding: Assume that at least one sub-partition has a non-zero CBF. Therefore, if n is the number of sub-partitions and the first n - 1 sub-partitions have produced zero CBF, the CBF of the nth sub-partition is presumed to be 1. – MPM Usage: The MPM flag will be presumed to be the flag in the blocks encoded and decoded by the ISP mode, and the MPM list is modified to exclude the DC mode and prioritize the vertical intra mode for ISP horizontal partitions and prioritize the vertical intra mode for vertical partitions. – Transformed Size Constraint: All ISP transforms with a length greater than 16 points use DCT-II. – PDPC: When the CU uses the ISP encoding / decoding mode, the PDPC filter will not be applied to the resulting sub-partitions. – MTS Flag: If the CU uses the ISP encoding / decoding mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests on the different available transforms for each resulting sub-partition. The transform selection for the ISP mode will be changed to be fixed and selected according to the intra mode used, the processing order, and the block size. Therefore, signaling is not required. For example, let t H and t V be the horizontal transform and the vertical transform selected for the w×h sub-partition respectively, where w is the width and h is the height. Then the transform is selected according to the following rules: – If w = 1 or h = 1, there is no horizontal transform or vertical transform respectively. – If w = 2 or w > 32, t H = DCT-II – If h = 2 or h > 32, t V = DCT-II – Otherwise, the transform is selected as shown in Table 4. Table 4 - Transform Selection Depends on Intra Mode In the ISP mode, all 67 intra modes are allowed. If the corresponding width and height are at least 4 samples long, PDPC is also applied. In addition, the conditions for selecting the intra interpolation filter no longer exist, and in the ISP mode, the cubic (DCT-IF) filter is always applied to the fractional position interpolation. 2.1.1.8 Matrix-Weighted Intra Prediction (MIP) The matrix-weighted intra prediction (MIP) method is a newly added intra prediction technique in VVC. To predict the samples of a rectangular block with width W and height H, the matrix-weighted intra prediction (MIP) takes one row of H reconstructed neighboring boundary samples on the left side of the block and one row of W reconstructed neighboring boundary samples above the block as inputs. If the reconstructed samples are not available, they are generated as in traditional intra prediction. As Figure 10 shown, the generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation. ● Average Neighboring Samples Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, according to a predefined rule depending on the block size, by averaging the neighboring boundary samples, the input boundary bdry top and bdry left are reduced to a smaller boundary and Then, the two reduced boundaries and are spliced into the reduced boundary vector bdry red , which has a size of 4 for a 4×4 shaped block and 8 for all other shaped blocks. If the mode refers to the MIP mode, this splicing is defined as follows: ● Matrix multiplication Taking the averaged samples as input, perform matrix-vector multiplication and then add an offset. The result is a reduced prediction signal on a subsampled set of samples in the original block. From the reduced input vector bdry red , a reduced prediction signal pred red is generated, which is a signal on a downsampled block with width W red and height H red . Here, W red and H red are defined as: The reduced prediction signal pred red is calculated by computing the matrix-vector product and adding an offset: pred red = A·bdry red + b Here, A is a matrix that has W red ·H red rows and has 4 columns if W = H = 4 and 8 columns in all other cases. b is a vector with size W red ·H red . The matrix A and the offset vector b are taken from one of the sets S0, S1, S2. One defines the index idx = idx(W,H) as follows: Here, each coefficient of the matrix A is represented with 8-bit precision. The set S0 consists of 16 matrices each having 16 rows and 4 columns, and 16 offset vectors The size of each offset vector is 16. The matrices and offset vectors of the set are for blocks of size 4×4. The set S1 consists of 8 matrices Each matrix has 16 rows and 8 columns, and 8 offset vectors Each offset vector has a size of 16. The set S2 consists of 6 matrices Each matrix has 64 rows and 8 columns, and 6 offset vectors Each offset vector has a size of 64. ● Interpolation The predicted signal at the remaining positions is generated by linear interpolation from the predicted signals on the subsampled set, which is a single-step linear interpolation in each direction. The interpolation is first performed in the horizontal direction and then in the vertical direction, independent of the shape or size of the block. ● Signaling of the MIP mode and coordination with other coding and decoding tools For each coding and decoding unit (CU) in the intra mode, a flag indicating whether to apply the MIP mode is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transpose flag (isTransposed) (which determines whether the mode is transposed) and the MIP mode Id (modeId) (which determines which matrix to be used for a given MIP mode) are derived as follows. isTransposed = predModeIntra & 1 modeId = predModeIntra >> 1 (2-6) The MIP coding and decoding mode is coordinated with other coding and decoding tools by considering the following aspects: – Enable LFNST for MIP on large blocks. The LFNST transform of the planar mode is used here. – The derivation of reference samples for MIP is performed entirely as the derivation of reference samples for traditional intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled samples. – Clipping is performed before upsampling instead of after upsampling. – MIP allows up to 64×64 regardless of the maximum transform size. – The number of MIP modes is 32 for sizeId = 0, 16 for sizeId = 1, and 12 for sizeId = 2. 2.1.1.9. Spatial GPM (SGPM) In the spatial GPM, a candidate list including split partitions and two intra prediction modes is constructed. Up to 11 MPMs of the intra prediction modes are used to form combinations, and the length of the candidate list is set to be equal to 16. The selected candidate index is transmitted by signal. Use Figure 11 The list is reordered using the template shown. The GPM mixing process is not used in the template, and the SAD between the prediction and reconstruction of the template is used for sorting. Figure 12 The GPM template is shown. Figure 13 The GPM split boundary is shown. The SGPM mode is applied to blocks whose width and height both satisfy the same constraints as in the inter-frame GPM. Consider the following items: ● Spatial GPM split mode: 26 predefined modes Adaptive derivation algorithm based on horizontal and vertical gradient ratios ● Intra prediction mode selection: IPM list with and without TIMD: For each split mode, the IPM list is derived for each part using the intra-inter GPM list derivation. The IPM list size is 3. In the list, the TIMD derivation mode is replaced by 2 derivation modes with horizontal and vertical orientations (using the top or left template) or the TIMD derivation mode is excluded. MPM list: A unified MPM list (up to 11 elements) is used for all split modes. ● Template size (left and above): 1 or 4 ● Extended block size: The extended spatial GPM is further applied to 4x8, 8x4, 4x16, and 16x4 blocks, which can be described as 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32. ● Adaptive mixing: Adaptive mixing is tested for the spatial GPM, where the mixing depth τ is derived as follows: ■ If min(width, height) == 4, then select 1 / 2τ. ■ Otherwise, if min(width, height) == 8, then select τ. ■ Otherwise, if min(width, height) == 16, then select 2τ. ■ Otherwise, if min(width, height) == 32, then select 4τ. ■ Otherwise, select 8τ. 2.1.2. Inter-frame prediction For each inter-predicted CU, motion parameters consisting of a motion vector, a reference picture index, and reference picture list use indices, as well as additional information required by the new decoding features of VVC, are used for inter-predicted sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded motion vector difference, or reference picture index. The Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in VVC. The Merge mode can be applied to any inter-predicted CU, not only for skip mode. The alternative to the Merge mode is the explicit transmission of motion parameters, where, for each reference picture list and motion vector, the corresponding reference picture index, along with flags and other required information, is explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC includes multiple new and refined inter-prediction coding tools listed as follows: – Extended Merge prediction – Merge mode with MVD (MMVD) – Symmetric MVD (SMVD) signaling – Affine motion compensation prediction – Sub-block based temporal motion vector prediction (SbTMVP) – Adaptive motion vector resolution (AMVR) – Motion field storage: 1 / 16 luminance sample MV storage and 8x8 motion field compression – Bi-directional prediction with CU-level weight (BCW) – Bi-directional optical flow (BDOF) – Decoder-side motion vector refinement (DMVR) – Geometric partition mode (GPM) – Intra-inter joint prediction (CIIP). The following text provides details on those inter-prediction methods specified in VVC. 2.1.2.1. Extended Merge prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVPs from spatially neighboring CUs. 2) Temporal MVPs from co-located CUs. 3) History-based MVPs from the FIFO table. 4) Pairwise-averaged MVPs. 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU code in the Merge mode, the index of the best Merge candidate is encoded using truncating unary binary (TU). The first binary digit (bin) of the Merge index is coded using context, while bypass coding is used for the other binary digits. The derivation process of the Merge candidates for each category is provided in this section. Similar to what is done in HEVC, VVC also supports parallel derivation of the Merge candidate lists for all CUs within a certain size region. 2.1.2.1.1. Spatial candidate derivation The derivation of the spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the Figure 14 shown positions, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are not available (e.g., because it belongs to another strip or slice) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs Figure 15 linked by arrows in are considered, and the candidate is added to the list only if the corresponding candidates used for the redundancy check do not have the same motion information. 2.1.2.1.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal Merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for deriving the co-located CU is signaled explicitly in the slice header. As shown by the dashed line in Figure 16 , the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. As shown in Figure 17As shown, the position of the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is unavailable, intra-coded / decoded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate. 2.1.2.1.3. History-based Merge candidate derivation After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of previously coded / decoded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a non-sub-block inter-coded / decoded CU, the associated motion information is added as a new HMVP candidate to the last entry of the table. The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is the same HMVP in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are shifted forward. HMVP candidates can be used in the Merge candidate list construction process. The last few HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. For spatial or temporal Merge candidates, a redundancy check is applied to the HMVP candidates. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches one less than the maximum allowed Merge candidates, the Merge candidate list construction process from HMVP is terminated. 2.1.2.1.4. Paired-average Merge candidate derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices of the Merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are both available in a list, they are averaged even if the two motion vectors point to different reference pictures; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, this list is kept invalid. When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached. 2.1.2.2. Merge Estimation Region The Merge Estimation Region (MER) allows for independent derivation of the Merge candidate list for a CU within the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for generating the Merge candidate list of the current CU. Additionally, the update process for the history-based motion vector prediction value candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2parMrglevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture, and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set in the form of log2_parallel_merge_level_minus2. 2.1.2.3. Merge Mode with MVD (MMVD) In addition to the Merge mode where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge Mode with Motion Vector Difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the Merge flag to specify whether MMVD is used for the CU. In MMVD, after selecting Merge candidates, they are further refined by MVD information transmitted via signals. The further information includes a Merge candidate flag, an index for specifying the motion amplitude, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The Merge candidate flag is transmitted via signals to specify which one to use. The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As Figure 18 shown, the offset is added to either the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 5. Table 5 - Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 6. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a non-predicted MV or a bi-predicted MV where both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 6 specify the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV with two MVs to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), the symbols in Table 6 specify the sign of the MV offset added to the list 0 MV component of the starting MV and the sign for the list 1 MV has an opposite value. Table 6 - Signs of MV offsets specified by the direction index Direction IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.4. Bi-directional prediction with CU-level weighting (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P0 + w * P1 + 4) >> 3 (2 - 7) In weighted average bi-prediction, five weights are allowed, w ∈ {-2, 3, 4, 5, 10}. For each bi-predicted CU, the weight w is determined in one of two ways: 1) for non-Merge CUs, the weight index is signaled after the motion vector difference; 2) for Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights will be used. For non-low-delay pictures, only 3 weights are used (w ∈ {3, 4, 5}). – At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-delay picture, unequal weights are only conditionally checked for 1-pixel and 4-pixel motion vector precisions. – When combined with affine, affine ME for unequal weights is performed if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked. – Unequal weights are not searched when certain conditions are met, depending on the POC distance between the current picture and its reference pictures, the coding / decoding QP, and the temporal level. The BCW weight index is decoded using one context-coded bit followed by bypass-coded bits. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, additional bits are signaled using bypass coding to indicate which unequal weight is used. Weighted Prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficient coding and decoding of video content in fading situations. Support for WP has also been added in the VVC standard. WP allows signaling of weighted parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets for the corresponding reference picture are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded and decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.2.5. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (previously known as BIO) was included in JEM. Compared to the JEM version, BDOF in VVC is a simpler version and requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met: – The CU is coded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures precedes the current picture in the display order and the other of the two reference pictures follows the current picture in the display order – The distances of the two reference pictures to the current picture (i.e., POC differences) are the same – Both reference pictures are short-term reference pictures – The CU is not coded and decoded using the affine mode or the ATMVP Merge mode – The CU has more than 64 luma samples – Both the CU height and the CU width are greater than or equal to 8 luma samples – The BCW weight index indicates equal weights – WP is not enabled for the current CU – The CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted sample and the L1 predicted sample. Then the motion refinement is used to adjust the dual-predicted sample values in the 4x4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the difference between two neighboring samples, and i.e., where I (k) (i, j) is the sample value at the prediction signal coordinates (i, j) in the list k, k = 0, 1, and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luminance bit depth bitDepth. Then, the autocorrelations and cross-correlations of the gradients S1, S2, S3, S5, and S6 are calculated as follows S5 = ∑ (i,j)∈Ω Abs(ψ y ), S6 = ∑ (i,j)∈Ω θ(i, j)·Sign(ψ y ) where where Ω is a 6×6 window around the 4×4 sub-block, and n a and n b are set to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, using the cross-correlation terms and autocorrelation terms, the motion refinement (v x , v y ) is derived using the following equation: where th′ BIO = 2 max(5,BD-7) . is the floor function, and Based on the motion refinement and the gradients, the following adjustment is calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bi - directional prediction samples in the following manner: pred BDOF (x,y) = (I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset ) >> shift (2 - 13) These values are selected to keep the multiplier in the BDOF process no more than 15 bits and the maximum bit - width of the intermediate parameters in the BDOF process within 32 bits. To derive the gradient values, some prediction samples I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 19 depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples (white positions) in the extended region are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8 - tap motion - compensated interpolation filter is used to generate the prediction samples (gray positions) inside the CU. These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma samples, it is divided into sub - blocks with width and / or height equal to 16 luma samples, and the sub - block boundaries are regarded as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub - block, the BDOF process can be skipped. When the SAD between the initial L0 prediction sample and the L1 prediction sample is less than the threshold, the BDOF process is not applied to the sub - block. The threshold is set to be equal to (8*W*(H>>1), where W represents the sub - block width and H represents the sub - block height. To avoid the additional complexity of SAD calculation, the SAD calculated in the DVMR process between the initial L0 prediction sample and the L1 prediction sample is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bi - directional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag of any one of the two reference pictures is 1, the BDOF is also disabled. When the CU is encoded / decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.2.6. Symmetric MVD Encoding and Decoding In VVC, in addition to the conventional unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, the symmetric MVD mode is applied to the bidirectional prediction MVD signaling. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not signaled but derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. – Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise BiDirPredFlag is set to 0. 2) At the CU level, if the CU is encoded / decoded by bidirectional prediction and BiDirPredFlag is equal to 1, then the symmetric mode flag indicating whether the symmetric mode is used is signaled explicitly. Figure 20 It is a figure for the symmetric MVD mode. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indexes for list 0 and list 1 are respectively set to be equal to the reference picture pair. MVD1 is set to be equal to (-MVD0). The final motion vector is as shown in the following formula. In the encoder, the symmetric MVD motion estimation starts from the initial MV evaluation. The initial MV candidate set includes the MVs obtained from the unidirectional prediction search, the MVs obtained from the bidirectional prediction search, and the MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.2.7. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the MVs in the Merge mode, the decoder-side motion vector refinement based on bilateral matching is applied in VVC. In the bidirectional prediction operation, the refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and list L1. As Figure 21As shown, the SAD between the red blocks for each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi - directional prediction signal. In VVC, DMVR can be applied to CUs encoded and decoded using the following patterns and features: – CU - level Merge mode with bi - directional prediction MVs – For the current picture, one reference picture is past and the other reference picture is future – The distances from the two reference pictures to the current picture (i.e., POC differences) are the same – Both reference pictures are short - term reference pictures – The CU has more than 64 luma samples – Both the CU height and CU width are greater than or equal to 8 luma samples – The BCW weight index indicates equal weights – WP is not enabled for the current block – The CIIP mode is not used for the current block. The refined MV derived through the DMVR process is used to generate inter - prediction samples and is also used for temporal motion vector prediction in future picture coding. While the original MV is used for the de - blocking process and is also used for spatial motion vector prediction in future CU coding. Additional features of DMVR are mentioned in the following sub - articles. 2.1.2.7.1. Search Scheme In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0+MV_offset (2 - 15) MV1′ = MV1 - MV_offset (2 - 16) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. Integer sample offset search adopts a 25-point full search. First, calculate the SAD of the initial MV pair. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. To reduce the influence of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidate reference reduces the SAD value by 1 / 4. After the integer sample search, there is fractional sample refinement. To save computational complexity, fractional sample refinement is derived by using the parametric error surface equation instead of an additional search using SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. When the integer sample search stage ends at the center with the minimum SAD in the first iteration or the second iteration search, fractional sample refinement is further applied. In sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C(2 - 17) where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (2 - 18) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) (2 - 19) The values of x min and y min are automatically restricted between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined differential MV. 2.1.2.7.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are centered around the initial fractional-pixel MV with integer-sample offsets, so for the DMVR search process, samples at these fractional positions need to be interpolated. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, within a 2-sample search range, compared with the normal motion compensation process, DVMR does not access more reference samples. After obtaining the refined MV using the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples in the normal MC process, samples will be filled from those available samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV. 2.1.2.7.3 Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it is further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.1.2.8. Intra-Inter Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the intra-inter joint prediction (CIIP) mode is applied to the current CU. Figure 22 The top and left neighboring blocks used in CIIP weight derivation are shown. As the name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived after the regular intra prediction process with planar mode. Then, a weighted average is used to combine the intra and inter prediction signals, where the weight values are calculated according to the encoding and decoding modes of the top and left neighboring blocks as follows: – If the top neighbor is available and intra-coded, then set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and intra-coded, then set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) equals 2, set wt to 3; – Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2; – Otherwise, set wt to 1. CIIP prediction is formed as follows: P CIIP = ((4 - wt)*P inter + wt*P intra + 2) >> 2 (2 - 20) 2.1.2.9. Multiple Hypothesis Prediction (MHP) On top of the inter - frame AMVP mode, the regular Merge mode, and the MMVD mode, up to two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated using each additional prediction signal. p n+1 = (1 - α n+1 )p n + α n+1 h n+1 The weight factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For the inter - frame AMVP mode, MHP is applied only when unequal weights in BCW are selected in the bi - directional prediction mode. 2.1.2.10. Overlapped - Block - based Motion Compensation (OBMC) When OBMC is applied, the motion information of neighboring blocks with weighted prediction is used to refine the top and left boundary pixels of the CU. The conditions for not applying OBMC are as follows: ● When OBMC is disabled at the SPS level ● When the current block has an intra - frame mode or an IBC mode ● When the current block applies LIC ● When the current luma block area is less than or equal to 32. Sub - block boundary OBMC is performed by applying the same blend to the top, left, bottom, and right sub - block boundary pixels using the motion information of neighboring sub - blocks. This enables the following sub - block - based coding / decoding tools: ● Affine AMVP mode; ● Affine Merge mode and Sub - block - based Temporal Motion Vector Prediction (SbTMVP); ● Sub - block - based bilateral matching. 2.1.2.11. Local Illumination Compensation (LIC) LIC is an inter - frame prediction technique that models the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template and the reference block template. The parameters of this function can be represented by a scale α and an offset β, which form a linear equation, i.e., α*p[x]+β, to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When looped motion compensation is enabled, the MV will be clipped, taking into account the loop offset. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for the AMVP mode to indicate the use of LIC. The local illumination compensation proposed in JVET - O0066 is used for unidirectional - predicted inter - frame CUs with the following modifications. ● Intra - frame neighboring samples can be used for LIC parameter derivation; ● LIC is disabled for blocks with fewer than 32 luma samples; ● For both non - sub - block and affine modes, LIC parameter derivation is performed based on the modulo - block samples corresponding to the current CU rather than the partial modulo - block samples corresponding to the top - left 16x16 unit. ● The samples of the reference block template are generated by using MC with block MVs without rounding them to integer - pixel precision. 2.1.2.12. Geometric Partitioning Mode (GPM) In VVC, the geometric partitioning mode is supported for inter - frame prediction. A CU - level flag is used as a kind of Merge mode to signal the geometric partitioning mode, and other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub - block Merge mode. For each possible CU size w×h = 2 m ×2 n , where m,n∈{3…6} excluding 8x64 and 64x8, the geometric partitioning mode supports a total of 64 partitions. When using this mode, the CU is divided into two parts by geometrically - located lines ( Figure 23 ). The position of the dividing line is mathematically derived from the angle and offset parameters of a specific partition. Each part in the geometric partition of the CU is inter - frame predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has a motion vector and a reference index. Unidirectional - prediction motion constraints are applied to ensure the same as traditional bi - directional prediction, and each CU only requires two motion - compensated predictions. If the geometric partitioning mode is used for the current CU, the geometric partitioning index (angle and offset) indicating the partitioning mode of the geometric partitioning and two Merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is signaled explicitly in the SPS, and the syntax binarization for the GPM Merge indices is specified. After predicting each part of the geometric partitioning, hybrid processing with adaptive weights is used to adjust the sample values along the geometric partitioning edge. This is the prediction signal for the entire CU, and the transform process and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.2.12.1 Unidirectional Prediction Candidate List Construction The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (X equals the parity of n) of the nth extended Merge candidate is used as the nth unidirectional prediction motion vector for the geometric partitioning mode. These motion vectors are marked with "x" in Figure 24 . If the corresponding LX motion vector of the nth extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partitioning mode. 2.1.2.12.2. Hybrid Along Geometric Partitioning Edge Figure 25 An example generation of the hybrid weight w0 using the geometric partitioning mode is shown. After predicting each part of the geometric partitioning using its own motion, the hybrid is applied to the two prediction signals to derive the samples around the geometric partitioning edge. The hybrid weight for each position of the CU is derived based on the distance between the independent position and the partitioning edge. The distance from the position (x,y) to the partitioning edge is derived as: where i,j are the indices of the angle and offset of the geometric partitioning, which depend on the signaled geometric partitioning index. ρ x,j and ρ y,j The signs of depend on the angle index i. The weights for each part of the geometric partitioning are derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) (2-25) w1(x,y) = 1 - w0(x,y) (2-27) The partIdx depends on the angular index i. An example of the weight w0 is shown below. 2.1.2.12.3. Motion field storage for geometric partitioning mode Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded in the geometric partitioning mode. The type of motion vector stored for each independent position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 - partIdx) : partIdx)(2 - 28) where motionIdx is equal to d(4x + 2, 4y + 2). The partIdx depends on the angular index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field. Otherwise, if sType is equal to 2, then the combined Mv from Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi - directional predicted motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, then only the unidirectional predicted motion Mv2 is stored. 2.1.2.12.4. GPM with inter - frame and intra - frame prediction (GPM inter - frame - intra - frame) With GPM inter - frame - intra - frame, in addition to the Merge candidates for each non - rectangular partition region in the CU for GPM applications, predefined intra - frame prediction modes for the geometric partitioning lines can also be selected. In the proposed method, the intra - frame or inter - frame prediction mode is determined for each GPM separation region with a flag from the encoder. When the inter - frame prediction mode, a unidirectional prediction signal is generated from the MVs in the Merge candidate list. On the other hand, when the intra - frame prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra - frame prediction mode specified by an index from the encoder. The variation of possible intra - frame prediction modes is restricted by the geometry. Finally, the two unidirectional prediction signals are blended in the same way as in ordinary GPM. 2.1.3 Screen content encoding and decoding tools 2.1.3.1. Intra - block copy (IBC) Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. As is well known, it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed within the current picture. The luminance block vectors of the CUs encoded and decoded by IBC have integer precision. The chrominance block vectors are also rounded to integer precision. When used in combination with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precisions. The CUs encoded and decoded by IBC are regarded as a third prediction mode in addition to the intra or inter prediction modes. The IBC mode is applicable to CUs with a width and height less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height not greater than 16 luminance samples. For non-Merge modes, first, a hash-based search is used to perform the block vector search. If the hash search does not return a valid candidate, a block-matching based local search will be performed. In the hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4x4 sub-blocks. For a current block of a larger size, when all the hash keys of all 4×4 sub-blocks match the hash keys at the corresponding reference positions, it is determined that the hash key matches the hash key of the reference block. If it is found that the hash keys of multiple reference blocks match the hash key of the current block, the block vector costs of each matching reference are calculated, and the one with the minimum cost is selected. In the block-matching search, the search range is set to cover the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using flags, and it can be signaled as the IBC AMVP mode or the IBC Skip / Merge mode, as follows: - IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC-encoded blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and paired candidates. - IBC AMVP mode: The block vector difference is encoded in the same way as the motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if it is IBC-encoded). When either neighbor is not available, the default block vector is used as the prediction value. A flag is signaled to indicate the block vector prediction value index. 2.1.3.1.1. IBC Reference Region To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstructed parts of predefined regions, which include the region of the current CTU and some regions of the left CTU. Figure 26 The reference regions for the IBC mode are shown, where each block represents a 64x64 luma sample unit. Depending on the position of the current coding / decoding CU position within the current CTU, the following applies: – If the current block falls within the upper left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, the CPR mode can be used to reference the reference samples in the lower right 64x64 block of the left CTU. The current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU. – If the current block falls within the upper right 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (0,64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also reference the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls within the lower left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to reference the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU. Otherwise, the current block can also use the CPR mode to reference the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls within the lower right 64x64 block of the current CTU, the current block can use the CPR mode to only reference the samples already reconstructed in the current CTU. This constraint allows the implementation of the IBC mode using local on-chip memory for hardware implementation. 2.1.3.1.2 IBC Interaction with Other Coding Tools The interaction between the IBC mode and other inter-frame coding tools in VVC (such as paired Merge candidates, history-based motion vector prediction (HMVP), intra / inter joint prediction mode (CIIP), Merge mode with motion vector difference (MMVD), and geometric partitioning mode (GPM)) is as follows: –IBC can be used with pairwise Merge candidates and HMVP. A new pairwise IBC Merge candidate can be generated by averaging two IBC Merge candidates. For HMVP, the IBC motion is inserted into the history buffer for future reference. –IBC cannot be combined with the following inter-frame tools: affine motion, CIIP, MMVD, and GPM. –When using DUAL_TREE partitioning, IBC is not allowed for chroma-coded blocks. Different from the HEVC screen content coding extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation process of the motion vector for the IBC mode excludes all neighboring blocks in the inter-frame mode and vice versa. The following IBC design aspects are applied: –IBC shares the same process as regular MV Merge (including pairwise Merge candidates and history-based motion prediction values), but TMVP and zero vectors are not allowed as they are invalid for the IBC mode. –Separate HMVP caches (each with 5 candidates) are used for regular MV and IBC. –The block vector constraints are implemented in the form of bitstream consistency constraints. The encoder needs to ensure that there are no invalid vectors in the bitstream and that if a Merge candidate is invalid (out of range or 0), the Merge is not used. Such bitstream consistency constraints are represented by the virtual buffer as described below. –For deblocking, IBC is treated as an inter-frame mode. –If the current block is coded using the IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to indicate only whether the MV is inter-pixel or 4-integer pixels. –The number of IBC Merge candidates can be signaled in the slice header separately from the number of regular, sub-block, and geometric Merge candidates. The virtual buffer concept is used to describe the allowed reference region for the IBC prediction mode and valid block vectors. Representing the CTU size as ctbSize, the virtual buffer ibcBuf has a width wIbcBuf = 128x128 / ctbSize and a height hIbcBuf = ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of the VPDU is min(ctbSize, 64) in each dimension, W v= min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows. – At the start of decoding each CTU row, the entire ibcBuf is flushed with the invalid value -1. – At the start of decoding the VPDU (xVPDU, yVPDU) relative to the top - left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU % wIbcBuf,..., xVPDU % wIbcBuf + W v -1; y = yVPDU % ctbSize,..., yVPDU % ctbSize + W v -1. – After decoding the CU containing (x, y) relative to the top - left corner of the picture, set ibcBuf[x % wIbcBuf][y % ctbSize] = recSample[x][y] For a block covering the coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid: ibcBuf[(x + bv[0]) % wIbcBuf][(y + bv[1]) % ctbSize] should not be equal to -1. 2.1.3.2. Block Differential Pulse - Coding Modulation (BDPCM) VVC supports Block Differential Pulse - Coding Modulation (BDPCM) for screen content coding and decoding. At the sequence level, the BDPCM enable flag is signaled in the SPS; this flag is signaled only when the transform skip mode (described in the next section) is enabled in the SPS. When BDPCM is enabled, if the CU size is less than or equal to MaxTsSize × MaxTsSize in terms of luma samples and if the CU is intra - coded, then a flag is sent at the CU level, where MaxTsSize is the maximum block size that allows the transform skip mode. This flag indicates whether to use regular intra - coding or BDPCM. If BDPCM is used, a BDPCM prediction direction flag is sent to indicate whether the prediction is horizontal or vertical. Then, a regular horizontal or vertical intra - prediction process with unfiltered reference samples is used to predict the block. The residual is quantized, and the difference between each quantized residual and its predicted value (i.e., the previously coded residual at the horizontal or vertical (depending on the BDPCM prediction direction) neighboring position) is coded. For a block of size M (height) × N (width), let r i,j , 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 be the prediction residual. Let Q(r i,j), where 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1 represents the quantized version of the residual r i,j The BDPCM is applied to the quantized residual values, thereby generating a modified M×N array with elements of where is predicted from its neighboring quantized residual values. For the vertical BDPCM prediction mode, for 0 ≤ j ≤ (N - 1), the following is used to derive For the horizontal BDPCM prediction mode, for 0 ≤ i ≤ (M - 1), the following is used to derive On the decoder side, the above process is reversed to calculate Q(r i,j ), where 0 ≤ i ≤ M - 1, 0 ≤ j ≤ N - 1, as follows: The inverse quantized residual Q -1 (Q(r i,j )) is added to the intra-block prediction value to generate the reconstructed sample value. Using the same residual coding and decoding process as in the transform skip mode residual coding and decoding, the predicted quantized residual values are sent to the decoder. For lossless coding and decoding, if the slice_ts_residual_coding_disabled_flag is set to 1, the quantized residual values are sent to the decoder using the regular transform residual coding and decoding. In terms of the MPM mode for future intra-mode coding and decoding, if the BDPCM prediction directions are horizontal or vertical respectively, the horizontal or vertical prediction mode is stored for the CU for BDPCM coding and decoding. For deblocking, if two blocks on both sides of a block boundary use BDPCM coding and decoding, then that specific block boundary is not deblocked. 2.1.3.3. Residual Coding and Decoding for Transform Skip Mode VVC allows the transform skip mode for luma blocks with sizes up to MaxTsSize×MaxTsSize, where the value of MaxTsSize is signaled in the PPS and is at most 32. When a CU is coded in the transform skip mode, its prediction residual is quantized and coded using the transform skip residual coding and decoding process. This process is modified from the transform coefficient coding and decoding process. In the transform skip mode, the residual of the TU is also coded in non-overlapping sub-blocks of size 4x4. For better coding and decoding efficiency, some modifications are made to customize the residual coding and decoding process for the characteristics of the residual signal. The following summarizes the differences between the transform skip residual coding and decoding and the regular transform residual coding and decoding: – The forward scan order is applied to sub - blocks within the scan - transformed block and positions within the sub - blocks; – There is no signaling of the final (x, y) position; – When all previous flags are equal to 0, coded_sub_block_flag is coded / decoded for each sub - block except the last sub - block; – sig_coeff_flag context modeling uses a reduced template, and the context model of sig_coeff_flag depends on top and left neighbor values; – The context model of the abs_level_gt1 flag also depends on left and top sig_coeff_flag values; – par_level_flag uses only one context model; – Additionally, more than 3, 5, 7, 9 flags are signaled to indicate coefficient levels, one context per flag; – For the binarization of the remaining values, Rice parameter derivation with a fixed order = 1 is used; – The context model of the sign flag is determined based on left and above neighbor values, and the sign flag is parsed after sig_coeff_flag to keep all context - coded bits together. For each sub - block, if coded_subblock_flag is equal to 1 (i.e., there is at least one non - zero quantized residual in the sub - block), the coding / decoding of the quantized residual levels is performed in three scan passes (see Figure 27 ): – First scan pass: The validity flag (sig_coeff_flag), sign flag (coeff_sign_flag), absolute level greater than 1 flag (abs_level_gtx_flag[0]), and parity (par_level_flag) are coded / decoded. For a given scan position, if sig_coeff_flag is equal to 1, then coeff_sign_flag is coded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[0] is equal to 1, then additionally par_level_flag is coded to specify the parity of the absolute level. – Scan passes greater than x: For each scan position with an absolute level greater than 1, up to four abs_level_gtx_flag[i] for i = 1 ··· 4 are coded to indicate whether the absolute level at the given position is greater than 3, 5, 7, or 9 respectively. – Remaining scan passes: The absolute remainder abs_remainder is coded and decoded in bypass mode. The absolute remainder is binarized using a fixed Rice parameter value of 1. The bits in scan passes #1 and #2 (the first scan pass and scan passes greater than x) are context-coded until the maximum number of context-coded bits in the TU has been exhausted. The maximum number of context-coded bits in the residual block is limited to 1.75 * block_width * block_height, or equivalently, 1.75 context-coded bits per sample position. The bits in the last scan pass (the remaining scan pass) are bypass-coded. The variable RemCcbs is first set to the maximum number of context-coded bits for the block and is decremented by 1 whenever a context-coded bit is coded. When RemCcbs is greater than or equal to 4, the context-coded bits are used to code the syntax elements including sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag in the first coding pass. If RemCcbs becomes less than 4 during the coding of the first pass, then the remaining coefficients not yet coded in the first pass are coded in the remaining scan pass (pass #3). After the first pass coding is complete, if remCcbs is greater than or equal to 4, the context-coded bits are used to code the syntax elements including abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag in the second coding pass. If RemCcbs becomes less than 4 during the coding of the second pass, then the remaining coefficients not yet coded in the second pass are coded in the remaining scan pass (pass #3). Figure 27 The transform skip residual coding process is shown. The star marks the position when the context-coded bits are exhausted, at which point bypass coding is used to code all the remaining bits. In addition, for blocks not encoded or decoded in the BDPCM mode, a level mapping mechanism is applied to transform skip residual encoding and decoding until the maximum number of context encoding bits has been reached. The level mapping uses the top and left neighboring coefficient levels to predict the current coefficient level in order to reduce the signaling cost. For a given residual position, let absCoeff represent the absolute coefficient level before mapping, and let absCoeffMod represent the coefficient level after mapping. Let X0 represent the absolute coefficient level at the left neighboring position, and let X1 represent the absolute coefficient level at the upper neighboring position. The level mapping is performed as follows: pred = max(X0, X1); if (absCoeff == pred) absCoeffMod = 1; else absCoeffMod = (absCoeff < pred)? absCoeff + 1 : absCoeff; Then, the absCoeffMod value is encoded and decoded as described above. After all context encoding bits have been exhausted, the level mapping is disabled for all remaining scan positions in the current block. 2.1.3.4 Palette Mode In VVC, the palette mode is used for screen content encoding and decoding in all chroma formats supported in the 4:4:4 profile (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome). When the palette mode is enabled, if the CU size is less than or equal to 64x64, then a flag is transmitted at the CU level, and the number of samples in the CU is greater than 16 to indicate whether the palette mode is used. Considering that applying the palette mode on small CUs introduces insignificant encoding and decoding gains and additional complexity on small blocks, the palette mode is disabled for CUs with less than or equal to 16 samples. The encoded and decoded coding unit (CU) via the palette is regarded as a prediction mode different from intra prediction, inter prediction, and intra block copy (IBC) modes. If the palette mode is utilized, then the sample values in the CU are represented by a set of representative color values. This set is called a palette. For positions with sample values close to the palette colors, the palette index is signaled. It is also possible to specify samples outside the palette by signaling an escape symbol. For samples within the CU encoded with the escape symbol, their component values are directly signaled using (possibly) quantized component values. This is shown in Figure 28 The quantized escape symbol is binary coded using a fifth-order exponential Golomb binary process (EG5). For the encoding and decoding of the palette, the palette prediction value is maintained. For non-wavefront cases, the palette prediction value is initialized to 0 at the start of each slice. For WPP cases, the palette prediction value at the start of each CTU row is initialized to the prediction value derived from the first CTU in the previous CTU row, such that the initialization scheme between the palette prediction value and CABAC synchronization is unified. For each entry in the palette prediction value, a reuse flag is signaled to indicate whether it is part of the current palette in the CU. The reuse flag is sent using run-length encoding of zeros. Thereafter, the number of new palette entries and the component values for the new palette entries are signaled. After encoding and decoding the palette-encoding CU, the palette prediction value is updated using the current palette, and the entries from the previous palette prediction value that are not reused in the current palette are added at the end of the new palette prediction value until the maximum allowed size is reached. An escape flag is signaled for each CU to indicate whether there is an escape symbol in the current CU. If there is an escape symbol, the palette table is augmented by one, and the last index is assigned to the escape symbol. In a manner similar to the coefficient groups (CGs) used in transform coefficient encoding, the CUs encoded using the palette mode are partitioned into multiple line-based coefficient groups, each coefficient group consisting of m samples (i.e., m = 16), where for each CG the index run, the palette index value, and the quantized color for the escape mode are encoded / parsed in sequence. As in HEVC, a horizontal or vertical traversal scan can be applied to scan the samples, as Figure 29 shown. The encoding order for palette run encoding and decoding in each segment is as follows: For each sample position, one context-encoded binary bit run_copy_flag = 0 is signaled to indicate whether the pixel has the same pattern as the previous sample position, i.e., whether the previously scanned sample and the current sample are both of run type COPY_ABOVE, or whether the previously scanned sample and the current sample are both of run type INDEX and have the same index value. Otherwise, run_copy_flag = 1 is signaled. If the current sample has a different pattern from the previous sample, then one context-encoded binary bit copy_above_palette_indices_flag is signaled to indicate the run type of the current sample, i.e., INDEX or COPY_ABOVE. Here, if the sample is in the first row (horizontal traversal scan) or in the first column (vertical traversal scan), the decoder does not have to parse the run type because the INDEX mode is used by default. In the same way, if the previously parsed run type is COPY_ABOVE, the decoder does not have to parse the run type. After the palette run encoding and decoding of the samples in one encoding and decoding pass, the index values (for the INDEX mode) and the quantized escape colors are grouped and encoded and decoded using CABAC bypass encoding and decoding in another encoding and decoding pass. This separation of context-encoded binary bits and bypass-encoded binary bits can improve the throughput within each row of CG. For a slice with a dual-luma / chroma tree, the palette is applied to the luma (Y component) and chroma (Cb and Cr components) separately, where the luma palette entries contain only Y values, and the chroma palette entries contain both Cb and Cr values. For a slice with a single tree, the palette is applied jointly over the Y, Cb, Cr components, i.e., each entry in the palette contains Y, Cb, Cr values unless when using local dual-tree encoded CUs, in which case the encoding and decoding of luma and chroma are handled separately. In this case, if the palette mode is used to encode and decode the corresponding luma block or chroma block, their palettes are applied in a way similar to the dual-tree case (this is related to non-4:4:4 encoding and will be further explained in 2.1.3.4.1). For a slice encoded using dual-tree encoding, the maximum palette prediction value size is 63, and the maximum palette table size for the encoding of the current CU is 31. For a slice encoded using dual-tree encoding, the maximum prediction value and palette table size are halved, i.e., for each of the luma palette and chroma palette, the maximum prediction value size is 31 and the maximum table size is 15. For deblocking, the palette-encoded blocks on the sides of the block boundary are not deblocked. 2.1.3.4.1. Palette Mode for Non-4:4:4 Content Support the palette mode in VVC in a manner similar to the palette mode in HEVC SCC for all chroma formats. For non-4:4:4 content, the following customizations are applied: 1. When signaling an escape value for a given sample position, if the sample position has only a luminance component but no chrominance component due to chroma subsampling, then only the luminance escape value is signaled. This is the same as in HEVC SCC. 2. For local double-tree blocks, the palette mode is applied to the block in the same way as the palette mode applied to a single-tree block with two exceptions: a. The process of updating the palette prediction value is slightly modified as follows. Since the local double-tree block contains only the luminance (or chrominance) component, the prediction value update process uses the value of the luminance (or chrominance) component signaled, and fills in the "missing" chrominance (or luminance) by setting it to the default value (1<<(component bit depth - 1)). Component. b. The maximum palette prediction value size remains 63 (since the strip is coded using a single tree), but the maximum palette table size for luminance / chrominance blocks remains 15 (since the block is coded using a single palette). 3. For the palette mode in monochrome format, the number of color components in the palette coded block is set to 1 instead of 3. 2.1.3.4.2. Encoder algorithm for the palette mode On the encoder side, the following steps are used to generate the palette table for the current CU. 1. First, to derive the initial entries in the palette table for the current CU, a simplified K-means clustering is applied. Initialize the palette table for the current CU to an empty table. For each sample position in the CU, calculate the SAD between the sample and each palette table entry, and obtain the minimum SAD among all palette table entries. If the minimum SAD is less than a predefined error limit errorLimit, the current sample is clustered with the palette table entry having the minimum SAD. Otherwise, a new palette table entry is created. The threshold errorLimit is QP-dependent and is retrieved from a lookup table containing 57 elements covering the entire QP range. After all the samples in the current CU have been processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entry after the 31st entry is discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroid of each cluster from step 1 or using one of the palette colors in the palette prediction values. The option with the lower rate - distortion cost is selected as the final color of the palette table. If a cluster has only a single sample point and the corresponding palette entry is not in the palette prediction values, then the corresponding sample point is converted to an escape symbol in the next step. 3. The resulting palette table contains some new entries from the centroids of the clusters in step 1, as well as some entries from the palette prediction values. Therefore, the table is reordered again so that all new entries (i.e., the centroids) are placed at the start of the table, followed by the entries from the palette prediction values. Given the palette table for the current CU, the encoder selects the palette index for each sample point position in the CU. For each sample point position, the encoder checks the RD cost for all index values corresponding to the palette table entries and the index representing the escape symbol, and selects the index with the minimum RD cost using the following equation: RD cost = distortion × (isChroma? 0.8:1) + lambda × by - pass encoding / decoding bits (2 - 33) After determining the index map for the current CU, each entry in the palette table is checked to see if it is used by at least one sample point position in the CU. Any unused palette entry is removed. After determining the index map for the current CU, grid RD optimization is applied to find the best values of run_copy_flag and run type for each sample point position by comparing the RD costs of the following three options: the same as the previously scanned position, run type COPY_ABOVE, or run type INDEX. When calculating the SAD value, the sample values are scaled down to 8 bits, unless the CU is encoded / decoded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. Additionally, in the case of lossless encoding / decoding, only the bitrate is used in the above rate - distortion optimization step (since lossless encoding / decoding does not cause distortion). 2.1.3.5. Adaptive Color Transformation In the HEVC SCC extension, an Adaptive Color Transformation (ACT) is applied to reduce the redundancy between the three color components in the 4:4:4 chroma format. ACT is also adopted in the VVC standard to enhance the encoding / decoding efficiency of 4:4:4 chroma format encoding / decoding. Similar to HEVC SCC, ACT performs a loop color space transformation in the prediction residual domain by adaptively transforming the residual from the input color space to the YCgCo space. Figure 30The decoding flowchart of applying ACT is shown. Two color spaces are adaptively selected by signaling an ACT flag at the CU level. When the flag equals 1, the residual of the CU is coded and decoded in the YCgCo space; otherwise, the residual of the CU is coded and decoded in the original color space. Additionally, similar to the HEVC ACT design, for inter and IBC CUs, ACT is enabled only when there is at least one non-zero coefficient in the CU. For intra CUs, ACT is enabled only when the chrominance component selects the same intra prediction mode as the luminance component (i.e., DM mode). 2.1.3.5.1. ACT mode In the HEVC SCC extension, ACT supports both lossless and lossy coding and decoding based on the lossless flag (i.e., cu_transquant_bypass_flag). However, there is no flag signaled in the bitstream to indicate whether lossy or lossless coding and decoding is applied. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown as follows. Since the YCgCo-R transform is not normalized, in order to compensate for the dynamic range change of the residual signal before and after the color transform, QP adjustments of (-5, 1, 3) are applied to the transform residuals of the Y, Cg, and Co components respectively. The adjusted quantization parameter only affects the quantization and inverse quantization of the residuals in the CU. For other coding and decoding processes (such as deblocking), the original QP is still applied. Additionally, because the forward and reverse color transforms require access to the residuals of all three components, the ACT mode is always disabled for split tree partitioning and ISP modes where the predicted block sizes of different color components are different. When ACT is applied, the transform skip (TS) and block differential pulse coding modulation (BDPCM) extended to the coded chrominance residuals are also enabled. 2.1.3.5.2 ACT fast coding algorithm To avoid brute-force R-D search in both the original color space and the transformed color space, the following fast coding algorithm is applied in the VTM reference software to reduce the encoder complexity when ACT is enabled. – The order of the RD check for enabling / disabling ACT depends on the original color space of the input video. For RGB videos, the RD cost of the ACT mode is checked first; for YCbCr videos, the RD cost of the non-ACT mode is checked first. The RD cost of the second color space is checked only if there is at least one non-zero coefficient in the first color space. – When a CU is obtained through different splitting paths, reuse the same ACT enable / disable decision. Specifically, when a CU is coded / decoded for the first time, the selected color space used to code / decoded the residual of a CU will be stored. Then, when the same CU is obtained by another splitting path, instead of checking the RD cost of two spaces, the stored color space decision will be directly reused. – The RD cost of the parent CU is used to determine whether to check the RD cost of the second color space for the current CU. For example, if the RD cost of the first color space is less than the RD cost of the second color space for the parent CU, then for the current CU, the second color space is not checked. – To reduce the number of tested coding / decoding modes, share the selected coding / decoding modes between two color spaces. Specifically, for intra modes, the preselected intra mode candidates based on SATD-based intra mode selection are shared between two color spaces. For inter and IBC modes, the block vector search or motion estimation is performed only once. The block vectors and motion vectors are shared by two color spaces. 2.1.3.6. Intra Template Matching (IntraTMP) Intra Template Matching Prediction (IntraTM) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, and its L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed part of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side. By matching the L-shaped causal neighbor of the current block with Figure 31 another block in the predefined search area in R1: the current CTU R2: the top-left CTU R3: the upper CTU R4: the left CTU. SAD is used as the cost function. Within each area, the decoder searches for the template with the minimum SAD relative to the current template and uses the corresponding block as the prediction block. The dimensions of all areas (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH where 'a' is a constant that controls the gain / complexity trade-off. In fact, 'a' is equal to 5. For CUs with dimensions of width and height less than or equal to 64, the intra-frame template matching tool is enabled. This maximum CU size for intra-frame template matching is configurable. When DIMD is not used for the current CU, the intra-frame template matching prediction mode is signaled at the CU level via a dedicated flag. 2.1.3.6.1. Using Block Vectors Derived from IntraTMP for IBC Block vectors (BVs) derived from intra-frame template matching prediction (IntraTMP) are used for intra-frame block copy (IBC). The stored IntraTMP BVs of neighboring blocks and the IBC BVs are used as spatial BV candidates in IBC candidate list construction. 2.1.3.6.2. Direct Block Vector (DBV) Mode for Chrominance Prediction For the chrominance component, when the chrominance dual-tree is activated in an intra-frame strip, if one of the luminance blocks (five positions) is encoded / decoded using MODE_IBC, then its block vector bvL is used and scaled to derive the chrominance block vector bvC. The scaling factor depends on the chrominance format sampling structure. Then, by using the position (xCb, yCb) of the current chrominance block and its bvC, the corresponding offset position (xCb + bvC[0], yCb + bvC[1]) is determined, and block copy prediction is performed. Figure 32 Five positions in the reconstructed luminance samples are shown. Figure 33 The prediction process of the DBV mode is shown. A CU-level flag is signaled to indicate whether the proposed DBV mode is applied, as shown in Table 7. Table 7 Binaryization Process for intra_chroma_pred_mode in the Proposed Method intra_chroma_pred_mode Binary bit Chroma intra mode 0 11100 List [0] 1 11101 List [1] 2 11110 List [2] 3 11111 List [3] 4 110 DIMD Chroma 5 10 DM 6 0 DBV 2.1.4 Transform and Quantization 2.1.4.1. Large Block Size Transform with High-Frequency Zeroing In VVC, large block size transforms with sizes up to 64x64 are enabled, which are mainly used for higher resolution videos such as 1080p and 4K sequences. For transform blocks with a size equal to 64 (width or height, or both width and height), the high-frequency transform coefficients are set to zero, so that only the lower frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the transform coefficients of the left 32 columns are retained. Similarly, when N equals 64, only the transform coefficients of the first 32 rows are retained. When the transform skip mode is used for large blocks, the entire block is used without zeroing any values. Additionally, the transform shift is removed in the transform skip mode. VTM also supports a configurable maximum transform size in the SPS, enabling the encoder to have the flexibility to select a transform size of up to 32 lengths or 64 lengths according to the requirements of a specific implementation. 2.1.4.2 Multiple Transform Selection (MTS) for Kernel Transforms In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of blocks coded inter- and intra-frame. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 8 shows the basis functions of the selected DST / DCT. Table 8 - Transform Basis Functions of DCT-II / VIII and DSTVII for N-Point Input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra- and inter-frame respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS only applies to luminance. The MTS signaling is skipped when one of the following conditions is met: - The position of the last significant coefficient of the luminance TB is less than 1 (i.e., only DC). - The last significant coefficient of the luminance TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 9. A unified transform selection for ISP and implicit MTS is used by eliminating the intra-mode and block shape dependencies. If the current block is in ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for the horizontal and vertical transform kernels. In terms of the transform matrix precision, 8-bit primary transform kernels are used. Thus, all the transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, for other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8), 8-bit primary transform kernels are used. Table 9 - Transform and Signaling Mapping Table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16×16 low-frequency region are retained. Similar to HEVC, the residual of a block can be encoded and decoded using the transform skip mode. To avoid redundancy in syntax encoding and decoding, when the MTS_CU_flag at the CU level is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for inter-coded blocks, implicit MTS can still be enabled. 2.1.4.3 Low-Frequency Non-Separable Transform (LFNST) In VVC, LFNST is applied between the forward primary transform and quantization (at the encoder) and between the de-quantization and inverse primary transform (at the decoder side). In LFNST, a 4x4 non-separable transform or an 8x8 non-separable transform is applied according to the block size. For example, 4x4 LFNST is applied to small blocks (i.e., min(width, height) < 8), while 8x8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). Figure 34 The low-frequency non-separable transform (LFNST) process is shown. The following uses the input as an example to describe the application of the non-separable transform used in LFNST. To apply 4x4 LFNST, the 4x4 input block X is first represented as a vector The non-separable transform is calculated as where denotes the transform coefficient vector, and T is a 16x16 transform matrix. Subsequently, the 16x1 coefficient vector is reorganized into 4x4 blocks using the scan order (horizontal, vertical, or diagonal) for the block. Coefficients with smaller indices are placed in the 4x4 coefficient block together with smaller scan indices. 2.1.4.3.1. Reduced non-separable transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on a direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix dimension to minimize the computational complexity and memory space for storing the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in the LFNST. The main idea of this reduced non-separable transform is to map an N-dimensional vector (for an 8x8 NSST, N is typically equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an NxN matrix, the RST matrix becomes an R×N matrix as follows: Among them, the R transformed rows are the R bases of the N-dimensional space. The inverse transformation matrix for RT is the transpose of its forward transformation. For the 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix (which is the conventional 8x8 non-separable transformation matrix size) is reduced to a 16x48 direct matrix. Therefore, a 48×16 inverse RST matrix is used on the decoder side to generate the core (main) transformation coefficients in the upper left 8×8 region. When applying a 16x48 matrix instead of a 16x64 with the same transformation set configuration, each of them obtains 48 input data from three 4x4 blocks in the upper left 8x8 block except for the lower right 4x4 block. With the reduced dimension, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, accompanied by a reasonable performance degradation. To reduce complexity, LFNST is restricted to be applicable only when all coefficients outside the first coefficient subgroup are insignificant. Therefore, when applying LFNST, all only the main transformation coefficients must be zero. This allows adjusting the LFNST index signaling at the last valid position and thus avoids the additional coefficient scanning in the current LFNST design, which requires checking for valid coefficients only at specific positions. The worst-case processing of LFNST (in terms of multiplications per pixel) is restricted to 8x16 and 8x48 transformations for non-separable transformations of 4x4 and 8x8 blocks respectively. In these cases, when applying LFNST, for other sizes less than 16, the last valid scan position must be less than 8. For blocks with shapes of 4xN and Nx4 where N>8, the proposed constraints imply that LFNST is now applied only once and only to the upper left 4×4 region. Since all only the main coefficients are zero when applying LFNST, the number of operations required for the main transformation is reduced in this case. From the perspective of the encoder, when testing the LFNST transformation, the quantization of coefficients is significantly simplified. Rate-distortion optimized quantization must be maximally completed for the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.1.4.3.2 LFNST Transform Selection A total of 4 transformation sets and 2 non-separable transformation matrices (kernels) for each transformation set are used in LFNST. As shown in Table 10, the mapping from the intra prediction mode to the transformation set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transformation set 0 is selected for the current chrominance block. For each transformation set, the selected non-separable quadratic transformation candidate is further specified by the LFNST index signaled explicitly. The index is signaled once per intra CU in the bitstream after the transformation coefficients. Table 10 - Transformation Selection Table 2.1.4.3.3. LFNST Index Signaling and Interaction with Other Tools Since LFNST is constrained to be applicable only when all coefficients outside the first coefficient subgroup are not significant, the LFNST index encoding and decoding depends on the position of the last significant coefficient. Additionally, the LFNST index is context - decoded but does not depend on the intra - prediction mode, and only the first binary bit is context - decoded. Furthermore, LFNST is applied to intra CUs in both intra and inter stripes and for both luminance and chrominance. If dual - tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter stripes (dual - tree disabled), a single LFNST index is signaled and used for both luminance and chrominance. Considering that due to the existing maximum transform size constraint (64x64), large CUs larger than 64x64 are implicitly partitioned (TU slicing), the LFNST index search can increase the data buffer by up to four times the number of decoding pipeline stages. Therefore, the maximum size allowing LFNST is constrained to 64x64. Note that LFNST only enables DCT2. The LFNST index signaling is placed before the MTS index signaling. The use of the scaling matrix for perceptual quantization is not obvious, and the scaling matrix specified for the main matrix can be used for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For the single - tree split mode, chrominance LFNST is not applied. 2.1.4.4. Sub - block Transform (SBT) In VTM, a sub - block transform is introduced for CUs in inter - prediction. In this transform mode, for a CU, only a sub - part of the residual block is encoded and decoded. When the cu_cbf of an inter - predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub - part of the residual block is encoded and decoded. For the former case, the inter - frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by an inferred adaptive transform while the other part of the residual block is zeroed. When SBT is used for inter - frame - decoded CUs, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 35As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in terms of luminance samples), a 1:3 / 3:1 partition along that dimension is not allowed. A CU can have at most 8 SBT modes. Position-dependent transform kernel selection is applied to the luminance transform blocks in SBT-V and SBT-H (chrominance TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are Figure 35 specified in. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transforms jointly specify the TU slice, cbf, and the horizontal and vertical kernel transform types of the residual block. SBT is not applied to CUs that utilize the inter-intra joint mode for coding and decoding. 2.1.4.5 Maximum Transform Size and Zeroing of Transform Coefficients The CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the maximum intra-coding block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, there is no canonical zeroing operation applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are normalized to zero. 2.1.4.6. Enhanced MTS for Intra-Coding In the current VVC design [1], for MTS, only the DST7 and DCT8 transform kernels are utilized, which are used for both intra- and inter-coding. Additional main transforms including DCT5, DST4, DST1, and the identity transform (IDT) are adopted. The MTS set also depends on the TU size and the intra-mode information. Sixteen different TU sizes are considered, and for each TU size, five different categories are considered according to the in-mode information. For each category, four different transform pairs are considered, which are the same as the transform pairs in VVC. Note that although a total of 80 different categories are considered, some of these different categories usually share exactly the same transform set. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For the angular mode, joint symmetry on the TU shape and intra prediction is considered. Thus, mode i (i > 34) with TU shape AxB will be mapped to the same class corresponding to mode j = (68 - i) with TU shape BxA. However, for each transform pair, the order of the horizontal and vertical transform kernels is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same class. However, the vertical and horizontal transform kernels are swapped. For the wide-angle mode, the closest regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to mode 80. The MTS indices [0,3] are signaled using 2-bit fixed-length coding and decoding. 2.1.4.7. Quadratic Transform: LFNST Extension with Large Kernels The LFNST design in VVC is extended as follows: ● The number of LFNST sets (S) and the number of candidates (C) are extended to S = 35 and C = 3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: ○ For predModeIntra < 2, lfnstTrSetIdx is equal to 2 ○ For predModeIntra in [0,34], lfnstTrSetIdx = predModeIntra ○ For predModeIntra in [35,66], lfnstTrSetIdx = 68 - predModeIntra ● Three different kernel LFNST4, LFNST8, and LFNST16 are defined to indicate the sets of LFNST kernels applied to 4xN / Nx4 (N ≥ 4), 8xN / Nx8 (N ≥ 8), and MxN (M, N ≥ 16), respectively. The kernel dimensions are specified as follows: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96). The forward LFNST is applied to the upper-left low-frequency region called the region of interest (ROI). When applying the LFNST, the main transform coefficients existing in the regions other than the ROI are set to zero, which is the same as in the VVC standard. The ROI for LFNST16 is at Figure 36It is shown in . It consists of six 4x4 sub-blocks, and these sub-blocks are consecutive in the scanning order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be Rx96. In this contribution, R is chosen to be 32, and accordingly 32 coefficients (two 4x4 sub-blocks) are generated from the forward LFNST16, and they are placed following the coefficient scanning order. The ROI for LFNST8 is in Figure 37 shown in . The forward LFNST8 matrix can be Rx64, and R is chosen to be 32. The generated coefficients are located in the same way as LFNST16. The mapping from the intra prediction mode to these sets is shown in the following table, Table 11. Mapping of Intra Prediction Mode to LFNST Set Index Intra prediction mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNST set index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNST set index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra prediction mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST set index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.1.4.8. Inseparable Primary Transform (NSPT) for Intra Coding and Decoding DCT-II + LFNST is replaced by NSPT for block sizes 4x4, 4x8, 8x4, and 8x8. NSPT follows the design of LFNST (i.e., 3 candidates and 35 sets) based on intra mode selection. The kernel sizes are as follows: ● NSPT4x4: 16x16; ● NSPT4x8 / NSPT8x4: 32x20; ● NSPT8x8: 64x32. Therefore, 12 and 32 coefficients are zeroed for NSPT4x8 / NSPT8x4 and NSPT8x8 respectively. 2.1.4.9. Sign Prediction The basic idea of the coefficient sign prediction methods (JVET-D0031 and JVET-J0021) is to calculate the reconstructed residuals for the negative and positive sign combinations of the applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal sign, the cost function is defined as the measurement of discontinuity across the Figure 38 block boundaries shown in . The measurement is performed for all hypotheses, and the hypothesis with the minimum cost is selected as the predicted value for the coefficient sign. The cost function is defined as the sum of the absolute second derivatives in the residual domain for the upper row and the left column, as follows: where R is the reconstruction neighbor, P is the prediction of the current block, and r is the residual hypothesis. The term (-R-1 + 2R0 - P1) can be calculated only once per block and only the residual hypothesis is subtracted. 3. Problem In ECM-7.0, OBMC can be applied to inter-frame coding and decoding blocks, whether they are inter-frame AMVP coded or inter-frame MERGE coded. For inter-frame AMVP coded blocks, a syntax flag indicating whether to apply OBMC can be signaled at the block level. For inter-frame MERGE coded blocks, it is implicitly assumed that OBMC is applied regardless of the block characteristics and the coding information of neighboring blocks. However, there may be cases where some inter-frame MERGE coded blocks do not prefer the OBMC mode. For example, a block containing sharp edges, or few gradients, or little color may not prefer the OBMC mode. Block-level adaptive OBMC that considers the prediction mode of neighboring blocks can bring higher coding and decoding gain. 4. Detailed solutions The following detailed solutions should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term "video unit" or "coding unit" or "block" can represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. In this disclosure, regarding "a block coded in mode N", here "mode N" can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards. Any change in coding tools is also applicable. 4.1. In one example, whether OBMC is applied to the current block can depend on the sample values of the samples within the current block (and / or neighboring the current block). 1) For example, the current block can be coded by inter-frame Merge. 2) For example, the current block can be coded by inter-frame AMVP. 3) For example, it can be based on the predicted samples of the current block (before OBMC). 4) For example, it can be based on the predicted samples neighboring the current block. 5) For example, it can be based on reconstructed samples adjacent to the current block. 6) For example, it can be based on the gradient / direction / angle (or histogram of gradient / direction / angle) of samples within (and / or adjacent to) the current block. a. For example, the current predicted samples before OBMC can be used. b. For example, adjacent reconstructed samples can be used. c. For example, the histogram of gradient / direction / angle can be calculated based on counting the gradients along a specific direction / angle. i. For example, specific directions / angles can be predefined. ii. For example, the specific directions / angles can be based on the directions of intra prediction angle modes in video coding / decoding. iii. For example, for a specific direction / angle, the gradient magnitude can be calculated based on counting the gradients (magnitudes) of at least one sample in the current block. 1. For example, the gradients (magnitudes) at a specific location (e.g., the center) in the current block can be counted. 2. For example, the gradients (magnitudes) at a series of specific locations in the current block can be counted. 3. For example, the gradients (magnitudes) of all samples in the current block can be counted. 4. For example, the gradients (magnitudes) of all samples except those in the first row, last row, first column, and last column of the current block can be counted. iv. For example, for a specific direction / angle, the gradient magnitude can be calculated based on counting the gradients (magnitudes) of at least one sample adjacent to the current block. 1. For example, the gradients (magnitudes) at a specific location adjacent to the current block can be counted. 2. For example, the gradients (magnitudes) of all samples adjacent to the current block (on the left and / or top) can be counted. d. For example, the histogram of gradient / direction / angle can be calculated based on dividing the entire range of direction / angle into a series of intervals / bits. direction / angle histogram. e. For example, the histogram of gradient / direction / angle can be calculated based on counting the gradients (magnitudes) in each interval / bit / direction / angle. 7) For example, it can be based on the color / brightness / intensity (or histogram of color / brightness / intensity) of samples within (and / or adjacent to) the current block. a. For example, the current predicted samples before OBMC can be used. b. For example, neighboring reconstructed samples can be used. c. For example, a histogram of color / brightness / intensity can be calculated based on counting the sample values in the Y and / or U and / or V (or, R and / or G and / or B) component domains. i. For example, the sample values at a series of specific positions in the current block can be counted. ii. For example, the sample values of all samples in the current block can be counted. iii. For example, the sample values at specific positions neighboring the current block can be counted. iv. For example, the sample values of all samples neighboring the current block (on the left side and / or the top) can be counted. d. For example, a histogram of color / brightness / intensity can be calculated based on dividing the entire range of color / brightness / intensity values into a series of intervals / bits. e. For example, a histogram of color / brightness / intensity can be calculated based on counting the number of samples in each interval / bit. 8) For example, it can be based on the number of main gradients / directions / angles / colors / brightness / intensities of samples within (and / or neighboring) the current block. a. For example, it can be calculated based on the predicted samples within the current block (before OBMC). b. For example, it can be calculated based on the reconstructed samples neighboring the current block. c. For example, the main gradient / direction / angle / color / brightness / intensity can be derived based on the histogram of gradient / direction / angle / color / brightness / intensity. d. For example, the main gradient / direction / angle / color / brightness / intensity can be derived based on how many intervals / bits in the histogram show values (e.g., gradient magnitude, color value, brightness value) greater than a threshold. e. For example, the main gradient / direction / angle / color / brightness / intensity can be derived based on how many intervals / bits in the histogram provide values (e.g., gradient magnitude, color value, brightness value) greater than the values of other intervals / bits. i. For example, the values of the intervals / bits in the histogram can be sorted first, assuming the sorted values (e.g., from largest to smallest) are represented by X0, X1, X2…, X n-2 、X n-1 where n intervals / bits are included in the histogram. If X i >= a*(X i+1) Then, the interval / binary bits from 0 to i can be regarded as the main gradient, where a represents a scaling factor (for example, a can be equal to a constant between 2 and 20). f. For example, if the number of main gradients / directions / angles / colors / brightness / intensities is less than a certain number (for example, 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8), then OBMC may not be applied to the block. 4.2. In one example, whether OBMC is applied to the current block may depend on the template cost. 1) For example, the current block can be coded / decoded by inter-frame Merge. 2) For example, the current block can be coded / decoded by inter-frame AMVP. 3) For example, it can be based on the first non-blended template cost and the second blended template cost. a. For example, the first template cost can be calculated based on the SAD between the current template and the reference template (where the reference template is identified by adding the current motion vector to the position of the current template). b. For example, the second template cost can be calculated based on the SAD between the current template and the blended reference template (where the blended reference template can be generated by blending the template identified by the current motion vector and the template identified by the neighboring motion vector). c. For example, if the non-blended template cost is smaller, then OBMC may not be applied to the block. i. Alternatively, if the blended template cost is smaller, then OBMC may be applied to the block. 4.3. In one example, whether OBMC is applied to the current block can depend on the motion vector precision of the current block. 1) For example, the current block can be coded / decoded by inter-frame Merge. 2) For example, the current block can be coded / decoded by inter-frame AMVP. 3) For example, it can be based on whether the motion vector of the block is an integer (rather than a fractional) precision motion vector. 4) For example, it can be based on whether the motion vector difference of the block is an integer (rather than a fractional) precision motion vector difference. 4.4. Whether and / or how to apply the above-disclosed method can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, for example, in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.5. Whether and / or how to apply the above-disclosed method can be signaled at the PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel. 4.6. Whether and / or how the above-disclosed method can be applied may depend on the transcoded decoding information, such as block size, color format, single / double tree segmentation, color component, stripe / picture type.
[0099] The term "video unit" or "coding / decoding unit" or "block" may represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. In the present disclosure, regarding "a block coded in mode N", here "mode N" may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding / decoding technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-frame - inter-frame, GPM intra-frame - intra-frame, GPM inter-frame - intra-frame, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.).
[0100] Figure 39 A flowchart of a method 3900 for video processing according to an embodiment of the present disclosure is shown. Method 3900 is implemented during the conversion between a target video block of a video and the bitstream of the video.
[0101] At block 3910, for the conversion between a video unit of a video and the bitstream of the video unit, it is determined whether overlapping block motion compensation (OBMC) is applied to the current block of the video unit based on at least one of the following: the sample values of the samples within the current block, the sample values of the samples adjacent to the current block, the template cost, or the motion vector accuracy of the current block. In some embodiments, the current block is coded by inter-frame Merge. Alternatively, the current block is coded by inter-frame advanced motion vector prediction (AMVP).
[0102] At block 3920, the conversion is performed based on the determination. In some embodiments, the conversion may include encoding a video unit from the bitstream. Alternatively or additionally, the conversion may include decoding a video unit from the bitstream. In this way, block-level adaptive OBMC considering block characteristics based on the decoded information can bring higher coding / decoding gain and improve coding / decoding efficiency.
[0103] In some embodiments, whether OBMC is applied to the current block is based on a first non - hybrid template cost and a second hybrid template cost. In some embodiments, the first non - hybrid template cost is determined based on the sum of absolute differences (SAD) between the current template and a reference template. In some embodiments, the reference template is identified by adding the current motion vector to the position of the current template.
[0104] In some embodiments, the second hybrid template cost is determined based on the SAD between the current template and a hybrid reference template. In some embodiments, the hybrid reference template is generated by mixing the template identified by the current motion vector and the template identified by a neighboring motion vector.
[0105] In some embodiments, if the first non - hybrid template cost is smaller, OBMC is applied to the current block. In some embodiments, if the second hybrid template cost is smaller, OBMC is applied to the current block.
[0106] In some embodiments, whether OBMC is applied to the current block is based on the predicted samples in the current block before OBMC. In some embodiments, whether OBMC is applied to the current block is based on the predicted samples neighboring the current block. In some embodiments, whether OBMC is applied to the current block is based on the reconstructed samples neighboring the current block.
[0107] In some embodiments, whether OBMC is applied to the current block is based on at least one of the following: the gradient of the samples within the current block, the direction of the samples within the current block, the angle of the samples within the current block, the histogram of the gradients of the samples within the current block, the histogram of the directions of the samples within the current block, the histogram of the angles of the samples within the current block, the gradient of the samples neighboring the current block, the direction of the samples neighboring the current block, the angle of the samples neighboring the current block, the histogram of the gradients of the samples neighboring the current block, the histogram of the directions of the samples neighboring the current block, or the histogram of the angles of the samples neighboring the current block.
[0108] In some embodiments, the current predicted samples before OBMC are used to determine whether OBMC is applied to the current block. In some embodiments, the neighboring reconstructed samples are used to determine whether OBMC is applied to the current block.
[0109] In some embodiments, the histogram of at least one of the gradient, direction, or angle is determined by counting the gradients along a target direction or angle. In some embodiments, the target direction or angle is predefined. Alternatively, the target direction or angle is based on the direction of the intra - prediction angle mode in video coding and decoding.
[0110] In some embodiments, for a target direction or angle, the gradient magnitude is determined based on counting the gradients or magnitudes of gradients of at least one sample point in the current block. In some embodiments, the gradient or magnitude of the gradient at a target position (e.g., the center) in the current block is counted.
[0111] In some embodiments, the gradients or magnitudes of gradients at a series of target positions in the current block are counted. In some embodiments, the gradients or magnitudes of gradients of all sample points in the current block are counted. In some embodiments, the gradients or magnitudes of gradients of all sample points except for the sample points in the first row, the last row, the first column, and the last column in the current block are counted.
[0112] In some embodiments, for a target direction or angle, the gradient magnitude is determined based on counting the gradients or magnitudes of gradients of at least one sample point adjacent to the current block. In some embodiments, the gradient or magnitude of the gradient at a target position adjacent to the current block is counted. In some embodiments, the gradients or magnitudes of gradients of all sample points adjacent to the current block are counted. In some embodiments, the sample points are on the left side and / or the top of the current block.
[0113] In some embodiments, a histogram of at least one of gradients, directions, and angles is determined based on dividing the entire range of directions or angles into a series of intervals or bits. In some embodiments, a histogram of at least one of gradients, directions, and angles is determined based on counting the gradients or magnitudes of gradients in each interval or each bit or each direction or each angle.
[0114] In some embodiments, whether OBMC is applied to the current block is based on at least one of the following: the color of the sample points in the current block, the luminance of the sample points in the current block, the intensity of the sample points in the current block, the histogram of the color of the sample points in the current block, the histogram of the luminance of the sample points in the current block, the histogram of the intensity of the sample points in the current block, the color of the sample points adjacent to the current block, the luminance of the sample points adjacent to the current block, the intensity of the sample points adjacent to the current block, the histogram of the color of the sample points adjacent to the current block, the histogram of the luminance of the sample points adjacent to the current block, or the histogram of the intensity of the sample points adjacent to the current block.
[0115] In some embodiments, the current predicted sample points before OBMC are used to determine whether OBMC is applied to the current block. In some embodiments, the neighboring reconstructed sample points are used to determine whether OBMC is applied to the current block.
[0116] In some embodiments, the histogram of at least one of color, luminance, or intensity is determined based on counting sample values in at least one of the Y, U, or V component domains. Alternatively or additionally, the histogram of at least one of color, luminance, or intensity is determined based on counting sample values in at least one of the R, G, or B component domains.
[0117] In some embodiments, the sample values at a series of target positions in the current block are counted. In some embodiments, the sample values of all samples in the current block are counted.
[0118] In some embodiments, the sample values at target positions adjacent to the current block are counted. In some embodiments, the sample values of all samples adjacent to the current block are counted. In some embodiments, the samples are on the left side and / or top of the current block.
[0119] In some embodiments, the histogram of color is determined based on dividing the entire range of color into a series of intervals or bins. Alternatively or additionally, the histogram of luminance is determined based on dividing the entire range of luminance into a series of intervals or bins. Alternatively or additionally, the histogram of intensity is determined based on dividing the entire range of intensity into a series of intervals or bins. In some embodiments, the histogram of at least one of color, luminance, or intensity is determined based on counting the number of samples in each interval or bin.
[0120] In some embodiments, whether OBMC is applied to the current block is based on at least one of the following: the number of main gradients of samples in the current block, the number of main directions of samples in the current block, the number of main angles of samples in the current block, the number of main colors of samples in the current block, the number of main luminances of samples in the current block, the number of main intensities of samples in the current block, the number of main gradients of samples adjacent to the current block, the number of main directions of samples adjacent to the current block, the number of main angles of samples adjacent to the current block, the number of main colors of samples adjacent to the current block, the number of main luminances of samples adjacent to the current block, or the number of main intensities of samples adjacent to the current block.
[0121] In some embodiments, the number of at least one of main gradient, direction, angle, color, luminance, or intensity is determined based on predicted samples in the current block before OBMC. In some embodiments, the number of at least one of main gradient, direction, angle, color, luminance, or intensity is determined based on reconstructed samples adjacent to the current block.
[0122] In some embodiments, at least one of a major gradient, direction, angle, color, luminance, or intensity is derived based on a histogram of at least one of a gradient, direction, angle, color, luminance, or intensity. In some embodiments, at least one of a major gradient, direction, angle, color, luminance, or intensity is derived based on how many bins or bits in the histogram exhibit values greater than a threshold.
[0123] In some embodiments, at least one of a major gradient, direction, angle, color, luminance, or intensity is derived based on how many bins or bits in the histogram provide values (e.g., gradient magnitude, color value, luminance value) that are greater than values of other bins or bits.
[0124] In some embodiments, the values of the bins or bits in the histogram are sorted. In some embodiments, if X i >= a*(X i+1 ), then the bins or bits from 0 to i are considered major gradients, where a represents a scaling factor, and the sorted values are represented by X0, X1, X2 …, X n-2 , X n-1 , and n bins / bits are included in the histogram. In some embodiments, a is a constant between 2 and 20. For example, the values of the bins / bits in the histogram can be sorted first. Assume the sorted values (e.g., from largest to smallest) are represented by X0, X1, X2 …, X n-2 , X n-1 , where n bins / bits are included in the histogram. If X i >= a*(X i+1 ), then the bins / bits from 0 to i can be considered major gradients, where a represents a scaling factor (e.g., a can be a constant between 2 and 20).
[0125] In some embodiments, if the number of at least one of a major gradient, direction, angle, color, luminance, or intensity is less than a threshold number, OBMC is not applied to the current block. In some embodiments, the threshold number is one of the following: 1, 2, 3, 4, 5, 6, 7, 8, or 9.
[0126] In some embodiments, whether OBMC is applied to the current block is based on whether the motion vector of the current block is an integer-precision motion vector. In some embodiments, whether OBMC is applied to the current block is based on whether the motion vector difference of the current block is an integer-precision motion vector difference.
[0127] In some embodiments, an indication of whether and / or how it is determined whether OBMC is applied to a current block is indicated at at least one of the following: sequence level, picture group level, picture level, slice level, or slice group level. In some embodiments, an indication of whether and / or how it is determined whether OBMC is applied to a current block is indicated in at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header. In some embodiments, an indication of whether and / or how it is determined whether OBMC is applied to a current block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, tile, sub-picture, or region containing more than one sample or pixel.
[0128] In some embodiments, method 3900 further includes: determining whether and / or how it is determined whether OBMC is applied to a current block based on codec information of a video unit, the codec information including at least one of the following: block size, color format, single-tree and / or dual-tree segmentation, color component, slice type, or picture type.
[0129] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing of a video. The method includes: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: sample values of samples within the current block, sample values of samples adjacent to the current block, template cost, or motion vector precision of the current block; and generating a bitstream based on the determination.
[0130] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: sample values of samples within the current block, sample values of samples adjacent to the current block, template cost, or motion vector precision of the current block; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium.
[0131] Embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.
[0132] Item 1. A method for video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, determining whether overlapping block motion compensation (OBMC) based on at least one of the following is applied to a current block of the video unit: a sample value of a sample within the current block, a sample value of a sample adjacent to the current block, a template cost, or a motion vector precision of the current block; and performing the conversion based on the determination.
[0133] Item 2. The method according to Item 1, wherein the current block is inter-frame Merge coded or decoded, or the current block is inter-frame advanced motion vector prediction (AMVP) coded or decoded.
[0134] Item 3. The method according to Item 1 or 2, wherein whether OBMC is applied to the current block is based on a first non-blended template cost and a second blended template cost.
[0135] Item 4. The method according to Item 3, wherein the first non-blended template cost is determined based on a sum of absolute differences (SAD) between a current template and a reference template.
[0136] Item 5. The method according to Item 4, wherein the reference template is identified by adding a current motion vector to a position of the current template.
[0137] Item 6. The method according to Item 3, wherein the second blended template cost is determined based on the SAD between the current template and a blended reference template.
[0138] Item 7. The method according to Item 6, wherein the blended reference template is generated by blending a template identified by a current motion vector and a template identified by an adjacent motion vector.
[0139] Item 8. The method according to Item 3, wherein if the first non-blended template cost is smaller, the OBMC is applied to the current block.
[0140] Item 9. The method according to Item 3, wherein if the second blended template cost is smaller, the OBMC is applied to the current block.
[0141] Item 10. The method according to Item 1 or 2, wherein whether OBMC is applied to the current block is based on predicted samples in the current block before the OBMC.
[0142] Item 11. The method according to Item 1 or 2, wherein whether OBMC is applied to the current block is based on predicted samples adjacent to the current block.
[0143] Item 12. The method according to Item 1 or 2, wherein whether the OBMC is applied to the current block is based on reconstructed samples adjacent to the current block.
[0144] Item 13. The method according to Item 1 or 2, wherein whether the OBMC is applied to the current block is based on at least one of the following: the gradient of samples within the current block, the direction of samples within the current block, the angle of samples within the current block, the histogram of the gradient of samples within the current block, the histogram of the direction of samples within the current block, the histogram of the angle of samples within the current block, the gradient of samples adjacent to the current block, the direction of samples adjacent to the current block, the angle of samples adjacent to the current block, the histogram of the gradient of samples adjacent to the current block, the histogram of the direction of samples adjacent to the current block, or the histogram of the angle of samples adjacent to the current block.
[0145] Item 14. The method according to Item 13, wherein the current predicted sample before the OBMC is used to determine whether the OBMC is applied to the current block.
[0146] Item 15. The method according to Item 13, wherein adjacent reconstructed samples are used to determine whether the OBMC is applied to the current block.
[0147] Item 16. The method according to Item 13, wherein the histogram of at least one of the gradient, direction, or angle is determined based on counting the gradient along a target direction or angle.
[0148] Item 17. The method according to Item 16, wherein the target direction or angle is predefined, or wherein the target direction or angle is based on the direction of the intra prediction angle mode in video coding and decoding.
[0149] Item 18. The method according to Item 16, wherein for the target direction or angle, the gradient magnitude is determined based on counting the gradient or the magnitude of the gradient of at least one sample in the current block.
[0150] Item 19. The method according to Item 18, wherein the gradient or the magnitude of the gradient at a target position in the current block is counted.
[0151] Item 20. The method according to Item 18, wherein the gradient or the magnitude of the gradient at a series of target positions in the current block is counted.
[0152] Item 21. The method according to Item 18, wherein the gradient or the magnitude of the gradient of all samples in the current block is counted.
[0153] Item 22. The method according to item 18, wherein the gradients or the magnitudes of the gradients of all the sample points except the first row, the last row, the first column, and the last column sample points in the current block are counted.
[0154] Item 23. The method according to item 16, wherein for the target direction or angle, the gradient magnitude is determined based on counting the gradients or the magnitudes of the gradients of at least one sample point adjacent to the current block.
[0155] Item 24. The method according to item 23, wherein the gradients or the magnitudes of the gradients at the target positions adjacent to the current block are counted.
[0156] Item 25. The method according to item 23, wherein the gradients or the magnitudes of the gradients of all the sample points adjacent to the current block are counted.
[0157] Item 26. The method according to item 25, wherein the sample points are on the left side and / or the top of the current block.
[0158] Item 27. The method according to item 13, wherein the histogram of at least one of the gradients, directions, and angles is determined based on dividing the entire range of the direction or angle into a series of intervals or bits.
[0159] Item 28. The method according to item 13, wherein the histogram of at least one of the gradients, directions, and angles is determined based on counting the gradients or the magnitudes of the gradients in each interval or each bit or each direction or each angle.
[0160] Item 29. The method according to item 1 or 2, wherein whether the OBMC is applied to the current block is based on at least one of the following: the color of the sample points in the current block, the luminance of the sample points in the current block, the intensity of the sample points in the current block, the histogram of the color of the sample points in the current block, the histogram of the luminance of the sample points in the current block, the histogram of the intensity of the sample points in the current block, the color of the sample points adjacent to the current block, the luminance of the sample points adjacent to the current block, the intensity of the sample points adjacent to the current block, the histogram of the color of the sample points adjacent to the current block, the histogram of the luminance of the sample points adjacent to the current block, or the histogram of the intensity of the sample points adjacent to the current block.
[0161] Item 30. The method according to item 29, wherein the current predicted sample points before the OBMC are used to determine whether the OBMC is applied to the current block.
[0162] Item 31. The method according to Item 29, wherein neighboring reconstructed samples are used to determine whether the OBMC is applied to the current block.
[0163] Item 32. The method according to Item 29, wherein the histogram of at least one of the color, luminance, or intensity is determined based on counting the sample values in at least one of the Y, U, or V component domains, and / or wherein the histogram of at least one of the color, luminance, or intensity is determined based on counting the sample values in at least one of the R, G, or B component domains.
[0164] Item 33. The method according to Item 32, wherein the sample values at a series of target positions in the current block are counted.
[0165] Item 34. The method according to Item 32, wherein the sample values of all samples in the current block are counted.
[0166] Item 35. The method according to Item 32, wherein the sample values at target positions neighboring the current block are counted.
[0167] Item 36. The method according to Item 32, wherein the sample values of all samples neighboring the current block are counted.
[0168] Item 37. The method according to Item 36, wherein the samples are on the left side and / or the top of the current block.
[0169] Item 38. The method according to Item 29, wherein the histogram of the color is determined based on dividing the entire range of the color into a series of intervals or bits, and / or wherein the histogram of the luminance is determined based on dividing the entire range of the luminance into a series of intervals or bits, and / or wherein the histogram of the intensity is determined based on dividing the entire range of the intensity into a series of intervals or bits.
[0170] Item 39. The method according to Item 29, wherein the histogram of at least one of the color, luminance, or intensity is determined based on counting the number of samples in each interval or bit.
[0171] Item 40. The method according to Item 1 or 2, wherein whether the OBMC is applied to the current block is based on at least one of the following: the number of main gradients of the samples within the current block, the number of main directions of the samples within the current block, the number of main angles of the samples within the current block, the number of main colors of the samples within the current block, the number of main luminances of the samples within the current block, the number of main intensities of the samples within the current block, the number of main gradients of the samples adjacent to the current block, the number of main directions of the samples adjacent to the current block, the number of main angles of the samples adjacent to the current block, the number of main colors of the samples adjacent to the current block, the number of main luminances of the samples adjacent to the current block, or the number of main intensities of the samples adjacent to the current block.
[0172] Item 41. The method according to Item 40, wherein the number of at least one of the main gradient, direction, angle, color, luminance or intensity is determined based on the predicted samples within the current block before the OBMC.
[0173] Item 42. The method according to Item 40, wherein the number of at least one of the main gradient, direction, angle, color, luminance or intensity is determined based on the reconstructed samples adjacent to the current block.
[0174] Item 43. The method according to Item 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on a histogram of at least one of the gradient, direction, angle, color, luminance or intensity.
[0175] Item 44. The method according to Item 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on how many bins or bits in the histogram exhibit values greater than a threshold.
[0176] Item 45. The method according to Item 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on how many bins or bits in the histogram provide values greater than the values of other bins or bits.
[0177] Item 46. The method according to Item 45, wherein the values of the bins or bits in the histogram are sorted.
[0178] Item 47. The method according to Item 46, wherein if X i >= a * (X i+1 ), then the bins or bits from 0 to i are considered the main gradient, where a represents a scaling factor, and the sorted values are X0, X1, X2…, X n-2, X n-1 indicates that n intervals / bits are included in the histogram.
[0179] Item 48. The method according to Item 47, wherein a is a constant between 2 and 20.
[0180] Item 49. The method according to Item 40, wherein if the number of at least one of the main gradient, direction, angle, color, brightness, or intensity is less than a threshold number, the OBMC is not applied to the current block.
[0181] Item 50. The method according to Item 49, wherein the threshold number is one of the following: 1, 2, 3, 4, 5, 6, 7, 8, or 9.
[0182] Item 51. The method according to Item 1 or 2, wherein whether the OBMC is applied to the current block is based on whether the motion vector of the current block is an integer-precision motion vector.
[0183] Item 52. The method according to Item 1 or 2, wherein whether the OBMC is applied to the current block is based on whether the motion vector difference of the current block is an integer-precision motion vector difference.
[0184] Item 53. The method according to any one of Items 1 to 52, wherein an indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at at least one of the following: sequence level, picture group level, picture level, slice level, or slice group level.
[0185] Item 54. The method according to any one of Items 1 to 52, wherein an indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated in at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.
[0186] Item 55. The method according to any one of Items 1 to 52, wherein an indication of whether and / or how to determine whether the OBMC is applied to the current block is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, picture, sub-picture, or a region containing more than one sample or pixel.
[0187] Item 56. The method according to any one of Items 1 to 52 further includes: determining whether and / or how to determine whether the OBMC is applied to the current block based on the codec information of the video unit, where the codec information includes at least one of the following: block size, color format, single-tree and / or double-tree segmentation, color component, stripe type, or picture type.
[0188] Item 57. The method according to any one of Items 1 to 56, where the conversion includes encoding the video unit into the bitstream.
[0189] Item 58. The method according to any one of Items 1 to 56, where the conversion includes decoding the video unit from the bitstream.
[0190] Item 59. An apparatus for video processing includes a processor and a non-transitory memory having instructions thereon, where the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 58.
[0191] Item 60. A non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to any one of Items 1 to 58.
[0192] Item 61. A non-transitory computer-readable recording medium stores a bitstream generated by a method executed by an apparatus for video processing for a video, where the method includes: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: sample values of samples within the current block, sample values of samples adjacent to the current block, template cost, or motion vector accuracy of the current block; and generating the bitstream based on the determination.
[0193] Item 62. A method for storing a bitstream of a video includes: determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: sample values of samples within the current block, sample values of samples adjacent to the current block, template cost, or motion vector accuracy of the current block; generating the bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable medium. Example device
[0194] Figure 40FIG. shows a block diagram of a computing device 4000 in which various embodiments of the present disclosure may be implemented. The computing device 4000 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0195] It should be understood that Figure 40 the computing device 4000 shown in is for illustrative purposes only and does not imply any limitation to the functions and scope of the embodiments of the present disclosure in any way.
[0196] As Figure 40 shown, the computing device 4000 includes a general computing device 4000. The computing device 4000 may at least include one or more processors or processing units 4010, a memory 4020, a storage unit 4030, one or more communication units 4040, one or more input devices 4050, and one or more output devices 4060.
[0197] In some embodiments, the computing device 4000 may be implemented as any user terminal or server terminal with computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is conceivable that the computing device 4000 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0198] The processing unit 4010 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in the memory 4020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4000. The processing unit 4010 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0199] Computing device 4000 generally includes various computer storage media. Such media can be any media accessible by computing device 4000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. Memory 4020 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. Storage unit 4030 can be any removable or non-removable media and can include machine-readable media, such as a memory, flash drive, disk, or other media that can be used to store information and / or data and can be accessed in computing device 4000.
[0200] Computing device 4000 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 40 it, a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In this case, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0201] Communication unit 4040 communicates with another computing device via a communication medium. Additionally, the functions of the components in computing device 4000 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Thus, computing device 4000 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0202] Input device 4050 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4060 can be one or more of various output devices, such as a display, speaker, printer, etc. With the aid of communication unit 4040, computing device 4000 can also communicate with one or more external devices (not shown), such as storage devices and display devices, computing device 4000 can also communicate with one or more devices that enable a user to interact with computing device 4000, or if needed, computing device 4000 can also communicate with any device that enables computing device 4000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication can be carried out via an input / output (I / O) interface (not shown).
[0203] In some embodiments, some or all components of computing device 4000 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure may provide services through shared data centers, although to a user, they appear as a single access point. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.
[0204] In an embodiment of the present disclosure, computing device 4000 may be used to implement video encoding / decoding. Memory 4020 may include one or more video codec modules 4025 having one or more program instructions. These modules are accessible and executable by processing unit 4010 to perform the functions of the various embodiments described herein.
[0205] In an example embodiment of performing video encoding, input device 4050 may receive video data as input 4070 to be encoded. The video data may be processed, for example, by video codec module 4025 to generate an encoded bitstream. The encoded bitstream may be provided as output 4080 via output device 4060.
[0206] In an example embodiment of performing video decoding, input device 4050 may receive the encoded bitstream as input 4070. The encoded bitstream may be processed, for example, by video codec module 4025 to generate decoded video data. The decoded video data may be provided as output 4080 via output device 4060.
[0207] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between a video unit of a video and the bitstream of the video unit, determining whether overlapping block motion compensation (OBMC) based on at least one of the following is applied to a current block of the video unit: The sample values of the samples within the current block, The sample values of the samples adjacent to the current block, The template cost, or The motion vector precision of the current block; and Performing the conversion based on the determination.
2. The method according to claim 1, wherein the current block is inter-frame Merge encoded and decoded, or the current block is inter-frame advanced motion vector prediction (AMVP) encoded and decoded.
3. The method according to claim 1 or 2, wherein whether OBMC is applied to the current block is based on a first non-blended template cost and a second blended template cost.
4. The method according to claim 3, wherein the first non-blended template cost is determined based on the sum of absolute differences (SAD) between a current template and a reference template.
5. The method according to claim 4, wherein the reference template is identified by adding a current motion vector to the position of the current template.
6. The method according to claim 3, wherein the second blended template cost is determined based on the SAD between a current template and a blended reference template.
7. The method according to claim 6, wherein the blended reference template is generated by blending a template identified by a current motion vector and a template identified by an adjacent motion vector.
8. The method according to claim 3, wherein if the first non-blended template cost is smaller, the OBMC is applied to the current block.
9. The method according to claim 3, wherein if the second blended template cost is smaller, the OBMC is applied to the current block.
10. The method according to claim 1 or 2, wherein whether OBMC is applied to the current block is based on predicted samples in the current block before the OBMC.
11. The method according to claim 1 or 2, wherein whether OBMC is applied to the current block is based on predicted samples adjacent to the current block.
12. The method according to claim 1 or 2, wherein whether OBMC is applied to the current block is based on reconstructed samples adjacent to the current block.
13. The method according to claim 1 or 2, wherein whether OBMC is applied to the current block is based on at least one of the following: The gradient of the samples within the current block, The direction of the samples within the current block, The angle of the samples within the current block, The histogram of the gradient of the samples within the current block, The histogram of the direction of the samples within the current block, The histogram of the angle of the samples within the current block, The gradient of the samples adjacent to the current block, The direction of the samples adjacent to the current block, The angle of the samples adjacent to the current block, The histogram of the gradient of the samples adjacent to the current block, The histogram of the direction of the samples adjacent to the current block, or The histogram of the angle of the samples adjacent to the current block.
14. The method according to claim 13, wherein a current predicted sample point before the OBMC is used to determine whether the OBMC is applied to the current block.
15. The method according to claim 13, wherein neighboring reconstructed sample points are used to determine whether the OBMC is applied to the current block.
16. The method according to claim 13, wherein a histogram of at least one of the gradient, direction, or angle is determined based on counting gradients along a target direction or angle.
17. The method according to claim 16, wherein the target direction or angle is predefined, or wherein the target direction or angle is based on the direction of an intra prediction angle mode in video coding and decoding.
18. The method according to claim 16, wherein for a target direction or angle, the gradient magnitude is determined based on counting the gradient or the magnitude of the gradient for at least one sample point in the current block.
19. The method according to claim 18, wherein the gradient or the magnitude of the gradient at a target position in the current block is counted.
20. The method according to claim 18, wherein the gradient or the magnitude of the gradient at a series of target positions in the current block is counted.
21. The method according to claim 18, wherein the gradient or the magnitude of the gradient for all sample points in the current block is counted.
22. The method according to claim 18, wherein the gradient or the magnitude of the gradient for all sample points except for the sample points in the first row, last row, first column, and last column in the current block is counted.
23. The method according to claim 16, wherein for a target direction or angle, the gradient magnitude is determined based on counting the gradient or the magnitude of the gradient for at least one sample point adjacent to the current block.
24. The method according to claim 23, wherein the gradient or the magnitude of the gradient at a target position adjacent to the current block is counted.
25. The method according to claim 23, wherein the gradient or the magnitude of the gradient for all sample points adjacent to the current block is counted.
26. The method according to claim 25, wherein the sample points are on the left side and / or top of the current block.
27. The method according to claim 13, wherein a histogram of at least one of the gradient, direction, angle is determined based on dividing the entire range of the direction or angle into a series of intervals or bits.
28. The method according to claim 13, wherein a histogram of at least one of the gradient, direction, angle is determined based on counting the gradient or the magnitude of the gradient in each interval or each bit or each direction or each angle.
29. The method according to claim 1 or 2, wherein whether the OBMC is applied to the current block is based on at least one of the following: the color of the sample points within the current block, the luminance of the sample points within the current block, the intensity of the sample points within the current block, the histogram of the color of the sample points within the current block, the histogram of the luminance of the sample points within the current block, The histogram of the intensities of the samples within the current block The color of the samples adjacent to the current block The luminance of the samples adjacent to the current block The intensity of the samples adjacent to the current block The histogram of the color of the samples adjacent to the current block The histogram of the luminance of the samples adjacent to the current block, or The histogram of the intensity of the samples adjacent to the current block 30. The method according to claim 29, wherein the current predicted sample before the OBMC is used to determine whether the OBMC is applied to the current block.
31. The method according to claim 29, wherein the neighboring reconstructed samples are used to determine whether the OBMC is applied to the current block.
32. The method according to claim 29, wherein the histogram of at least one of the color, luminance or intensity is determined based on counting the sample values in at least one of the Y, U or V component domains, and / or wherein the histogram of at least one of the color, luminance or intensity is determined based on counting the sample values in at least one of the R, G or B component domains.
33. The method according to claim 32, wherein the sample values at a series of target positions in the current block are counted.
34. The method according to claim 32, wherein the sample values of all the samples in the current block are counted.
35. The method according to claim 32, wherein the sample values at the target positions adjacent to the current block are counted.
36. The method according to claim 32, wherein the sample values of all the samples adjacent to the current block are counted.
37. The method according to claim 36, wherein the samples are on the left side and / or the top of the current block.
38. The method according to claim 29, wherein the histogram of the color is determined based on dividing the entire range of the color into a series of intervals or bits, and / or wherein the histogram of the luminance is determined based on dividing the entire range of the luminance into a series of intervals or bits, and / or wherein the histogram of the intensity is determined based on dividing the entire range of the intensity into a series of intervals or bits.
39. The method according to claim 29, wherein the histogram of at least one of the color, luminance or intensity is determined based on counting the number of samples in each interval or bit.
40. The method according to claim 1 or 2, wherein whether the OBMC is applied to the current block is based on at least one of the following: The number of main gradients of the samples within the current block The number of main directions of the samples within the current block The number of main angles of the samples within the current block The number of main colors of the samples within the current block The number of main luminances of the samples within the current block The number of main intensities of the samples within the current block The number of main gradients of the samples adjacent to the current block The number of main directions of the samples adjacent to the current block The number of main angles of the samples adjacent to the current block The number of main colors of the samples adjacent to the current block The number of main luminances of the samples adjacent to the current block, or the number of main intensities of the samples adjacent to the current block.
41. The method according to claim 40, wherein the number of at least one of the main gradient, direction, angle, color, luminance or intensity is determined based on the predicted samples within the current block before the OBMC.
42. The method according to claim 40, wherein the number of at least one of the main gradient, direction, angle, color, luminance or intensity is determined based on the reconstructed samples adjacent to the current block.
43. The method according to claim 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on a histogram of at least one of the gradient, direction, angle, color, luminance or intensity.
44. The method according to claim 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on how many bins or bits in the histogram exhibit values greater than a threshold.
45. The method according to claim 40, wherein at least one of the main gradient, direction, angle, color, luminance or intensity is derived based on how many bins or bits in the histogram provide values greater than the values of other bins or bits.
46. The method according to claim 45, wherein the values of the bins or bits in the histogram are sorted.
47. The method according to claim 46, wherein if X i >= a * (X i+1 ), then the interval or binary bit from 0 to i is regarded as the main gradient, where a represents a scaling factor, and the sorted values are represented by X0, X1, X2…, X n-2 , X n-1 , and n intervals / binary bits are included in the histogram.
48. The method according to claim 47, wherein a is a constant between 2 and 20.
49. The method according to claim 40, wherein if the number of at least one of the main gradient, direction, angle, color, luminance or intensity is less than a threshold number, the OBMC is not applied to the current block.
50. The method according to claim 49, wherein the threshold number is one of the following: 1, 2, 3, 4, 5, 6, 7, 8 or 9.
51. The method according to claim 1 or 2, wherein whether the OBMC is applied to the current block is based on whether the motion vector of the current block is an integer-precision motion vector.
52. The method according to claim 1 or 2, wherein whether the OBMC is applied to the current block is based on whether the motion vector difference of the current block is an integer-precision motion vector difference.
53. The method according to any one of claims 1 to 52, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated at least at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.
54. The method according to any one of claims 1 to 52, wherein the indication of whether and / or how to determine whether the OBMC is applied to the current block is indicated in at least one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or Slice group header.
55. The method according to any one of claims 1 to 52, wherein the indication of whether and / or how it is determined whether the OBMC is applied to the current block is included in one of the following: Prediction block (PB), Transform block (TB), Codec block (CB), Prediction unit (PU), Transform unit (TU), Codec unit (CU), Virtual pipeline data unit (VPDU), Codec tree unit (CTU), CTU row, Strip, Slice, Sub-picture, or Region containing more than one sample or pixel.
56. The method according to any one of claims 1 to 52, further comprising: Determining whether and / or how it is determined whether the OBMC is applied to the current block based on the codec information of the video unit, the codec information including at least one of the following: Block size, Color format, Single-tree and / or dual-tree segmentation, Color component, Strip type, or Picture type.
57. The method according to any one of claims 1 to 56, wherein the transformation includes encoding the video unit into the bitstream.
58. The method according to any one of claims 1 to 56, wherein the transformation includes decoding the video unit from the bitstream.
59. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 58.
60. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 58.
61. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method includes: Determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: Sample values of samples within the current block, Sample values of samples adjacent to the current block, Template cost, or Motion vector precision of the current block; and Generating the bitstream based on the determination.
62. A method for storing a bitstream of a video, comprising: Determining whether motion compensation based on overlapping sub-blocks (OBMC) is applied to a current block of a video unit of the video based on at least one of the following: Sample values of samples within the current block, Sample values of samples adjacent to the current block, Template cost, or Motion vector precision of the current block; Generating the bitstream based on the determination; and Storing the bitstream in a non-transitory computer-readable medium.