Method and device for video processing and medium
By applying the codec information of the second block and limiting its position in the conversion between the video unit and the bitstream of the video unit, and performing block-level adaptive OBMC in combination with the codec information or predefined rules, the problem of improving codec efficiency in the prior art is solved, and higher codec gain and performance improvements are achieved.
Patent Information
- Application Number
- CN202380090365.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2023-12-28
- Publication Date
- 2025-08-08
AI Technical Summary
The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in block-level adaptive overlap subblock motion compensation (OBMC), which is difficult to further improve the encoding and decoding gain and decoding performance.
In the conversion between the video unit and the bitstream of the video unit, the codec information of the second block is applied, and its position is restricted based on predefined rules, and block-level adaptive OBMC is performed based on decoded information, and mode decision is determined in combination with the codec information or predefined rules, and the conversion is performed to consider block characteristics.
Improves the codec gain and codec efficiency, and improves the codec performance.
Smart Images

Figure CN120457676A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to block-level adaptive overlapped sub-block motion compensation (OBMC) in video coding. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. Various video compression technologies have been proposed for video encoding and decoding, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, further improvements in the encoding and decoding efficiency of video encoding and decoding technologies are often desired. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: applying codec information of a second block to a codec process of a current block associated with the video unit, wherein the position of the second block is restricted based on predefined rules, for conversion between a video unit and a bitstream of the video unit; and performing the conversion based on the codec process. In this way, block-level adaptive OBMC, which considers block characteristics based on decoded information, can achieve higher codec gains. Furthermore, it can improve codec efficiency and performance.
[0005] In a second aspect, another method for video processing is proposed. The method includes: determining a set of prediction samples used for mode decision for a video unit based on at least one of codec information or a predefined rule for conversion between a video unit and a bitstream of the video unit; and performing conversion based on the set of prediction samples. In this way, block-level adaptive OBMC, which considers block characteristics based on decoded information, can achieve higher codec gains. Furthermore, it can improve codec efficiency and performance.
[0006] In a third aspect, a device for video processing is provided. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect or the second aspect of the present disclosure.
[0007] In a fourth aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions for causing a processor to execute the method according to the first aspect or the second aspect of the present disclosure.
[0008] In a fifth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; and generating a bitstream based on the codec process.
[0009] In a sixth aspect, a method for storing a bitstream of a video is provided. The method includes: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; generating a bitstream based on the codec process; and storing the bitstream in a non-transitory computer-readable medium.
[0010] In a seventh aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining, for conversion between a video unit and a bitstream of the video unit, a set of prediction samples used for mode decision for the video unit based on at least one of: codec information or a predefined rule; and generating a bitstream based on the set of prediction samples.
[0011] In an eighth aspect, a method for storing a video bitstream is provided. The method includes: determining, for conversion between a video unit of the video and a bitstream of the video unit, a set of prediction samples used for mode decision of the video unit based on at least one of: codec information or a predefined rule; generating a bitstream based on the set of prediction samples; and storing the bitstream in a non-transitory computer-readable medium.
[0012] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0014] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0015] Figure 2shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0016] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0017] Figure 4 Intra prediction mode is shown;
[0018] Figure 5 The reference samples used for wide-angle intra prediction are shown;
[0019] Figure 6 The problem of discontinuity is shown in the case of orientations exceeding 45°;
[0020] Figure 7A Schematic diagram showing the definition of samples used by PDPC for diagonal top right mode applied to diagonal and adjacent angle intra modes;
[0021] Figure 7B Schematic diagram showing the definition of samples used by PDPC for diagonal bottom left mode applied to diagonal and adjacent angle intra modes;
[0022] Figure 7C Schematic diagram showing the definition of samples used by PDPC for adjacent diagonal top right mode applied to diagonal and adjacent angle intra modes;
[0023] Figure 7D Schematic diagram showing the definition of samples used by PDPC for adjacent diagonal bottom left mode applied to diagonal and adjacent angle intra modes;
[0024] Figure 8 An example of four reference rows adjacent to a prediction block is shown;
[0025] Figure 9A A schematic diagram showing the process of sub-segmentation depending on the block size;
[0026] Figure 9B A schematic diagram showing the process of sub-segmentation depending on the block size;
[0027] Figure 10 shows the matrix-weighted intra prediction process;
[0028] Figure 11 The spatial GPM candidates are shown;
[0029] Figure 12 A GPM template is shown;
[0030] Figure 13 GPM mixing is shown;
[0031] Figure 14 Shows the location of the spatial merge candidate;
[0032] Figure 15 The candidate pairs considered for redundancy check of spatial Merge candidates are shown;
[0033] Figure 16 A diagram showing motion vector scaling for temporal Merge candidates is shown;
[0034] Figure 17 The candidate positions for the time domain Merge candidate, C0 and C1, are shown;
[0035] Figure 18 The MMVD search point is shown;
[0036] Figure 19 shows the extended CU area used in BDOF;
[0037] Figure 20 A diagram for a symmetric MVD pattern is shown;
[0038] Figure 21 Decoding side motion vector refinement is shown;
[0039] Figure 22 Shown are the top and left neighboring blocks used in the derivation of CIIP weights;
[0040] Figure 23 An example of GPM partitioning grouped at the same angle is shown;
[0041] Figure 24 Unidirectional prediction MV selection for geometric partitioning mode is shown;
[0042] Figure 25 shows an exemplary generation of warp weights w0 using a geometric segmentation pattern;
[0043] Figure 26 The current CTU processing order and its available reference samples in the current CTU and the left CTU are shown;
[0044] Figure 27 The residual encoding and decoding passes for a transform skip block are shown;
[0045] Figure 28 An example of a block encoded and decoded in palette mode is shown;
[0046] Figure 29 shows sub-block based index map scanning for a palette, left for horizontal scanning and right for vertical scanning;
[0047] Figure 30Shown is a decoding flow chart with ACT;
[0048] Figure 31 The intra-frame template matching search area used is shown;
[0049] Figure 32 Five locations in the reconstructed luminance samples are shown;
[0050] Figure 33 The prediction process of DBV mode is shown;
[0051] Figure 34 The low-frequency non-separable transform (LFNST) process is shown;
[0052] Figure 35 The SBT position, type and transformation type are shown;
[0053] Figure 36 The ROI for LFNST16 is shown;
[0054] Figure 37 The ROI for LFNST8 is shown;
[0055] Figure 38 Discontinuity measurements are shown;
[0056] Figure 39 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0057] Figure 40 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown;
[0058] Figure 41 A block diagram is shown of a computing device in which various embodiments of the present disclosure may be implemented.
[0059] Throughout the drawings, same or similar reference numbers generally refer to same or similar elements. DETAILED DESCRIPTION
[0060] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0061] In the following description and claims, unless defined otherwise, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0062] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, whether or not explicitly described, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in relation to other embodiments.
[0063] It should be understood that although the terms "first" and "second" and the like can be used to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0064] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises," "includes," and / or "having," when used herein, indicate the presence of the described features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0065] Figure 1 is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0066] Video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0067] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a coded picture and associated data. The coded picture is a coded representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be transmitted directly to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0068] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0069] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0070] Figure 2 is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0071] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0072] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0073] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, wherein at least one reference picture is a picture in which the current video block is located.
[0074] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0075] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0076] The mode selection unit 203 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0077] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0078] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are not dependent on macroblocks in the same picture.
[0079] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0080] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0081] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0082] In one example, motion estimation unit 204 may indicate to video decoder 300 a value in a syntax structure associated with the current video block that indicates the current video block has the same motion information as another video block.
[0083] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0084] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0085] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0086] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0087] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0088] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0089] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0090] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0091] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.
[0092] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0093] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0094] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0095] exist Figure 3 In the example of FIG, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to video encoder 200.
[0096] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data and reference images of adjacent PBs. The motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from adjacent blocks in the spatial or temporal domain.
[0097] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.
[0098] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to generate a prediction block.
[0099] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the coded video sequence, partition information describing how each macroblock of the picture of the coded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0100] The intra prediction unit 303 can form a prediction block from spatially adjacent blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0101] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0102] Some exemplary embodiments of the present disclosure are described in detail below. It should be understood that the section titles used in this document are for ease of understanding and do not limit the embodiments disclosed in the section to only that section. In addition, although certain embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding and decoding steps in detail, it is understood that the corresponding decoding steps will be implemented by a decoder, which de-encodes. In addition, the term "video processing" includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding technology. Specifically, it relates to overlapping sub-block motion compensation (OBMC) and related technologies in image / video coding. It can be applied to existing video coding standards such as HEVC and VVC. It can also be applied to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 Video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held concurrently every quarter. The new video codec standard was officially named the Versatile Video Codec (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. The VVC working draft and test model (VTM) are updated after each meeting. The VVC project achieved technical completion (FDIS) at the July 2020 meeting. 2.1. Existing Codec Tools Intra-frame prediction 2.1.1.1. Intra-mode codec with 67 intra-prediction modes To capture arbitrary edge directions present in natural videos, the number of directional intra modes in VVC is extended to 65 from the 33 used in HEVC. Figure 4 The red dashed arrows in the figure depict new directional modes that are not present in HEVC, and the planar and DC modes remain unchanged. These more dense directional intra prediction modes are applicable to all block sizes and for luma and chroma intra prediction. In VVC, for non-square blocks, multiple normal-angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid division operations for DC prediction, only the longer sides are used to calculate the average value of non-square blocks. 2.1.1.2. Intra-mode encoding and decoding In order to keep the complexity of the most probable mode (MPM) list generation low, an intra mode codec with 6 MPMs is used by considering two available adjacent intra modes. The following three aspects are considered when constructing the MPM list: – Default intra mode; – Neighborhood intra mode; – Derived intra mode. A unified 6-MPM list is used for intra blocks, regardless of whether the MRL and ISP codecs are applied. The MPM list is constructed based on the intra modes of the neighboring blocks on the left and above. Assuming the mode on the left is denoted as Left and the mode of the block above is denoted as Above, the unified MPM list is constructed as follows: – When a neighboring block is not available, its intra mode is set to planar by default. – If both Left and Above modes are non-angle modes: – MPM list → {plane, DC, V, H, V-4, V+4}. – If one of the modes Left and Above is an angular mode and the other is a non-angular mode: – Set mode Max to the larger of Left and Above – MPM list → {Planar, Max, DC, Max-1, Max+1, Max-2}. – If Left and Above are both angle modes and they are different: – Set mode Max to the larger of Left and Above – If the difference between the patterns Left and Above is in the range of 2 to 62 (inclusive) – MPM list → {Plane, Left, Above, DC, Max-1, Max+1} -otherwise – MPM list → {Flat, Left, Above, DC, Max-2, Max+2}. – If Left and Above are both angle modes and they are the same: – MPM list → {plane, left, left-1, left+1, DC, left-2}. In addition, the first binary bit of the mpm index codeword is CABAC context coded. A total of three contexts are used, corresponding to whether the current intra block is MRL-enabled, ISP-enabled, or a normal intra block. During the 6MPM list generation process, deduplication is used to remove repeated patterns so that only unique patterns can be included into the MPM list.For entropy coding of the 61 non-MPM modes, truncated binary code (TBC) is used. 2.1.1.3. Wide-angle intra prediction for non-square blocks Normal-angle intra prediction directions are defined as running clockwise from 45 degrees to -135 degrees. In VVC, multiple normal-angle intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original mode index, which is remapped to the wide-angle mode index after parsing. The total number of intra prediction modes remains unchanged at 67, and the intra mode encoding and decoding methods remain unchanged. To support these prediction directions, define a top reference of length 2W+1 and a left reference of length 2H+1, such as Figure 5 shown. The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra-frame prediction modes are shown in Table 1. Table 1 – Intra prediction modes replaced by Wide mode like Figure 6 As shown in Figure 2, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap Δp. αnegative impact. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted using these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fractional modes in the normal prediction mode and the wide-angle mode. In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the mapping table entries to more accurately convert the prediction angles for chroma blocks. 2.1.1.4. Mode Dependent Intra-Frame Smoothing (MDIS) A four-tap intra interpolation filter is used to improve the accuracy of directional intra prediction. In HEVC, a two-tap linear interpolation filter has been used to generate intra prediction blocks in directional prediction modes (i.e., excluding planar and DC prediction values). In VVC, a simplified 6-bit 4-tap Gaussian interpolation filter is used only for directional intra modes. The non-directional intra prediction process is not modified. The 4-tap filter is selected according to the MDIS condition for directional intra prediction modes providing non-fractional displacement, i.e., all directional modes except the following modes: 2, HOR_IDX, DIA_IDX, VER_IDX, 66. Depending on the intra prediction mode, the following reference sample processing is performed: – Directional intra prediction modes are classified into one of the following groups: – vertical or horizontal mode (HOR_IDX, VER_IDX), – diagonal mode (2, DIA_IDX, VDIA_IDX) that represents angles that are multiples of 45 degrees, – remaining directional patterns; – If the directional intra prediction mode is classified as belonging to group A, no filter is applied to the reference samples to generate the prediction samples; – Otherwise, if the mode falls into group B, the [1,2,1] reference sample filter may be applied (according to the MDIS condition) to the reference samples to further copy these filtered values into the intra prediction values according to the selected direction, However, no interpolation filter is applied; Otherwise, if the mode is classified as belonging to group C, only the intra reference sample interpolation filter is applied to the reference samples to generate prediction samples that fall between the reference samples at fractional or integer positions according to the selected direction (no reference sample filtering is performed). 2.1.1.5. Position-dependent intra prediction combination In VVC, the results of intra prediction for DC, planar, and multiple angular modes are further modified by the Position Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that uses a combination of unfiltered boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, horizontal, vertical, bottom-left angular mode and its eight adjacent angular modes, and top-right angular mode and its eight adjacent angular modes. The prediction sample pred(x',y') is predicted using a linear combination of the intra prediction mode (DC, planar, angular) and the reference samples according to the following equation 3-8: pred(x',y')=(wL×R -1,y '+wT×R x’,-1 -wTL×R -1,-1 +(64-wL-wT+wTL)×pred(x',y')+32)>>6(2-1) where R x,-1 、R -1,y Respectively represent the reference sample points at the top and left boundaries of the current sample point (x, y), and R -1,-1 Indicates the reference sample located at the upper left corner of the current block. If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filters are required, as required in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is the same for DC and planar modes, and the clipping operation is avoided. For angular modes, the PDPC scaling factor is adjusted so that no range check is required, and the angle condition for enabling PDPC is removed (using scaling >= 0). In addition, in all angular mode cases, the PDPC weights are based on 32. The PDPC weights depend on the prediction mode, as shown in Table 2. PDPC is applied to blocks with a width and height both greater than or equal to 4. Figures 7A-7D The reference sample points (R x,-1 、R -1,y and R -1,-1 ). The prediction sample pred(x',y') is located at (x',y') in the prediction block. For example, for the diagonal mode, the reference sample Rx,-1 The coordinate x of is given by the following formula: x=x'+y'+1, and the reference point R -1,y The coordinate y of is similarly given by the following formula: y = x' + y' + 1. For other annular patterns, the reference point R x,-1 and R -1,y Can be at a fractional sample position. In this case, the sample value at the nearest integer sample position is used. Table 2 - Example of PDPC weights according to prediction mode 2.1.1.6. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Figure 8 In
[15] , an example of 4 reference lines is depicted, where the samples of segments A and F are not obtained from reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines are used (reference line 1 and reference line 3). The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row idx greater than 0, only additional reference row modes are included in the MPM list, and only the mpm index is signaled, while the remaining modes are not signaled. The reference row index is signaled before the intra prediction mode, and if a non-zero reference row index is signaled, planar mode is excluded from the intra prediction mode. For the first row of blocks inside a CTU, MRL is disabled to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in the DC intra prediction mode for non-zero reference row indices is aligned with the derivation of reference row index 0. MRL requires 3 adjacent luminance reference rows to be stored with the CTU to generate the prediction. The Cross-Component Linear Model (CCLM) tool also requires 3 adjacent luminance reference rows for its downsampling filter. Using the same 3-row MLR definition is aligned with CCLM to reduce the storage requirements for the decoder. 2.1.1.7. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size of ISP is 4x8 (or 8x4). If the block size is larger than 4x8 (or 8x4), the corresponding block is divided into 4 sub-partitions. It has been noted that M×128 (with M≤64) and 128×N (with N≤64) ISP blocks may cause potential problems with 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luma TB will be divided into four M×32TBs (only horizontal division is possible), each TB is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not divided. Therefore, both chroma components will have a size larger than a 32×32 block. Similarly, using ISP with a 128×N CU can cause a similar situation. Therefore, these two situations are problems for a 64×64 decoder pipeline. For this reason, the CU size that can use ISP is limited to a maximum of 64×64. Figure 9A and Figure 9B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples. In ISP, 1xN / 2xN sub-block predictions are not allowed to rely on the reconstructed values of previously decoded 1xN / 2xN sub-blocks of the codec block, resulting in a minimum prediction width of four samples for each sub-block. For example, an 8xN (N>4) codec block encoded using ISP with vertical partitioning is partitioned into two prediction regions, each of size 4xN, and four transforms of size 2xN. Furthermore, a 4xN codec block encoded using ISP with vertical partitioning is predicted using a full 4xN block; four 1xN transforms are used. While both 1xN and 2xN transform sizes are allowed, it is specified that transforms for these blocks within a 4xN region can be performed in parallel. For example, when a 4xN prediction region contains four 1xN transforms, no transform is performed horizontally; the transform in the vertical direction can be performed as a single 4xN transform in the vertical direction. Similarly, when a 4xN prediction region contains two 2xN transform blocks, the transform operations for the two 2xN blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, processing these smaller blocks does not increase latency compared to processing intra blocks in a 4x4 regular codec. Table 3 – Entropy codec coefficient group sizes Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-partition, a reconstructed sample is obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization and inverse transformation. Therefore, the reconstructed sample value of each sub-partition can be used to generate a prediction for the next sub-partition, and each sub-partition is processed repeatedly. In addition, the first sub-partition to be processed is the one that contains the upper left sample of the CU, and then continues downward (horizontal partitioning) or to the right (vertical partitioning). As a result, the reference samples used to generate the sub-partition prediction signal are only located to the left and above the row. All sub-partitions share the same intra mode. The following is a summary of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: The size of the entropy coding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 3. Note that the new size only affects blocks generated by ISP where one dimension is less than 4 samples. In all other cases, the coefficient group remains 4×4 in dimension. –CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions and the first n-1 subpartitions produce zero CBF, the CBF of the nth subpartition is assumed to be 1. – MPM usage: In blocks encoded and decoded in ISP mode, the MPM flag will be inferred to be 1, and the MPM list is modified to exclude DC mode and prioritize horizontal intra mode for ISP horizontal partitioning and vertical intra mode for ISP vertical partitioning. – Transform size restriction: All ISP transforms with length greater than 16 points use DCT-II. –PDPC: When the CU uses ISP codec mode, the PDPC filter will not be applied to the generated sub-splits. -MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests for the different available transforms for each generated sub-split. Instead, the transform selection for ISP mode will be fixed and selected according to the utilized intra mode, processing order and block size. Therefore, no signaling is required. For example, making t H and t V are the horizontal and vertical transforms selected for the w×h sub-partition, respectively, where w is the width and h is the height. The transforms are selected according to the following rules: If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w=2 or w>32, t H =DCT-II – If h = 2 or h > 32, t V =DCT-II – Otherwise, select the transformation as shown in Table 4. Table 4 – Transform selection depending on intra mode In ISP mode, all 67 intra modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the conditions for intra interpolation filter selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.1.1.8. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a newly added intra prediction technique in VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, they are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation, as shown in Figure 10 shown. ●Average of neighboring points Among the boundary samples, four samples or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging the adjacent boundary samples according to a predefined rule depending on the block size. top and bdry left Reduced to a smaller boundary and Then, the two reduced boundaries and is spliced to the reduced boundary vector bdry red , so for blocks of shape 4×4 its size is 4, and for blocks of all other shapes its size is 8. If mode refers to a MIP mode, the splicing is defined as follows: Matrix multiplication Take the averaged samples as input, perform matrix-vector multiplication, and then add the offset. The result is a reduced prediction signal on the downsampled set of samples in the original block. From the reduced input vector bdry red Generate a reduced prediction signal pred red, , which has a width of W red and height Hred Here, W red and H red is defined as: Reduced prediction signal pred red It is calculated by taking the matrix-vector product and adding the offset: pred red =A·bdry red +b. Here, A is a matrix with W red ·H red rows, and if W=H=4, then it has 4 columns, and in all other cases has 8 columns. b is the size of W red ·H red The matrix A and the offset vector b are taken from one of the sets S0, S1, and S2. The index idx = idx(W,H) is defined as follows: Here, each coefficient of matrix A is represented with 8-bit precision. Set S0 consists of 16 matrices Each matrix has 16 rows and 4 columns, and 16 offset vectors Each offset vector has a size of 16. The set of matrices and offset vectors is used for blocks of size 4×4. Set S1 consists of 8 matrices Each matrix has 16 rows and 8 columns, and 8 offset vectors The size of each offset vector is 16. Set S2 consists of 6 matrices Each matrix has 64 rows and 8 columns, and 6 offset vectors of size 64 Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the downsampled set by linear interpolation, which is a single-step linear interpolation in each direction. Regardless of the block shape or block size, the interpolation is performed first in the horizontal direction and then in the vertical direction. ●MIP mode signaling and coordination with other codec tools For each codec unit (CU) in intra mode, a flag is sent indicating whether the MIP mode is to be applied. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transposed flag (isTransposed) that determines whether the mode is transposed and the MIP mode ID (modeId) that determines which matrix to use for a given MIP mode are derived as follows. isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-6) The MIP codec mode is coordinated with other codec tools by taking into account the following aspects: – Enable LFNST for MIPs on large blocks. Here, the planar LFNST transform is used. - Reference sample derivation for MIP is performed in exactly the same way as for regular intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled reference samples. – Clipping is performed before upsampling, rather than after upsampling. – Regardless of the maximum transform size, MIPs are allowed to be at most 64x64. - For sizeId=0, the number of MIP modes is 32, for sizeId=1, the number of MIP modes is 16, and for sizeId=2, the number of MIP modes is 12. 2.1.1.9. Airspace GPM (SGPM) In the spatial domain GPM, a candidate list including partitioning and two intra prediction modes is constructed. An MPM of up to 11 intra prediction modes is used to form a combination, and the length of the candidate list is set to be equal to 16. The selected candidate index is transmitted through the signal. use Figure 11 The templates shown in
[14] reorder the list. The GPM blending process is not used in the templates and the SAD between the prediction and reconstruction of the templates is used for the ranking. Figure 12 A GPM template is shown. Figure 13 GPM partition boundaries are shown. SGPM mode is applied to blocks whose width and height meet the same restrictions as in inter-frame GPM. Consider the following projects: ●Airspace GPM segmentation mode: 26 predefined modes Adaptive inference algorithm based on the ratio of horizontal gradient to vertical gradient ●Intra-frame prediction mode selection: List of IPMs with and without TIMD: For each segmentation mode, an IPM list is derived for each part using intra-inter GPM list derivation. The IPM list size is 3. In the list, TIMD-derived patterns are replaced by 2 derived patterns with horizontal and vertical orientations (using top or left template), or TIMD-derived patterns are excluded. MPM List: A unified MPM list (maximum 11 elements) is used for all segmentation modes. Template size (left and top): 1 or 4 ●Extended block size: Spatial GPM is extended to be further applied to 4x8, 8x4, 4x16 and 16x4 blocks, which can be described as 4<=width<=64, 4<=height<=64, width<height*8, height<width*8, width*height>=32. Adaptive Hybrid: Adaptive mixing is tested for spatial GPM, where the mixing depth τ is derived as follows: ■If min(width, height) == 4, then choose 1 / 2τ ■ Otherwise, if min(width, height) == 8, then select τ ■ Otherwise, if min(width, height) == 16, then choose 2τ ■ Otherwise, if min(width, height) == 32, then choose 4τ ■Otherwise, select 8τ. 2.1.2. Inter-frame prediction For each inter-predicted CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list usage index, as well as additional information for inter-prediction sample generation required by the new codec features of VVC. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector increments or reference picture indices. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, and from the additional scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not just for skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list usage flag for each reference picture list, as well as other required information, are explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC also includes many new and refined inter-frame prediction coding tools listed below: – Extended Merge prediction –Merge mode with MVD (MMVD) – Symmetrical MVD (SMVD) signaling – Affine motion compensated prediction – Sub-block based temporal motion vector prediction (SbTMVP) – Adaptive Motion Vector Resolution (AMVR) – Motion field storage: 1 / 16 luminance sample MV storage and 8x8 motion field compression – Bidirectional prediction with CU-level weights (BCW) – Bidirectional Optical Flow (BDOF) –Decoder-side motion vector refinement (DMVR) – Geometric Partitioning Mode (GPM) – Joint inter-frame and intra-frame prediction (CIIP). The following text provides details of those inter prediction methods specified in VVC. 2.1.2.1. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Airspace MVP from airspace adjacent CU 2) Temporal MVP from the same CU 3) History-based MVP from FIFO table 4) Paired Average MVP 5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU codec in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded with context, and bypass coding is used for the remaining binary bits. This section provides the derivation process of various types of Merge candidates. As done in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a specific size area. 2.1.2.1.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Figure 14Among the candidates at the positions shown, select up to four Merge candidates. The derivation order is B 0, 、A 0, 、B 1, , A1 and B2. Position B1 is considered only when one or more CUs at positions B2, A0, B0 and A1 are not available (for example, because they belong to another slice or piece) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 15 The pairs are linked by arrows in , and a candidate is added to the list only if the corresponding candidates used for redundancy check do not have the same motion information. 2.1.2.1.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Figure 16 As shown by the dotted line in , the scaled motion vector for the temporal merge candidate is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to 0. like Figure 17 As shown, the position for the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, is intra-coded, or is outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. 2.1.2.1.3. History-based Merge Candidate Derivation History-based MVP (HMVP) Merge candidates are added to the Merge list after the spatial MVP and TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (cleared). As long as there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find whether the same HMVP exists in the table. If found, the same HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. For spatial or temporal merge candidates, HMVP candidates are subjected to redundancy check. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates used for Merge list generation is set to (N<=4)?M:(8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate minus 1, the Merge candidate list construction process from HMVP is terminated. 2.1.2.1.4. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list. The predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge index in the merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference pictures. If only one motion vector is available, that motion vector is used directly. If no motion vector is available, the list remains invalid. When the merge list is not full after adding pairwise average merge candidates, a zero MVP will be inserted at the end until the maximum number of merge candidates is reached. 2.1.2.2.Merge Estimation Region Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). Candidate blocks in the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. In addition, only when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>> Log2ParMrgLevel and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>> The update process of the history-based motion vector predictor candidate list is only updated when (log2ParMrgLevel) is less than or equal to 0.001, where (xCb, yCb) is the top left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled as log2_parallel_merge_level_minus2 in the sequence parameter set. 2.1.2.3. Merge Mode with MVD (MMVD) In addition to the Merge mode in which the implicitly derived motion information is directly used for prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is also introduced in VVC. The MMVD flag is transmitted by signal immediately after the Skip flag and Merge flag are sent to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it is further refined by signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. A merge candidate flag is signaled to specify which candidate is used. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 18 As shown, the offset is added to the horizontal component or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 5. Table 5 – Relationship between distance index and predefined offsets The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 6. It should be noted that the meaning of the MVD symbol can change according to the information of the starting MV. When the starting MV is an unpredicted MV or a bidirectionally predicted MV, and both lists point to the same side of the current picture (that is, the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 6 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectionally predicted MV, and the two MVs point to different sides of the current picture (that is, the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the symbol in Table 6 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 6 – Sign of MV offset specified by direction index Direction IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.4. Bidirectional Prediction with CU-Level Weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-7) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) For non-merge CUs, the weight index is signaled after the motion vector difference; 2) For merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights are conditionally checked only for 1-pixel and 4-pixel motion vector precision if the current picture is a low-latency picture. – When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – Unequal weights are only conditionally checked when the two reference pictures in bidirectional prediction are the same. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using one context codec bit, followed by a bypass codec bit. The first context codec bit indicates whether equal weights are used; if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with fading effects. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets for the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. This can be applied to both normal merge mode and inherited affine merge mode. For constructed affine merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index of a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly for a CU. When a CU is encoded or decoded in CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.1.2.5. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: – The CU is coded using “true” bi-prediction mode, i.e., one of the two reference pictures precedes the current picture in display order and the other follows the current picture in display order – The distance from the two reference images to the current image (i.e., POC difference) is the same – Both reference images are short-term reference images. –CU is not encoded or decoded using Affine mode or ATMVP Merge mode –CU has more than 64 luma samples –CU height and CU width are both greater than or equal to 8 luminance samples –BCW weight index indicates equal weight – WP is not enabled for the current CU –CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 sub-block, the motion refinement (v x ,v y ). Motion refinement is then used to adjust the bidirectional prediction sample values in the 4x4 sub-block. The following steps are applied in the BDOF process. First, the horizontal gradient and vertical gradient of the two prediction signals and k = 0, 1 is calculated by directly calculating the difference between two adjacent sample points, that is, Among them I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k, k=0, 1, and shift1 is calculated as shift1=max(6, bitDepth-6) based on the luma bit depth bitDepth. Then, the autocorrelations and cross-correlations S1, S2, S3, S5, and S6 of the gradients are calculated as S1=∑ (i,j)∈Ω Abs(ψ x (i,j)),S3=∑ (i,j)∈Ω θ(i,j)·Sign(ψ x (i,j)) (2-9) S5=∑ (i,j)∈Ω Abs(ψ y (i,j)),S6=∑ (i,j)∈Ω θ(i,j)·Sign(ψ y (i,j)) in θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ) where Ω is a 6x6 window around the 4x4 sub-block, and n a and n b The values of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. Motion refinement (v x ,v y ) and then the cross-correlation and autocorrelation terms are derived using: in th′ BIO =2 max(5,BD-7) . is a floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+ο offset )>>shift (2-13). These values are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits. In order to derive the gradient value, some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j) needs to be generated. Figure 19As shown, BDOF in VVC uses an extended row / column around the CU boundary. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly obtaining the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are filled (i.e. repeated) with samples and gradient values from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are considered as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8* W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded in symmetric MVD mode or CIIP mode. 2.1.2.6. Symmetric MVD encoding and decoding In VVC, in addition to the normal unidirectional prediction and bidirectional prediction mode MVD signaling, a symmetric MVD mode for bidirectional prediction MVD signaling is also applied. In symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not transmitted through the signal but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a pair of forward and backward reference pictures or a pair of backward and forward reference pictures, then BiDirPredFlag is set to 1, and the reference pictures in list 0 and list 1 are both short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, a symmetric mode flag is explicitly signaled to indicate whether the symmetric mode is used. Figure 20 Figure 1 shows the symmetric MVD mode. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of List 0 and List 1 are set equal to a pair of reference pictures, respectively. MVD1 is set equal to (-MVD0). The resulting motion vector is shown in the following formula. In the encoder, symmetric MVD motion estimation starts with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.2.7. Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of MV in Merge mode, decoder-side motion vector refinement based on bilateral matching is applied in VVC. In bidirectional prediction operation, the refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between two candidate blocks in reference picture list L0 and list L1. Figure 21 As shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs coded and decoded with the following modes and features: – CU-level Merge mode with bi-predictive MV – With respect to the current picture, one reference picture is in the past and the other reference picture is in the future – The distances from both reference pictures to the current picture (i.e., POC differences) are the same – Both reference images are short-term reference images –CU has more than 64 luma samples –CU height and CU width are both greater than or equal to 8 luminance samples –BCW weight index indicates equal weight – WP is not enabled for the current block – CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction for future picture encoding and decoding. The original MV is used in the deblocking process and is also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.1.2.7.1. Search Scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR (represented by the candidate MV pair (MV0, MV1)) obeys the following two equations: MV0′=MV0+MV_offset (2-15) MV1′=MV1-MV_offset (2-16) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples away from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied for integer sample offset search. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. In order to reduce the impact of DMVR refinement uncertainty, it is proposed to bias the previous original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation rather than an additional search with SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates with the center having the minimum SAD in either the first or second iteration of the search. In the sub-pixel offset estimation based on the parameter error surface, the center position cost and the costs at the four neighboring positions from the center are used to fit the two-dimensional parabolic error surface equation of the following form E(x,y)=A(xx min ) 2 +B(yymin ) 2 +C (2-17) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the value of the minimum cost. By solving the above equation using the cost values of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-18) y min =(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-19). x min and y min The value of is automatically constrained to be between -8 and 8, since all cost values are positive and the minimum is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement delta MV. 2.1.2.7.2. Bilinear interpolation and sample filling In VVC, the resolution of the MV is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points surround the original fractional pixel MV with integer sample offsets, so those fractional samples need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter and a 2-sample search range, DVMR does not access more reference samples than the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV are padded with samples from those available. 2.1.2.7.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.1.2.8. Joint Inter-Frame and Intra-Frame Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the inter / intra joint prediction (CIIP) mode is applied to the current CU. Figure 22 The top and left neighboring blocks used in CIIP weight derivation are shown. As the name implies, CIIP prediction combines the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal P in CIIP mode inter The same inter-frame prediction process as the conventional Merge mode is used to derive the intra-frame prediction signal P intra The conventional intra prediction process with planar mode is derived. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values are calculated based on the codec modes of the top and left neighboring blocks as shown below: – If the top neighboring block is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighboring block is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) is equal to 2, then wt is set to 3; – Otherwise, if (isIntraLeft + isIntraTop) is equal to 1, wt is set to 2; – Otherwise, set wt to 1. The CIIP forecast is formed as follows: P CIIP =((4-wt)*P inter +wt*P intra +2)>>2 (2-20). 2.1.2.9. Multiple Hypothesis Prediction (MHP) In inter-AMVP mode, normal Merge mode and MMVD mode, up to two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the following table: add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. 2.1.2.10. Overlapped Sub-Block Motion Compensation (OBMC) When OBMC is applied, the top and left boundary pixels of the CU are refined using motion information of neighboring blocks with weighted prediction. The conditions under which OBMC should not be applied are as follows: When OBMC is disabled at the SPS level ●When the current block has intra mode or IBC mode ●When the current block applies LIC ●When the current luminance block area is less than or equal to 32. Sub-block boundary OBMC is performed by applying the same blending to the top, left, bottom and right sub-block boundary pixels using the motion information of the neighboring sub-blocks. It enables for sub-block based codecs: ●Affine AMVP mode; Affine Merge mode and sub-block based temporal motion vector prediction (SbTMVP); ●Sub-block based bilateral matching. 2.1.2.11. Local Illumination Compensation (LIC) LIC is an inter-frame prediction technique that models the local illumination variation between the current block and its prediction block as a function of the local illumination variation between the current block template and the reference block template. The parameters of this function can be represented by a scale α and an offset β, which form a linear equation, namely α*p[x]+β, to compensate for illumination variation, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. When surround motion compensation is enabled, the surround offset must be taken into account to limit the MV. Since α and β can be derived based on the current block template and the reference block template, they do not require signaling overhead, except for the signaling of the LIC flag for AMVP mode to indicate the use of LIC. The local illumination compensation proposed in JVET-O0066 is used for unidirectionally predicted inter CUs with the following modifications. ● Neighboring samples within the frame can be used to derive LIC parameters; ● Disable LIC for blocks with fewer than 32 luma samples; For both non-subblock mode and affine mode, LIC parameter derivation is performed based on the template block samples corresponding to the current CU, rather than based on the partial template block samples corresponding to the first top-left 16x16 unit; • The samples of the reference block template are generated by using MC with the block MV without rounding it to integer pixel precision. 2.1.2.12. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode for inter-frame prediction is supported. The geometric partitioning mode is signaled as a Merge mode using a CU level flag. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h=2 m ×2 n , the geometric partitioning mode supports a total of 64 partitions, where m,n∈{3…6} excludes 8x64 and 64x8. When this mode is used, the CU is divided into two parts by a geometrically positioned straight line ( Figure 23 ). The position of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each part has a motion vector and a reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion-compensated predictions are required per CU, as with regular bidirectional prediction. If the geometric partitioning mode is used for the current CU, a geometric partitioning index (angle and offset) and two Merge indexes (one for each partition) indicating the partitioning mode of the geometric partitioning are further transmitted by signal. The number of maximum GPM candidate sizes is explicitly transmitted by signal in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric partitioning, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partitioning edge. This is a prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored. 2.1.2.12.1. One-way prediction candidate list construction The unidirectional prediction candidate list is directly derived from the merge candidate list constructed according to the extended merge prediction process. Let n be the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are Figure 24 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partition mode instead. 2.1.2.12.2. Blending Along Geometric Partition Edges Figure 25An exemplary generation of warping weights w0 using geometric partitioning mode is shown. After each part of the geometric partition is predicted using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edge. The blending weight for each position of the CU is derived based on the distance between the individual position and the partition edge. The distance from the segmentation edge to the position (x, y) is derived as: where i,j are the indices of the angle and offset for the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,j The sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-25) w1(x,y)=1-w0(x,y) (2-27). partIdx depends on the angle index i. An example of the weight w0 is shown below. 2.1.2.12.3. Motion Field Storage for Geometric Partitioning Mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combination Mv of Mv1 and Mv2 are stored in the motion field of the CU coded in the geometric partition mode. The type of motion vector stored for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-28) where motionIdx is equal to d(4x+2,4y+2). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.1.2.12.4. GPM with Inter and Intra Prediction (GPM Inter-Intra) With GPM Inter-Intra, in addition to the Merge candidates for each non-rectangular partitioned area in the CU to which GPM is applied, a predefined intra prediction mode for the geometric partition line can also be selected. In the proposed method, for each GPM separated area, a flag from the encoder is used to determine whether it is intra prediction mode or inter prediction mode. When in inter prediction mode, a unidirectional prediction signal is generated by the MV from the Merge candidate list. On the other hand, when in intra prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra prediction mode specified by the index from the encoder. The variation of possible intra prediction modes is limited by the geometric shape. Finally, the two unidirectional prediction signals are mixed in the same way as ordinary GPM. 2.1.3. Screen content encoding and decoding tools 2.1.3.1. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC is in integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode different from the intra or inter prediction mode. The IBC mode is applicable to CUs with a width and height less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check on blocks with a width or height of no more than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on a 4x4 sub-block. For larger-sized current blocks, a hash key is determined to match the hash key of a reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the smallest cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighboring block and one from the upper neighboring block (if encoded with IBC). When either neighboring block is unavailable, the default block vector is used as the predictor. A flag is signaled to indicate the block vector predictor index. 2.1.3.1.1.IBC Reference Area To reduce memory consumption and decoder complexity, IBC in VVC only allows the reconstruction of a predefined area, which includes the area of the current CTU and some areas of the left CTU. Figure 26 The reference area of the IBC mode is shown, where each block represents a 64x64 luma sample unit. Depending on the location of the current codec CU position within the current CTU, the following applies: – If the current block falls into the upper left 64x64 block of the current CTU, in addition to the reconstructed samples in the current CTU, the CPR mode can also be used to reference the reference samples in the lower right 64x64 block of the left CTU. The current block can also use the CPR mode to reference the reference samples in the lower left 64x64 block of the left CTU and the reference samples in the upper right 64x64 block of the left CTU. – If the current block falls into the upper right 64x64 block of the current CTU, in addition to the samples that have been reconstructed in the current CTU, if the luma position (0,64) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the lower left 64x64 block and the lower right 64x64 block of the left CTU; otherwise, the current block can also refer to the reference samples in the lower right 64x64 block of the left CTU. – If the current block falls into the lower left 64x64 block of the current CTU, in addition to the samples already reconstructed in the current CTU, if the luma position (64,0) relative to the current CTU has not been reconstructed, the current block can also use the CPR mode to refer to the reference samples in the upper right 64x64 block and the lower right 64x64 block of the left CTU. Otherwise, the current block can also use the CPR mode to refer to the reference samples in the lower right 64x64 block of the left CTU. If the current block falls into the lower right 64x64 block of the current CTU, it can only use the CPR mode to refer to the samples that have been reconstructed in the current CTU. This restriction allows the IBC mode to be implemented for hardware implementations using local on-chip memory. 2.1.3.1.2. Interaction between IBC and other codecs The interaction between IBC mode and other inter-frame codecs in VVC (such as paired merge candidates, history-based motion vector predictor (HMVP), intra / inter joint prediction mode (CIIP), merge mode with motion vector difference (MMVD) and geometric partitioning mode (GPM)) is as follows: – IBC can be used with pairwise merge candidates and HMVP. A new pairwise IBC merge candidate can be generated by averaging two IBC merge candidates. For HMVP, the IBC motion is inserted into the history cache for future reference. – IBC cannot be used in combination with the following interframe tools: Affine Motion, CIIP, MMVD, and GPM. – When using DUAL_TREE partitioning, IBC is not allowed for chroma codec blocks. Unlike the HEVC screen content codec extension, the current picture is no longer included as one of the reference pictures in reference picture list 0 for IBC prediction. The derivation process of motion vectors for IBC mode excludes all neighboring blocks in inter mode, and vice versa. The following IBC design aspects are applied: – IBC shares the same process as regular MV Merge, including pairwise merge candidates and history-based motion prediction values, but TMVP and zero vectors are not allowed because they are invalid for IBC mode. – Separate HMVP buffers (5 candidates each) are used for traditional MV and IBC. – Block vector constraints are implemented as bitstream consistency constraints. The encoder needs to ensure that no invalid vectors exist in the bitstream and that a merge should not be used if the merge candidate is invalid (out of range or 0). As described below, this bitstream consistency constraint is expressed in terms of a virtual buffer. – For deblocking, IBC is handled as inter mode. If the current block is coded using IBC prediction mode, AMVR does not use quarter pixels; instead, AMVR is signaled to only indicate whether the MV is inter pixels or 4 integer pixels. - The number of IBC Merge candidates may be signaled in the slice header separately from the number of normal Merge candidates, sub-block Merge candidates, and geometric Merge candidates. The concept of virtual buffer is used to describe the allowable reference area and valid block vectors for IBC prediction mode. Denoting the CTU size as ctbSize, the virtual buffer ibcBuf has a width wIbcBuf=128x128 / ctbSize and a height hIbcBuf=ctbSize. For example, for a CTU size of 128x128, the size of ibcBuf is also 128x128; for a CTU size of 64x64, the size of ibcBuf is 256x64; and for a CTU size of 32x32, the size of ibcBuf is 512x32. The size of the VPDU is min(ctbSize, 64) in each dimension, W v =min(ctbSize, 64). The virtual IBC buffer ibcBuf is maintained as follows. – At the start of decoding each CTU line, flush the entire ibcBuf with an invalid value of -1. – At the start of decoding the VPDU (xVPDU, yVPDU) relative to the top left corner of the picture, set ibcBuf[x][y] = -1, where x = xVPDU % wIbcBuf, ..., xVPDU % wIbcBuf + W v -1; y= yVPDU%ctbSize,…,yVPDU%ctbSize+W v -1. – After decoding the CU containing (x, y) relative to the top left corner of the picture, set ibcBuf[x%wIbcBuf][y%ctbSize]=recSample[x][y]. For a block covering coordinates (x, y), it is valid if the following is true for the block vector bv = (bv[0], bv[1]); otherwise, it is invalid: ibcBuf[(x+bv[0])%wIbcBuf][(y+bv[1])%ctbSize] should not be equal to -1. 2.1.3.2. Block Differential Pulse Coded Modulation (BDPCM) VVC supports Block Differential Pulse Coded Modulation (BDPCM) for screen content encoding and decoding. At the sequence level, the BDPCM enable flag is signaled in the SPS; this flag is signaled only when transform skip mode (described in the next section) is enabled in the SPS. When BDPCM is enabled, if the CU size in terms of luma samples is less than or equal to MaxTsSize by MaxTsSize, and if the CU is intra-coded, a flag is transmitted at the CU level, where MaxTsSize is the maximum block size for which transform skip mode is allowed. The flag indicates whether conventional intra-coding or BDPCM is used. If BDPCM is used, a BDPCM prediction direction flag is transmitted to indicate whether the prediction is horizontal or vertical. The block is then predicted using a conventional horizontal or vertical intra prediction process with unfiltered reference samples. The residuals are quantized, and the difference between each quantized residual and its predicted value (i.e., the residual of the previously coded residual at a neighboring position horizontally or vertically (depending on the BDPCM prediction direction)) is coded. For a block of size M (height) × N (width), such that r i,j ,0≤i≤M-1,0≤j≤N-1 is the prediction residual. i,j ), 0≤i≤M-1,0≤j≤N-1 represents the residual r i,j quantized version of . BDPCM is applied to the quantized residual values, resulting in a quantized version of The modified M×N array in is predicted from its neighboring quantized residual values. For vertical BDPCM prediction mode, for 0≤j≤(N-1), the following is used to derive For the horizontal BDPCM prediction mode, for 0≤i≤(M-1), the following is used to derive At the decoder side, the above process is reversed to calculate Q(r i,j ), 0≤i≤M-1,0≤j≤N-1, as shown below: Dequantized residual Q -1 (Q(r i,j )) is added to the intra block prediction value to produce the reconstructed sample value. Predicted quantized residual value The residual codec is sent to the decoder using the same residual codec process as in transform skip mode residual codec. For lossless codecs, if slice_ts_residual_coding_disabled_flag is set to 1, the quantized residual values are sent to the decoder using regular transform residual codec. In terms of MPM modes for future intra mode codecs, the horizontal or vertical prediction mode is stored for the BDPCM-coded CU if the BDPCM prediction direction is horizontal or vertical, respectively. For deblocking, if both blocks on either side of a block boundary are coded using BDPCM, then that particular block boundary is not deblocked. 2.1.3.3. Residual coding and decoding for transform skip mode VVC allows transform skip mode to be used for luminance blocks of size up to MaxTsSize by MaxTsSize, where the value of MaxTsSize is transmitted by signal in the PPS and can be up to 32. When a CU is encoded and decoded in transform skip mode, its prediction residual is quantized and encoded using the transform skip residual coding and decoding process. This process is modified from the transform coefficient coding and decoding process. In transform skip mode, the residual of the TU is also encoded and decoded in units of non-overlapping sub-blocks of size 4x4. For better coding and decoding efficiency, some modifications are made to customize the residual coding and decoding process to the characteristics of the residual signal. The following summarizes the differences between transform skip residual coding and conventional transform residual coding and decoding: – Forward scan order is applied to scan sub-blocks within a converted block and positions within sub-blocks; – There is no through signaling of the last (x, y) position; – when all previous flags are equal to 0, coded_sub_block_flag is encoded and decoded for each sub-block except the last sub-block; –sig_coeff_flag context modeling uses a reduced template, and the context model of sig_coeff_flag depends on the top and left neighboring values; – The context model of the abs_level_gt1 flag also depends on the sig_coeff_flag values on the left and top; –par_level_flag uses only one context model; – Additional flags greater than 3, 5, 7, 9 are signaled to indicate coefficient levels, one context per flag; – Binarization of the remainder values using fixed order = 1 Rice parameter derivation; – The context model for the sign flag is determined based on the neighboring values to the left and above, and the sign flag is parsed after sig_coeff_flag to keep all context-coded bins together. For each subblock, if coded_subblock_flag is equal to 1 (i.e., there is at least one non-zero quantized residual in the subblock), encoding and decoding the quantized residual level is performed in three scanning passes (see Figure 27 ): – First scan pass: Encode and decode the significance flag (sig_coeff_flag), the sign flag (coeff_sign_flag), the absolute level greater than 1 flag (abs_level_gtx_flag[0]), and the parity check (par_level_flag). For a given scan position, if sig_coeff_flag is equal to 1, then coeff_sign_flag is encoded and decoded, followed by abs_level_gtx_flag[0] (which specifies whether the absolute level is greater than 1). If abs_level_gtx_flag[0] is equal to 1, then par_level_flag is also encoded and decoded to specify the parity check of the absolute level. – Greater than x scan passes: For each scan position where the absolute level is greater than 1, up to four abs_level_gtx_flag[i] (i=1...4) are encoded to indicate whether the absolute level at the given position is greater than 3, 5, 7 or 9, respectively. – Remainder scan pass: The remainder of the absolute level abs_remainder is encoded and decoded in bypass mode. The remainder of the absolute level is binarized using a fixed Rice parameter value of 1. The bins in scan passes #1 and #2 (the first scan pass and greater than x scan passes) are context coded until the maximum number of context coded bins in the TU has been exhausted. The maximum number of context coded bins in the residual block is limited to 1.75*block_width*block_height, or equivalently, an average of 1.75 context coded bins per sample position. The bins in the last scan pass (the remainder scan pass) are bypass coded. The variable RemCcbs is first set to the maximum number of context coded bins for the block and is decremented by 1 each time a context coded bin is coded. When RemCcbs is greater than or equal to 4, the syntax elements in the first codec pass are coded using context coded bins, including sig_coeff_flag, coeff_sign_flag, abs_level_gt1_flag, and par_level_flag. If RemCcbs becomes smaller than 4 when encoding and decoding the first pass, the remaining coefficients that have not been encoded in the first pass are encoded and decoded in the remainder scan pass (pass #3). After completing the first pass of encoding and decoding, if RemCcbs is greater than or equal to 4, the syntax elements in the second pass of encoding and decoding are encoded and decoded using the bins of the context encoding, including abs_level_gt3_flag, abs_level_gt5_flag, abs_level_gt7_flag, and abs_level_gt9_flag. If RemCcbs becomes less than 4 during the encoding and decoding of the second pass, the remaining coefficients in the second pass that have not been encoded and decoded are encoded and decoded in the remainder scan pass (pass #3). Figure 27 The transform skip residual encoding process is shown. The asterisks mark the locations where the context coded bits are exhausted and all remaining bits are encoded and decoded using the bypass codec. In addition, for blocks that are not coded in BDPCM mode, a level mapping mechanism is applied to transform skip residual coding until the maximum number of context coding bins is reached. Level mapping uses the top and left neighboring coefficient levels to predict the current coefficient level in order to reduce signaling cost. For a given residual position, denote absCoeff as the absolute coefficient level before mapping and absCoeffMod as the coefficient level after mapping. Let X0 denote the absolute coefficient level of the left neighboring position and let X1 denote the absolute coefficient level of the upper neighboring position. Level mapping is performed as follows: The absCoeffMod value is then encoded and decoded as described above.After all context-encoded bins are exhausted, level mapping is disabled for all remaining scan positions in the current block. 2.1.3.4. Palette Mode In VVC, palette mode is used for encoding and decoding screen content in all chroma formats supported in the 4:4:4 profile (i.e., 4:4:4, 4:2:0, 4:2:2, and monochrome). When palette mode is enabled, if the CU size is less than or equal to 64x64 and the number of samples in the CU is greater than 16, a flag is transmitted at the CU level to indicate whether palette mode is used. Considering that applying palette mode on small CUs introduces insignificant coding gain and brings additional complexity to small blocks, palette mode is disabled for CUs with less than or equal to 16 samples. A palette-encoded codec unit (CU) is treated as a prediction mode different from intra prediction, inter prediction, and intra block copy (IBC) mode. If palette mode is used, the sample values in the CU are represented by a set of representative color values. This set is called a palette. For positions where the sample values are close to the palette colors, the palette index is transmitted through the signal. It is also possible to specify samples outside the palette by signaling a jump symbol. For samples within the CU that are encoded and decoded using the jump symbol, their component values are directly transmitted through the signal using (possibly) quantized component values. This is in Figure 28 The quantized escape symbols are binarized using the fifth-order Exp-Golomb binarization process (EG5). In order to encode and decode the palette, the palette prediction value is maintained. In the non-wavefront case, the palette prediction value is initialized to 0 at the beginning of each stripe. For the WPP case, the palette prediction value at the beginning of each CTU row is initialized to the prediction value derived from the first CTU in the previous CTU row, so that the initialization scheme between the palette prediction value and the CABAC synchronization is unified. For each entry in the palette prediction value, the reuse flag is transmitted by signal to indicate whether it is part of the current palette in the CU. The reuse flag is sent using a run length codec of zero. Thereafter, the number of new palette entries and the component values of the new palette entries are transmitted by signal. After encoding the palette-encoded CU, the palette prediction value will be updated using the current palette, and entries from the previous palette prediction value that are not reused in the current palette will be added to the end of the new palette prediction value until the maximum allowed size is reached. A jump flag is transmitted by signal for each CU to indicate whether a jump symbol exists in the current CU. If there is a jump symbol, the palette table is incremented by 1 and the last index is assigned as the jump symbol. In a manner similar to the coefficient groups (CGs) used in transform coefficient coding, a CU coded using palette mode is divided into multiple row-based coefficient groups, each consisting of m samples (i.e., m=16), where the index run, palette index value, and quantized color for jump mode are encoded / parsed in sequence for each CG. As in HEVC, horizontal or vertical traversal scanning can be applied to scan samples, such as Figure 29 shown. The coding order for palette run-length coding in each segment is as follows: For each sample position, one context-coded binary bit run_copy_flag=0 is signaled to indicate whether the pixel has the same mode as the previous sample position, i.e., whether the previously scanned sample and the current sample are both of run type COPY_ABOVE, or whether the previously scanned sample and the current sample are both of run type INDEX and have the same index value. Otherwise, run_copy_flag=1 is signaled. If the current sample and the previous sample have different modes, one context-coded binary bit copy_above_palette_indices_flag is signaled to indicate the run type of the current sample, i.e., INDEX or COPY_ABOVE. Here, if the sample is in the first row (horizontal traversal scan) or the first column (vertical traversal scan), the decoder does not have to resolve the run type because INDEX mode is used by default. In the same way, if the previously resolved run type is COPY_ABOVE, the decoder does not have to resolve the run type. After palette-encoding the samples in one pass, the index values (for INDEX mode) and quantized jump colors are grouped and encoded using CABAC bypass encoding in another pass. This separation of context-encoded bits and bypass-encoded bits can improve throughput within each row CG. For slices with dual luma / chroma trees, the palettes are applied separately to luma (Y component) and chroma (Cb and Cr components), where the luma palette entries contain only Y values and the chroma palette entries contain both Cb and Cr values. For slices with a single tree, the palette will be applied jointly to the Y, Cb, Cr components, i.e., each entry in the palette contains Y, Cb, Cr values, except when the CU is coded using a local dual tree, in which case the coding of luma and chroma is handled separately. In this case, if the corresponding luma or chroma blocks are coded using palette mode, their palettes are applied in a manner similar to the dual tree case (this is related to non-4:4:4 codecs and will be further explained in 2.1.3.4.1). For slices coded with dual-tree, the maximum palette predictor size is 63, and the maximum palette table size for coding the current CU is 31. For slices coded with dual-tree, the maximum predictor size and palette table size are halved, i.e., for each entry in the luma palette and chroma palette, the maximum predictor size is 31, and the maximum table size is 15. For deblocking, palette-coded blocks on the side of a block boundary are not deblocked. 2.1.3.4.1. Palette Mode for Non-4:4:4 Content The palette mode in VVC supports all chroma formats in a similar way to the palette mode in HEVC SCC. For non-4:4:4 content, the following customizations apply: 1. When the pop-out value for a given sample position is signaled, if the sample position has only luma components and no chroma components due to chroma downsampling, only the luma pop-out value is signaled. This is the same as in HEVC SCC. 2. For local dual-tree blocks, the palette mode is applied to the block in the same way as the palette mode is applied to single-tree blocks, with the following two exceptions: a. The palette prediction value update process is slightly modified as follows. Since the local dual-tree block contains only luma (or chroma) components, the prediction value update process uses the signaled value of the luma (or chroma) component and fills in the "missing" chroma (or luma) component by setting it to the default value (1 << (component bit depth - 1)). b. The maximum palette prediction value size is kept at 63 (because slices are coded using a single tree), but the maximum palette table size for luma / chroma blocks is kept at 15 (because blocks are coded using separate palettes). 3. For palette mode in monochrome format, the number of color components in a palette-encoded block is set to 1 instead of 3. 2.1.3.4.2. Encoder Algorithm for Palette Mode On the encoder side, the following steps are used to generate the palette table of the current CU: 1. First, in order to derive the initial entries in the palette table of the current CU, a simplified K-means clustering is applied. The palette table of the current CU is initialized to an empty table. For each sample position in the CU, the SAD between the sample and each palette table entry is calculated, and the minimum SAD among all palette table entries is obtained. If the minimum SAD is less than the predefined error limit errorLimit, the current sample is clustered with the palette table entry with the minimum SAD. Otherwise, a new palette table entry is created. The threshold errorLimit is QP-dependent and is obtained from a lookup table containing 57 elements covering the entire QP range. After all samples of the current CU are processed, the initial palette entries are sorted according to the number of samples clustered with each palette entry, and any entries after the 31st entry are discarded. 2. In the second step, the initial palette table colors are adjusted by considering two options: using the centroid of each cluster from step 1 or using one of the palette colors in the palette prediction value. The option with the lower rate-distortion cost is selected as the final color of the palette table. If a cluster has only a single sample and the corresponding palette entry is not in the palette prediction value, the corresponding sample is converted to a jump symbol in the next step. 3. The palette table thus generated contains some new entries from the centroids of the clusters in step 1 and some entries from the palette predictions. Therefore, the table is reordered again so that all new entries (i.e. centroids) are placed at the beginning of the table, followed by entries from the palette predictions. Given the palette table of the current CU, the encoder selects a palette index for each sample position in the CU. For each sample position, the encoder checks the RD cost of all index values corresponding to the palette table entries and the RD cost of the index representing the escape symbol, and selects the index with the smallest RD cost using the following equation: RD cost = distortion × (isChroma?0.8:1) + lambda × bits of bypass coding (2-33). After determining the index map for the current CU, each entry in the palette table is checked to see if it is used by at least one sample position in the CU. Any unused palette entry will be deleted. After the index map of the current CU is determined, grid RD optimization is applied to find the best value of run_copy_flag and run type for each sample position by comparing the RD cost of three options: the same as the previously scanned position, run type COPY_ABOVE or run type INDEX. When calculating the SAD value, the sample value is reduced to 8 bits unless the CU is coded in lossless mode, in which case the actual input bit depth is used to calculate the SAD. In addition, in the case of lossless codecs, only the rate is used in the above-mentioned rate-distortion optimization step (because lossless codecs do not produce distortion). 2.1.3.5. Adaptive Color Transformation In the HEVC SCC extension, Adaptive Color Transform (ACT) is applied to reduce redundancy between the three color components in the 444 chroma format. ACT has also been adopted into the VVC standard to improve the codec efficiency of the 444 chroma format. As in HEVC SCC, ACT performs an in-loop color space conversion in the prediction residual domain by adaptively converting the residual from the input color space to the YCgCo space. Figure 30The decoding flow chart for applying ACT is shown. The two color spaces are adaptively selected by signaling an ACT flag at the CU level. When the flag is equal to 1, the residual of the CU is encoded and decoded in the YCgCo space; otherwise, the residual of the CU is encoded and decoded in the original color space. In addition, similar to the HEVC ACT design, for inter-frame and IBC CUs, ACT is enabled only when there is at least one non-zero coefficient in the CU. For intra-frame CUs, ACT is enabled only when the chroma component selects the same intra-frame prediction mode (i.e., DM mode) as the luminance component. 2.1.3.5.1.ACT mode In the HEVC SCC extension, ACT supports both lossless and lossy codecs based on the lossless flag (i.e., cu_transquant_bypass_flag). However, no flag is signaled in the bitstream to indicate whether lossy or lossless codec is applied. Therefore, the YCgCo-R transform is applied as ACT to support both lossy and lossless cases. The YCgCo-R reversible color transform is shown below. Because the YCgCo-R transform is not normalized, QP adjustments (-5, 1, 3) are applied to the transformed residuals of the Y, Cg, and Co components to compensate for the dynamic range changes of the residual signals before and after color transformation. The adjusted quantization parameters only affect the quantization and inverse quantization of the residual within the CU. For other codec processes (such as deblocking), the original QP is still applied. In addition, because the forward and inverse color transforms require access to the residuals of all three components, ACT mode is always disabled for ISP modes with separate tree partitioning and different prediction block sizes for different color components. When ACT is applied, transform skip (TS) and block differential pulse codec modulation (BDPCM) that are extended to codec the chroma residual are also enabled. 2.1.3.5.2.ACT Fast Encoding Algorithm To avoid brute force RD search in both the original and converted color spaces, the following fast encoding algorithm is applied in the VTM reference software to reduce encoder complexity when ACT is enabled. – The order of enabling / disabling RD checking for ACT depends on the original color space of the input video. For RGB video, the RD cost of ACT mode is checked first; for YCbCr video, the RD cost of non-ACT mode is checked first. The RD cost of the second color space is checked only if there is at least one non-zero coefficient in the first color space. – Reuse the same ACT enable / disable decision when a CU is obtained through a different split path. Specifically, when the CU is first encoded and decoded, the selected color space used to encode the residual of a CU is stored. Then, when the same CU is obtained through another split path, instead of checking the RD cost of the two spaces, the stored color space decision is directly reused. The RD cost of the parent CU is used to determine whether to check the RD cost of the second color space for the current CU. For example, if the RD cost of the first color space for the parent CU is less than the RD cost of the second color space, the second color space is not checked for the current CU. To reduce the number of codec modes tested, the selected codec mode is shared between the two color spaces. Specifically, for intra mode, pre-selected intra mode candidates based on SATD-based intra mode selection are shared between the two color spaces. For inter and IBC modes, block vector search or motion estimation is performed only once. Block vectors and motion vectors are shared between the two color spaces. 2.1.3.6. Intra-frame Template Matching (IntraTMP) Intra Template Matching (IntraTM) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. Within a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed portion of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. By combining the L-shaped causal neighbors of the current block with Figure 31 The prediction signal is generated by matching another block in a predefined search area in the image to generate a prediction signal, where the predefined search area includes: R1: Current CTU R2: Upper left CTU R3: Upper CTU R4: left CTU. SAD is used as the cost function. In each region, the decoder searches for the template with the smallest SAD relative to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW SearchRange_h=a*BlkH Where 'a' is a constant that controls the gain / complexity tradeoff. In practice, 'a' is equal to 5. The intra template matching tool is enabled for CUs with width and height dimensions less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.1.3.6.1. Using Block Vectors Derived from IntraTMP for IBC The block vector (BV) derived from intra template matching prediction (IntraTMP) is used for intra block copy (IBC). The stored IntraTMP BVs of neighboring blocks are used together with the IBC BVs as spatial BV candidates in IBC candidate list construction. 2.1.3.6.2. Direct Block Vector (DBV) Mode for Chroma Prediction For chroma components, when chroma dual-tree is activated in intra slices, if one of the luma blocks (five positions) is coded with MODE_IBC, its block vector bvL is used and scaled to derive the chroma block vector bvC. The scaling factor depends on the chroma format sampling structure. Then, by using the position (xCb, yCb) of the current chroma block and its bvC, the corresponding offset position (xCb+bvC[0], yCb+bvC[1]) is determined, and block copy prediction is performed. Figure 32 Five positions in the reconstructed luminance samples are shown. Figure 33 The prediction process of the DBV model is shown. As shown in Table 7, a CU level flag is signaled to indicate whether the proposed DBV mode is applied. Table 7 Binarization process for intra_chroma_pred_mode in the proposed method intra_chroma_pred_mode Binary string Chroma Intra mode 0 11100 List[0] 1 11101 List[1] 2 11110 List[2] 3 11111 List[3] 4 110 DIMD chromaticity 5 10 DM 6 0 DBV 2.1.4. Transformation and Quantization 2.1.4.1. Large Block Size Transformation with High-Frequency Zeroing In VVC, large block size transforms with a maximum size of 64×64 are enabled, which are mainly used for higher resolution videos, such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high-frequency transform coefficients are zeroed so that only the low-frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M is equal to 64, only the left 32 columns of transform coefficients are retained. Similarly, when N is equal to 64, only the first 32 rows of transform coefficients are retained. When transform skip mode is used for large blocks, the entire block is used without zeroing any values. In addition, transform shifts are removed in transform skip mode. VTM also supports a configurable maximum transform size in SPS, so that the encoder can flexibly select a transform size of up to 32 lengths or 64 lengths according to the needs of the specific implementation. 2.1.4.2. Multi-Transform Selection (MTS) for Core Transformations In addition to the DCT-II used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding of both inter-frame and intra-frame coded blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DCT-VII and DCT-VIII. Table 8 shows the selected DST / DCT basis functions. Table 8 – Transform basis functions for DCT-II / VIII and DSTVII for N-point input To maintain orthogonality of the transform matrix, the transform matrix is quantized more accurately than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within 16 bits, all coefficients have 10 bits after horizontal transform and after vertical transform. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is applied only to luma. MTS signaling is skipped when one of the following conditions applies: – The position of the last significant coefficient of the luma TB is less than 1 (i.e., DC only) – The last significant coefficient of the luminance TB lies within the MTS null region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, the other two flags are transmitted by signal to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 9. A unified transform selection for ISP and implicit MTS is used by removing the intra mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra block and both the intra explicit MTS and the inter explicit MTS are turned on, only DST7 is used for both the horizontal transform core and the vertical transform core. When it comes to transform matrix accuracy, an 8-bit main transform core is used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform cores (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8) use an 8-bit main transform core. Table 9 – Transformation and signalling mapping table In order to reduce the complexity of large-size DST-7 and DCT-8, the high-frequency transform coefficients are zeroed for DST-7 blocks and DCT-8 blocks with size (width or height, or both width and height) equal to 32. Only the coefficients in the 16x16 low-frequency area are retained. As in HEVC, the residual of the block can be encoded and decoded using the transform skip mode. In order to avoid redundancy in syntax encoding and decoding, the transform skip flag is not transmitted through the signal when the CU level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. When MTS is enabled for inter-frame codec blocks, implicit MTS can also be enabled. 2.1.4.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, LFNST is applied between the forward main transform and quantization (at the encoder) and between inverse quantization and inverse main transform (at the decoder side). In LFNST, a 4x4 non-separable transform or an 8x8 non-separable transform is applied depending on the block size. For example, 4x4 LFNST is applied for small blocks (i.e., min(width, height) < 8), and 8x8 LFNST is applied for larger blocks (i.e., min(width, height) > 4). Figure 34 The low frequency non-separable transform (LFNST) process is shown. Using the input as an example, the application of the non-separable transform used in LFNST is described as follows. To apply 4x4 LFNST, the 4x4 input block X First, it is represented as a vector The non-separable transform is computed as where indicates the transform coefficient vector, and T is a 16x16 transform matrix. Subsequently, the 16x1 coefficient vector is reorganized into 4x4 blocks using the scan order (horizontal, vertical, or diagonal) for that block. Coefficients with smaller indices are placed in positions with smaller scan indices in the 4x4 coefficient blocks. 2.1.4.3.1. Reduced Non-Separable Transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on a direct matrix multiplication method such that it is implemented in a single pass without multiple iterations. However, the non-separable transform matrix dimension needs to be reduced to minimize the computational complexity and the storage space for storing the transform coefficients. Therefore, the Reduced Non-Separable Transform (or RST) method is used in the LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (for 8x8 NSST, N is usually equal to 64) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, the RST matrix is not an NxN matrix but an R×N matrix as follows: The R rows of the transform are the R basis of the N-dimensional space. The inverse transform matrix for RT is the transpose of its forward transform. For 8x8 LFNST, a reduction factor of 4 is applied, and the 64x64 direct matrix (which is the conventional 8x8 non-separable transform matrix size) is reduced to a 16x48 direct matrix. Therefore, the 48×16 inverse RST matrix is used on the decoder side to generate the core (main) transform coefficients in the 8x8 upper left region. When a 16x48 matrix is applied instead of a 16x64 matrix with the same transform set configuration, each matrix takes 48 input data from three 4x4 blocks in the upper left 8x8 block, excluding the lower right 4x4 block. With the reduced dimensions, the memory usage for storing all LFNST matrices is reduced from 10KB to 8KB, while the performance degradation is reasonable. To reduce complexity, LFNST is only applied when all coefficients outside the first coefficient subgroup are insignificant. Therefore, when LFNST is applied, all main transform coefficients must be zero. This allows LFNST index signaling to be conditional on the last significant position, avoiding the extra coefficient scan required in current LFNST designs, which is only required when checking for significant coefficients at specific positions. The worst-case processing of LFNST (in terms of per-pixel multiplications) restricts the non-separable transforms for 4x4 and 8x8 blocks to 8x16 and 8x48 transforms, respectively. In these cases, when applying LFNST, the last significant scan position must be less than 8, and less than 16 for other sizes. For blocks of shape 4xN and Nx4 with N > 8, the proposed restriction means that LFNST is now applied only once, and only to the top-left 4x4 region. Since all primary-only coefficients are zero when LFNST is applied, the number of operations required for the primary transform is reduced in this case. From the encoder's perspective, coefficient quantization is significantly simplified when testing the LFNST transform. At most, the first 16 coefficients (in scan order) must be rate-distortion-optimized quantized; the remaining coefficients are forced to zero. 2.1.4.3.2. LFNST Transform Selection A total of 4 transform sets are used in LFNST, and each transform set uses 2 inseparable transform matrices (kernels). As shown in Table 10, the mapping from intra prediction modes to transform sets is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), transform set 0 is selected for the current chroma block. For each transform set, the selected inseparable secondary transform candidate is further specified by an LFNST index that is explicitly signaled. This index is signaled once in the bitstream for each intra CU after the transform coefficients. Table 10 – Transformation selection table 2.1.4.3.3.LFNST Index Signaling and Interaction with Other Tools Since LFNST is only applied when all coefficients outside the first coefficient subgroup are insignificant, the LFNST index encoding depends on the position of the last significant coefficient. In addition, the LFNST index is context-coded, but does not depend on the intra prediction mode, and only the first binary bit is context-coded. In addition, LFNST is applied to intra CUs in both intra and inter slices, and to both luma and chroma. If dual-tree is enabled, the LFNST indices for luma and chroma are signaled separately. For inter slices (dual-tree disabled), a single LFNST index is signaled and used for both luma and chroma. Considering that large CUs larger than 64x64 are implicitly partitioned (TU slices) due to the existing maximum transform size limit (64x64), LFNST index searches can quadruple the data cache for a given number of decoding pipeline stages. Therefore, the maximum size allowed for LFNST is limited to 64x64. Note that LFNST is only enabled with DCT2. LFNST index signaling is placed before MTS index signaling. The use of the scaling matrix for perceptual quantization is not obvious, because the scaling matrix specified for the main matrix can be used for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For single-tree partitioning mode, chroma LFNST is not applied. 2.1.4.4. Sub-block Transform (SBT) In VTM, sub-block transform is introduced for inter-frame predicted CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When cu_cbf is equal to 1 for an inter-frame predicted CU, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded and decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded using the inferred adaptive transform, and the other part of the residual block is zeroed. When SBT is used for inter-coded CU, SBT type and SBT location information are signaled in the bitstream. Figure 35As shown, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is like a binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is like an asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 in luminance samples, the 1:3 / 3:1 partitioning along that dimension is prohibited. A CU has a maximum of 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are selected in Figure 35 For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DCT-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slice, cbf, and horizontal and vertical core transform types of the residual block. SBT is not applied to CUs coded in inter-intra combination mode. 2.1.4.5. Maximum Transform Size and Zeroing of Transform Coefficients The CTU size and maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the largest intra-frame codec block can have a size of 128x128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform, no normal zeroing operation is applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are normalized to zero. 2.1.4.6. Enhanced MTS for intra-frame coding and decoding In the current VVC design, for MTS, only DST7 and DCT8 transform cores for intra-frame coding and inter-frame coding are utilized. Additional main transforms including DCT5, DST4, DST1 and identity transform (IDT) are adopted. In addition, the MTS set also depends on the TU size and intra-frame mode information. 16 different TU sizes are considered, and for each TU size, 5 different categories are considered according to the intra-frame mode information. For each category, 4 different transform pairs are considered, the same as for VVC. Note that although a total of 80 different categories are considered, some of these different categories often share exactly the same transform set. Therefore, there are 58 (less than 80) unique entries in the resulting LUT. For angle modes, the joint symmetry of TU shape and intra prediction is taken into account. Therefore, mode i (i>34) with TU shape AxB will be mapped to the same category corresponding to mode j=(68-i) with TU shape BxA. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16x4 block with mode 18 (horizontal prediction) and a 4x16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For wide-angle mode, the nearest conventional angle mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. The MTS index [0,3] is signaled using a 2-bit fixed length codec. 2.1.4.7. Secondary Transform: LFNST Extension with Large Kernel The LFNST design in VVC is extended as follows: The number of LFNST sets (S) and candidates (C) is extended to S=35 and C=3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: o For predModeIntra<2, lfnstTrSetIdx is equal to 2 οlfnstTrSetIdx=predModeIntra,for predModeIntra in [0,34] οlfnstTrSetIdx=68–predModeIntra, for predModeIntra in [35,66] • Three different kernels LFNST4, LFNST8, and LFNST16 are defined to indicate a set of LFNST kernels, which are applied to 4xN / Nx4 (N≥4), 8xN / Nx8 (N≥8), and MxN (M, N≥16), respectively. The kernel dimensions are specified as follows: (LFSNT4, LFNST8*, LFNST16*) = (16x16, 32x64, 32x96). Forward LFNST is applied to the upper left low frequency region, which is called the region of interest (ROI). When LFNST is applied, the main transform coefficients present in the region other than the ROI are zeroed, which is unchanged from the VVC standard. The ROI of LFNST16 is Figure 36It is depicted in . It consists of six 4x4 sub-blocks that are consecutive in scan order. Since the number of input samples is 96, the transform matrix for forward LFNST16 can be Rx96. In this contribution, R is chosen to be 32, and accordingly, 32 coefficients (two 4x4 sub-blocks) are generated from forward LFNST16, which are placed to follow the coefficient scan order. The ROI of LFNST8 is Figure 37 The forward LFNST8 matrix can be Rx64, and R is chosen to be 32. The generated coefficients are positioned in the same way as LFNST16. The mapping of intra prediction modes to these sets is shown in the table below. Table 11. Mapping of intra prediction modes to LFNST set indices Intra prediction mode -14 13 -12 -11 10 -9 8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 1 16 17 LFNST collection index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 1 16 17 Intra prediction mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 4 47 48 49 LFNST collection index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra prediction mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNST collection index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 22 2 2 2 2.1.4.8. Non-separable primary transform (NSPT) for intra-frame coding and decoding For block sizes 4x4, 4x8, 8x4, and 8x8, DCT-II+LFNST is replaced by NSPT. NSPT follows the design of LFNST, i.e., 3 candidates and 35 sets, which are selected based on intra mode. The kernel sizes are as follows: NSPT4x4: 16x16; NSPT4x8 / NSPT8x4: 32x20; NSPT8x8: 64x32. Therefore, 12 and 32 coefficients are zeroed for NSPT4x8 / NSPT8x4 and NSPT8x8, respectively. 2.1.4.9. Sign Prediction of Coefficients The basic idea of the sign prediction method is to compute the reconstructed residuals for both negative and positive sign combinations for the applicable transform coefficients and to select the hypothesis that minimizes the cost function. To derive the optimal symbol, the cost function is defined as Figure 38 Discontinuity measures for block boundaries are shown. The cost function is measured for all hypotheses, and the hypothesis with the smallest cost is selected as the predicted value for the coefficient sign. The cost function is defined as the sum of the absolute second-order derivatives in the residual domain of the upper row and left column as follows: where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. -1 +2R0-P1) can only be calculated once for each block, and only the residual hypothesis is subtracted. 2.2 Block-Level Adaptive OBMC in Video Coding and Decoding – Neighboring Prediction Mode Correlation The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In the present disclosure, regarding a “block coded in mode N”, “mode N” here may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or codec techniques (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, Affine, CIIP, GPM, Spatial GPM, SGPM, GPM Inter-Inter, GPM Intra-Intra, GPM Inter-Intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, Deblocking, SAO, Bilateral Filter, LMCS and corresponding variants, etc.). It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 2.2.1. In one example, whether OBMC is applied to the current block may depend on the prediction modes of spatial / temporal neighboring blocks that are adjacent / non-adjacent to the current block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on whether there are neighboring blocks that are encoded and decoded using IBC. 4) For example, it can be based on whether there are neighboring blocks that are coded using PLT. 5) For example, it can be based on whether there are neighboring blocks that are coded using intraTMP. 6) For example, it can be based on whether there are neighboring blocks that are coded using BDPCM. 7) For example, it can be based on whether there are neighboring blocks that are coded using transform skipping. 8) For example, the neighboring block may be adjacent to the current block. 9) For example, the neighboring block may be non-adjacent to the current block. 10) For example, the neighboring block may be a spatial neighboring block within the current picture. 11) For example, the neighboring block may be a time domain block in a reference picture. 12) For example, the neighboring block may be a sub-block (eg, 4x4 or 8x8) smaller than the current block. 13) For example, the neighboring block may be a video unit that is larger than or equal to the current block. 14) For example, the neighboring blocks may be sample locations. 15) For example, a series of adjacent neighboring blocks / sub-blocks to the left and / or above the current block may be checked one by one (eg, following a predefined position and a predefined checking order). a. For example, if there is a neighbor that is coded using a specific mode, the process is terminated and OBMC is considered not to apply to the current block. b. For example, if there is a neighbor that is coded using inter-frame mode, it is further checked whether its reference block is coded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded neighbor and the position of such inter-frame coded neighbor), and if the reference block is coded using a specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in a reference picture. c. For example, if there is a neighbor that is encoded using intraTMP, it can be further checked whether its reference block is encoded using a specific mode (for example, the reference block is identified by adding the block vector associated with such intraTMP-encoded neighbor and the position of such intraTMP-encoded neighbor), and if the reference block is encoded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in the current picture. d. For example, the specific mode may be IBC and / or PLT. e. For example, the specific mode may be intraTMP. f. For example, the specific mode may be BDPCM. g. For example, the specific mode may be transform skip. 16) For example, a series of non-adjacent neighboring blocks / sub-blocks in the coded area of the current picture may be checked one by one (eg, following a predefined position and a predefined checking order). h. For example, if there is a neighbor that is coded using a specific mode, the process is terminated and OBMC is considered not to apply to the current block. i. For example, if there is a neighbor that is coded using inter-frame mode, then its reference block is further checked whether it is coded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded neighbor and the position of such inter-frame coded neighbor), and if the reference block is coded using a specific mode, then the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in a reference picture. j. For example, if there is a neighbor that is encoded using intraTMP, it is further checked whether its reference block is encoded using a specific mode (for example, the reference block is identified by adding the block vector associated with such intraTMP-encoded neighbor and the position of such intraTMP-encoded neighbor), and if the reference block is encoded using the specific mode, the process is terminated and OBMC is considered not to be applied to the current block. i. For example, the reference block is in the current picture. k. For example, the specific mode may be IBC and / or PLT. 1. For example, the specific mode may be intraTMP. m. For example, the specific mode may be BDPCM. n. For example, the specific mode may be transform skip. 17) For example, a series of temporal blocks / sub-blocks in a reference picture may be checked one by one (eg, following a predefined position and order). o. For example, if there is a time domain block that is coded using a specific mode, the process is terminated and OBMC is considered not to apply to the current block. p. For example, if there is a time domain block coded and decoded using the inter-frame mode, it is further checked whether the reference block is coded and decoded using a specific mode (for example, the reference block is identified by adding the motion vector associated with such inter-frame coded time domain block and the position of such inter-frame coded time domain block), and if the reference block is coded and decoded using the specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. q. For example, if there is a time domain block that is coded using intraTMP, it is further checked whether its reference block is coded using a specific mode (for example, the reference block is identified by adding a block vector associated with such intraTMP coded time domain block and the position of such intraTMP coded time domain block), and if the reference block is coded using the specific mode, the process is terminated and it is considered that OBMC is not applied to the current block. r. For example, the specific mode may be IBC and / or PLT. s. For example, a specific mode may be intraTMP. t. For example, the specific mode may be BDPCM. u. For example, the specific mode may be transform skip. 2.2.2. In one example, whether OBMC is applied to the current block may depend on the prediction mode of the reference block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, the reference block may be a block / sub-block identified based on adding a displacement (eg, predefined, or based on a motion vector, or based on a block vector) to the position of the first block. a. For example, the first block may be the current block. b. For example, the first block may be a neighboring block adjacent to the current block. c. For example, the first block may be a neighboring block that is non-adjacent to the current block. d. For example, the first block may be a reference block for the current block. e. For example, the first block may be a reference block of a neighboring block. f. For example, the reference block may be identified based on the position of the current block encoded in inter-frame mode and its motion information associated with the current block (eg, motion vector and reference index). i. For example, in this case, the reference block is in a reference picture. g. For example, the reference block may be identified based on the position of the current block encoded in the intraTMP mode and its motion information (eg, block vector) associated with the current block. i. For example, in this case, the reference block is in the current picture. h. For example, the reference block may be identified based on the position of the inter-mode coded neighboring block and its motion information (eg, motion vector and reference index) associated with the inter-mode coded neighboring block. i. For example, in this case, the reference block is in a reference picture. i. For example, a reference block may be identified based on the location of a neighboring block of an intraTMP mode codec and its motion information (eg, a block vector) associated with the neighboring block of the intraTMP mode codec. i. For example, in this case, the reference block is in the current picture. j. For example, the reference block may be identified based on the position of the inter-mode coded reference block and its motion information (eg, motion vector and reference index) associated with the inter-mode coded reference block. i. For example, in this case, the reference block is in another reference picture (not the reference picture where the reference block of inter-mode coding is located). k. For example, the reference block may be identified based on the position of the reference block encoded in the intraTMP mode and its motion information (eg, block vector) associated with the reference block encoded in the intraTMP mode. i. For example, in this case, the reference block is in the same reference picture as the reference block encoded in the intraTMP mode. 4) For example, the reference block may be checked when the neighboring block is coded using the inter mode. a. If the neighbor is inter-coded, the reference block is identified by adding the motion vector associated with such inter-coded neighbor and the position of such inter-coded neighbor, and if the reference block is coded using a specific mode, OBMC is considered not to apply to the current block. b. For example, the specific mode may be IBC and / or PLT. c. For example, the specific mode may be intraTMP. d. For example, the specific mode may be BDPCM. e. For example, the specific mode may be transform skip. 5) For example, the reference block may be checked when the neighboring block is coded using the intraTMP mode. a. If the neighbor block is intraTMP coded, the reference block is identified by adding the block vector associated with such intraTMP coded neighbor and the position of such intraTMP coded neighbor, and if the reference block is coded using a specific mode, OBMC is considered not to apply to the current block. b. For example, the specific mode may be IBC and / or PLT. c. For example, the specific mode may be BDPCM. d. For example, the specific mode may be transform skip. 2.2.3. In one example, whether OBMC is applied to the current block may depend on the prediction mode of the reference block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, the reference block of the reference block may be identified by adding the motion vector associated with the inter-mode coded reference block and the position of such inter-coded reference block. 4) For example, the reference block of the reference block may be identified by adding a block vector associated with the reference block coded in the intraTMP mode and the position of such intraTMP coded reference block. 5) For example, historical / propagated prediction patterns can be stored in a cache. a. For example, if the block itself is coded using a specific mode or it once had a reference block coded using a specific mode, the specific mode may be stored in a cache associated with the block information, indicating that it has history / propagation information using the specific mode. 6) For example, the specific mode may be IBC and / or PLT. 7) For example, the specific mode may be intraTMP. 8) For example, the specific mode may be BDPCM. 9) For example, the specific mode may be transform skip. 2.2.4 In one example, whether a block in the current picture is coded into a specific mode may be stored in a buffer. 1) For example, the specific mode may be IBC. 2) For example, the specific mode may be PLT. 3) For example, the specific mode may be intraTMP. 4) For example, the specific mode may be BDPCM. 5) For example, the specific mode may be transform skip. 6) For example, a shared parameter / cache may be used to store whether a block is encoded or decoded using IBC or PLT. a. For example, storage may require a single parameter / cache. 7) For example, a separate parameter / buffer may be used to store whether a block is coded using IBC or PLT or intraTMP or BDPCM. b. For example, storage may require multiple parameters / buffers. 8) For example, it can be stored at a granularity of M×M (eg, M=4 or 8) sub-blocks. 2.2.5 In one example, whether a block and / or its reference blocks are coded into a specific mode may be stored in a cache. 1) For example, the specific mode may be IBC. 2) For example, the specific mode may be PLT. 3) For example, the specific mode may be intraTMP. 4) For example, the specific mode may be BDPCM. 5) For example, the specific mode may be transform skip. 6) For example, the block or its reference blocks are encoded using IBC or PLT, a parameter equal to true (eg, indicating that it is a historical / propagated screen content block) may be stored in the cache. a. Alternatively, in another aspect, a parameter equal to false (eg, indicating that it is not a historical / propagated block of screen content) may be stored in the cache. 7) For example, whether a block is coded using IBC or PLT and whether the reference blocks of such block are coded using IBC or PLT may be stored as separate parameters and in separate buffers. 8) For example, it can be stored at a granularity of M×M (eg, M=4 or 8) sub-blocks. 2.2.6. In one example, whether to enable OBMC can be coupled with whether to enable a specific tool. 1) For example, the specific tool may be an IBC. 2) For example, the specific tool may be a PLT. 3) For example, the specific tool may be intraTMP. 4) For example, the specific tool may be BDPCM. 5) For example, a specific tool may be a transform skip. 6) In one example, whether a specific tool is applied may be controlled by a first syntax element (SE), such as in a VPS / SPS / PPS / slice header / CTU / CU / etc. 7) In one example, whether to apply OBMC may be controlled by a second syntax element (SE), such as in VPS / SPS / PPS / slice header / CTU / CU / etc. 8) In one example, it may be constrained that if the first SE instructs to enable a specific tool, the second SE must instruct to disable OBMC. 9) In one example, if the first SE indicates that a specific tool is enabled, a second SE may be set at the encoder to indicate that OBMC must be disabled. 2.2.7. Whether and / or how to apply the methods disclosed above may be signaled at the sequence level / GOP level / picture level / slice level / slice group level, such as in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 2.2.8. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. 2.2.9. Whether and / or how to apply the method disclosed above may depend on codec information, such as block size, color format, single-tree partitioning / dual-tree partitioning, color component, slice / picture type. 3. Question In ECM-7.0, OBMC can be applied to inter-frame codec blocks, regardless of whether it is inter-frame AMVP coded or inter-frame MERGE coded. For inter-frame AMVP coded blocks, a syntax flag can be signaled at the block level to indicate whether OBMC is applied. For inter-frame merged coded blocks, OBMC is implicitly inferred to be applied, regardless of block characteristics and codec information of neighboring blocks. However, there may be some cases where OBMC mode is not preferred for inter-frame MERGE coded blocks. For example, blocks containing sharp edges, or a small amount of gradient, or a small amount of color may not prefer OBMC mode. Block-level adaptive OBMC that considers the prediction modes of neighboring blocks can bring higher coding gain. 4. Detailed solution The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. The term “video unit” or “codec unit” or “block” may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB. In the present disclosure, regarding “a block encoded and decoded in mode N”, “mode N” here can be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a coding and decoding technology (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS and corresponding variants, etc.). It should be noted that the terms mentioned below are not limited to the specific terms defined in existing standards, and any changes in codec tools are also applicable. 4.1. In one example, whether OBMC is applied to the current block may depend on sample values of samples within the current block (and / or adjacent to the current block). 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on the prediction samples of the current block (before OBMC). 4) For example, it can be based on prediction samples that are adjacent to the current block. 5) For example, it can be based on the reconstructed samples that are adjacent to the current block. 6) For example, it can be based on the gradient / direction / angle (or histogram of gradient / direction / angle) of samples inside the current block (and / or adjacent to the current block). a. For example, the current prediction sample point before OBMC can be used. b. For example, adjacent reconstruction points can be used. c. For example, a histogram of gradients / directions / angles can be calculated based on the count of gradients along a specific direction / angle. i. For example, a specific direction / angle may be predefined. ii. For example, the specific direction / angle may be based on the direction of an intra prediction angle mode in a video codec. iii. For example, for a specific direction / angle, the gradient magnitude may be calculated based on the count of the gradient (magnitude) of at least one sample point in the current block. 1. For example, the (magnitude of) the gradient at a specific location (eg, the center) in the current block can be counted. 2. For example, the gradients (magnitudes) at a series of specific locations in the current block can be counted. 3. For example, the gradients (magnitudes) of all samples in the current block may be counted. 4. For example, the gradients (magnitudes) of all sample points in the current block except the sample points in the first row, last row, first column, and last column may be counted. iv. For example, for a specific direction / angle, the gradient magnitude may be calculated based on the count of the gradient (magnitude) of at least one sample point adjacent to the current block. 1. For example, the gradient (magnitude) at a specific location adjacent to the current block may be counted. 2. For example, the gradients (magnitudes) of all samples adjacent to (on the left and / or above) the current block may be counted. d. For example, a histogram of gradients / directions / angles can be calculated based on dividing the entire range of directions / angles into a series of bins / bits. e. For example, a histogram of gradients / directions / angles may be computed based on the counts of (magnitudes of) gradients in each bin / bin / direction / angle. 7) For example, it may be based on the color / brightness / intensity (or histogram of color / brightness / intensity) of samples inside (and / or adjacent to) the current block. a. For example, the current prediction sample point before OBMC can be used. b. For example, adjacent reconstruction points can be used. c. For example, a histogram of color / brightness / intensity may be calculated based on counts of sample values in the Y and / or U and / or V (or, R and / or G and / or B) component domains. i. For example, the sample values at a series of specific positions in the current block can be counted. ii. For example, the sample values of all samples in the current block may be counted. iii. For example, the sample values at a specific location adjacent to the current block may be counted. iv. For example, the sample values of all samples adjacent to the current block (on the left and / or above) may be counted. d. For example, a histogram of color / brightness / intensity can be calculated based on dividing the entire range of color / brightness / intensity values into a series of bins / bits. e. For example, a histogram of color / brightness / intensity can be calculated based on counting the number of samples in each bin / bit. 8) For example, it can be based on the number of dominant gradients / directions / angles / colors / brightness / intensities of samples inside (and / or adjacent to) the current block. a. For example, it can be calculated based on the prediction samples inside the current block (before OBMC). b. For example, it can be calculated based on the reconstructed samples adjacent to the current block. c. For example, the main gradient / direction / angle / color / brightness / intensity can be derived based on a histogram of gradient / direction / angle / color / brightness / intensity. d. For example, the dominant gradient / direction / angle / color / brightness / intensity can be derived based on how many bins / bits in the histogram show values greater than a threshold (e.g., gradient magnitude, color value, brightness value). e. For example, the dominant gradient / direction / angle / color / brightness / intensity can be derived based on how many bins / bins in the histogram provide values (e.g., gradient magnitude, color value, brightness value) that are much larger than the values of other bins / bins. i. For example, the values of the bins / bits in the histogram can be sorted first - assuming that the sorted (e.g., from largest to smallest) values are X0, X1, X2, ..., X n-2 、X n-1 Indicates that n bins / bits are included in the histogram - if Xi > = a*(X i+1 ), then the interval / bin from 0 to i can be regarded as the main gradient, where a represents the scaling factor (for example, a can be equal to a constant between 2 and 20). f. For example, if the number of dominant gradients / directions / angles / colors / brightness / intensities is less than a certain number (eg, 1 or 2 or 3 or 4 or 5 or 6 or 7 or 8 or 9), OBMC may not be applied to the block. 4.2. In one example, whether OBMC is applied to the current block may depend on the template cost. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on a first non-hybrid template cost and a second hybrid template cost. a. For example, the first template cost may be calculated based on the SAD between the current template and a reference template (where the reference template is identified by adding the current motion vector to the position of the current template). b. For example, the second template cost may be calculated based on the SAD between the current template and a hybrid reference template (where the hybrid reference template may be generated by mixing a template identified by the current motion vector and a template identified by a neighboring motion vector). c. For example, if the non-hybrid template cost is smaller, OBMC may not be applied to the block. i. Alternatively, OBMC can be applied to the block if the hybrid template is less expensive. 4.3. In one example, whether OBMC is applied to the current block may depend on the motion vector accuracy of the current block. 1) For example, the current block may be inter-frame Merge coded. 2) For example, the current block may be inter-frame AMVP coded. 3) For example, it can be based on whether the motion vector of the block is an integer (rather than fractional) precision motion vector. 4) For example, it may be based on whether the motion vector differences of the blocks are integer (rather than fractional) precision motion vector differences. 4.4 In one example, when the encoding and decoding information of the second block is used in the encoding and decoding process of the current block, the position of the second block may be restricted based on a specific rule. 1) For example, the encoding and decoding process of the current block may refer to at least one of the following: a. Mode decision b. Motion Candidate Derivation c. Exercise list generation d. Block vector candidate derivation e. Block vector list generation f. Intra-mode candidate derivation g. Intra-frame brightness MPM list generation h. Intra-frame chroma block vector candidate derivation under dual tree i. Intra-frame chroma mode candidate derivation under dual-tree j. Block-level OBMC on / off decision. 2) For example, the second block may be a reference block of the current block. 3) For example, the second block may be a reference block of a reference block of the current block. 4) For example, the second block can be a co-located or non-co-located luminance block with the current chrominance block. 5) For example, the second block may be required not to exceed the valid range. a. For example, the effective search range may be predefined. b. For example, the effective search range can be based on CTU size / information. c. For example, the effective search range may be based on the VPDU size / information. d. For example, the effective search range can be based on the tile size / information. e. For example, the effective search range can be based on sub-picture size / information. 6) For example, when the second block represents a reference block derived from a motion vector (e.g., from inter mode), the requirement for the location of the reference block may be based on the location of the CTU / CTU row / slice / sub-picture where the current block is located. a. For example, the location of the reference block may be required to be no more than the co-located CTU (ie, the CTU co-located with the current CTU in the reference picture) and X (eg, X=3) sample columns adjacent to the co-located CTU on the right. b. For example, the location of the reference block may be required to be no larger than the co-located CTU and the CTU adjacent to the right of the co-located CTU. c. For example, the location of the reference block may be required to be no more than the co-located CTU row. d. For example, the location of the reference block may be required to be no more than Y (eg, Y=3) sample rows above the co-located CTU. e. For example, the location of the reference block may be required to be no more than the co-located sub-picture. f. For example, the location of the reference block may be required to not exceed the co-located slice. 7) For example, when the second block represents the reference block B of the reference block A of the current block, the requirement for the position of the reference block B may be based on the position of the CTU / CTU row / slice / sub-picture where the current block is located. a. For example, the position of the reference block B may be required to be no more than the co-located CTU and X (eg, X=3) sample columns adjacent to the right of the co-located CTU. b. For example, the location of the reference block may be required to be no larger than the co-located CTU and the CTU adjacent to the right of the co-located CTU. c. For example, the location of the reference block may be required to be no more than the co-located CTU row. d. For example, the location of the reference block may be required to be no more than Y (eg, Y=3) sample rows above the co-located CTU. e. For example, the location of the reference block may be required to be no larger than the co-located CTU and the CTU adjacent to the co-located CTU on the left. f. For example, the location of the reference block may be required to be no more than the co-located CTU and the CTU adjacent to the co-located CTU on the left and the CTU adjacent to the co-located CTU on the right. g. For example, the location of the reference block may be required to be no more than the co-located sub-picture. h. For example, the location of the reference block may be required to not exceed the co-located slice. 8) For example, when the second block is the luma block of the current chroma block, the position of the luma block may be required not to exceed the co-located luma CU. a. Alternatively, the position of the luma block may be required to not exceed the current luma CTU. b. Alternatively, the location of the luma block may be required to be no more than the current luma CTU and a CTU adjacent to the right of the current luma CTU. c. Alternatively, the position of the luma block may be required to not exceed the current luma CTU row. 9) For example, when the second block exceeds the valid range (or, is outside the required position range), the second block may be deemed unavailable. a. For example, in this case, predefined codec information may be used instead. b. For example, in this case, the codec information of the second block is not used. 4.5 For example, how many and / or which prediction samples are used for mode decision may be based on codec information and / or predefined rules. 1) For example, the mode decision may refer to at least one of the following: a. OBMC on / off determination based on gradient / DIMD b. Gradient / DIMD-based transformation kernel determination c. Gradient / DIMD-based intra mode derivation for main transform d. Gradient / DIMD-based intra-mode derivation for secondary transform e. Gradient / DIMD-based Intra Mode Derivation for Separable Transform f. Gradient / DIMD-based Intra Mode Derivation for Non-separable Transforms g. LFNST kernel derivation for a specific mode (e.g., MIP mode) h. NSPT kernel derivation for a specific mode (eg, MIP mode). 2) For example, not all prediction samples in the current block are used for mode decision. 3) For example, some prediction samples within the current block are used for mode decision. 4) For example, the prediction samples used for mode decision can be sub-sampled. 5) For example, the prediction samples within the current block may be subsampled by a subsampling factor. a. For example, the subsampling factor in the width direction can be equal to 1 or 2 or 4 or 8. b. For example, the subsampling factor in the height direction may be equal to 1 or 2 or 4 or 8. c. For example, the values of the subsampling factors in the width and / or height directions may be derived based on the block width and / or height. i. For example, larger subsampling factors can be used for larger blocks. ii. For example, smaller subsampling factors can be used for smaller blocks. 6) For example, whether the prediction sample is subsampled can be determined based on block information. a. For example, it can be based on the number of samples in the block. b. For example, it can be based on block width. c. For example, it can be based on block height. d. For example, the subsampling method can be based on at least one threshold. 7) For example, the first row and / or last row and / or first column and / or last column of samples in the prediction block may not be used for mode decision. a. For example, the prediction block can be subsampled. b. For example, the prediction block may not be subsampled. 8) For example, partial / subsampled samples can be used for mode decision. 9) For example, gradients and / or histograms of gradients / colors / brightness / intensity can be calculated based on partial / subsampled samples. 4.6. Whether and / or how the methods disclosed above are applied can be signaled at the sequence level / group of pictures level / picture level / slice level / slice group level, such as in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. 4.7. Whether and / or how to apply the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / slice / sub-picture / other types of regions containing more than one sample or pixel. 4.8. Whether and / or how to apply the method disclosed above may depend on the coded information, such as block size, color format, single-tree partitioning / dual-tree partitioning, color component, slice / picture type.
[0103] The term "video unit" or "codec unit" or "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, or a TB. In the present disclosure, with respect to a "block coded in mode N", "mode N" here may be a prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.) or a codec technique (e.g., AMVP, SMVD, Merge, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, spatial GPM, SGPM, GPM inter-inter, GPM intra-intra, GPM inter-intra, MHP, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, LIC, OBMC, DIMD, TIMD, PDPC, CCLM, CCCM, GLM, intraTMP, ALF, deblocking, SAO, bilateral filter, LMCS, and corresponding variants, etc.).
[0104] Figure 39 FIG39 is a flowchart of a method 3900 for video processing according to an embodiment of the present disclosure. The method 3900 is implemented during conversion between a target video block of a video and a bitstream of the video.
[0105] At block 3910, for conversion between a video unit of a video and a bitstream of the video unit, codec information of the second block is applied to a codec process of a current block associated with the video unit. In this case, a location of the second block is restricted based on a predefined rule. In some embodiments, the codec process of the current block includes at least one of the following: mode decision, motion candidate derivation, motion list generation, block vector candidate derivation, block vector list generation, intra mode candidate derivation, intra luma most probable mode (MPM) list generation, intra chroma block vector candidate derivation under dual tree, intra chroma mode candidate derivation under dual tree, or block-level overlapping sub-block based motion compensation (OBMC) on / off decision.
[0106] At block 3920, conversion is performed based on the encoding and decoding process. Alternatively or additionally, conversion may include decoding the video unit from the bitstream. In this way, block-level adaptive OBMC that considers block characteristics based on already decoded information can lead to higher encoding and decoding gains and improve encoding and decoding efficiency.
[0107] In some embodiments, the second block is a reference block for the current block. In some other embodiments, the second block is a first reference block for a second reference block of the current block. Alternatively, the second block is a luma block that is collocated with the current chroma block, or wherein the second block is a luma block that is not collocated with the current chroma block.
[0108] In some embodiments, the second block does not exceed the effective search range. For example, the effective search range is predefined. Alternatively, the effective search range is based on the codec tree unit (CTU) size or CTU information. In some other embodiments, the effective search range is based on the virtual pipeline data unit (VPDU) size or VPDU information. As another example, the effective search range is based on the slice size or slice information. In some embodiments, the effective search range is based on the sub-picture size or sub-picture information.
[0109] In some embodiments, if the second block is a reference block derived from a motion vector, the requirement for the location of the reference block is based on the location of one of the following: CTU, CTU row, slice, or sub-picture where the current block is located. For example, the motion vector may be from an inter mode.
[0110] In some embodiments, the reference block is positioned no further than a co-located CTU and a certain number of sample columns adjacent to the right of the co-located CTU. For example, the co-located CTU is in a reference picture and is co-located with the current CTU. In some embodiments, the certain number of sample columns includes three sample columns.
[0111] In some embodiments, the reference block is located no further than the co-located CTU and the CTU adjacent to the co-located CTU on the right side thereof. In some other embodiments, the reference block is located no further than the co-located CTU row.
[0112] In some embodiments, the reference block is located no more than a certain number of sample rows above the co-located CTU, for example, the certain number of sample rows includes 3 sample rows.
[0113] In some embodiments, the reference block is positioned no further than the co-located sub-picture. In some other embodiments, the reference block is positioned no further than the co-located slice.
[0114] In some embodiments, if the second block is the first reference block of the second reference block of the current block, the position requirement of the second reference block is based on the position of the current block in one of the following: CTU, CTU row, slice, or sub-picture. In some embodiments, the position of the second reference block does not exceed the co-located CTU and a certain number of sample columns adjacent to the right of the co-located CTU. For example, the certain number of sample columns includes 3 sample columns.
[0115] In some embodiments, the second reference block is positioned no further than the co-located CTU and the CTU adjacent to the co-located CTU on the right. In some embodiments, the second reference block is positioned no further than the co-located CTU row.
[0116] In some embodiments, the second reference block is located no more than a certain number of sample rows above the co-located CTU, for example, the certain number of sample rows includes 3 sample rows.
[0117] In some embodiments, the second reference block is positioned no further than the co-located CTU and the CTU adjacent to the left of the co-located CTU. In some other embodiments, the second reference block is positioned no further than the co-located CTU and the CTU adjacent to the left of the co-located CTU and the CTU adjacent to the right of the co-located CTU.
[0118] In some embodiments, the position of the second reference block does not exceed the co-located sub-picture. In some other embodiments, the position of the second reference block does not exceed the co-located slice.
[0119] In some embodiments, if the second block is a luma block of the current chroma block, the luma block is positioned no further than a co-located luma codec unit (CU). In some other embodiments, if the second block is a luma block of the current chroma block, the luma block is positioned no further than a current luma codec tree unit (CTU). In some other embodiments, if the second block is a luma block of the current chroma block, the luma block is positioned no further than the current luma CTU and the CTU adjacent to the right of the current luma CTU. Alternatively, if the second block is a luma block of the current chroma block, the luma block is positioned no further than a row of the current luma CTU.
[0120] In some embodiments, if the second block exceeds the valid range or the second block is outside the required position range, the second block is deemed unusable. For example, predefined codec information is used for the codec process of the current block. As another example, the codec information of the second block is not used.
[0121] In some embodiments, an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is indicated at one of the following: sequence level, group of picture level, picture level, slice level, or slice group level. In some embodiments, an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or slice group header. In some embodiments, an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or an area containing more than one sample or pixel.
[0122] In some embodiments, method 3900 further includes determining, based on the coded information of the video unit, whether to apply the coded information of the second block to the coding process of the current block and / or how to apply the coded information of the second block to the coding process of the current block. The coded information may include at least one of the following: block size, color format, single-tree partitioning and / or dual-tree partitioning, color component, slice type, or picture type.
[0123] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; and generating a bitstream based on the codec process.
[0124] According to yet other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; generating a bitstream based on the codec process; and storing the bitstream in a non-transitory computer-readable medium.
[0125] Figure 40FIG4 is a flow chart of a method 4000 for video processing according to an embodiment of the present disclosure. The method 4000 is implemented during conversion between a target video block of a video and a bitstream of the video.
[0126] At block 4010, for conversion between a video unit of a video and a bitstream of the video unit, a set of prediction samples to be used for mode decision for the video unit is determined based on at least one of codec information or a predefined rule. In some embodiments, the number of prediction samples in the set of prediction samples is determined based on at least one of the codec information or the predefined rule. Alternatively or additionally, determining which prediction samples are in the set of prediction samples is based on at least one of the codec information or the predefined rule.
[0127] At block 4020, conversion is performed based on the set of prediction samples. In some embodiments, conversion may include encoding the video unit from a bitstream. Alternatively or additionally, conversion may include decoding the video unit from the bitstream. In this manner, block-level adaptive OBMC that considers block characteristics based on already decoded information can result in higher codec gain and improved codec efficiency.
[0128] In some embodiments, the mode decision comprises at least one of: gradient-based overlapped sub-block based motion compensation (OBMC) on / off determination, decoder-side intra mode derivation (DIMD)-based OBMC on / off determination, gradient-based transform kernel determination, DIMD-based transform kernel determination, gradient-based intra mode derivation for primary transform, DIMD-based intra mode derivation for primary transform, gradient-based intra mode derivation for secondary transform, DIMD-based intra mode derivation for secondary transform, gradient-based intra mode derivation for separable transform, DIMD-based intra mode derivation for separable transform, gradient-based intra mode derivation for non-separable transform, DIMD-based intra mode derivation for non-separable transform, low-frequency non-separable transform (LFNST) kernel derivation for codec mode, or non-separable primary transform (NSPT) kernel derivation for intra coding for codec mode. In some embodiments, the codec mode is matrix-weighted intra prediction (MIP) mode.
[0129] In some embodiments, not all prediction samples in the current block are used for mode decision. In some other embodiments, some prediction samples in the current block are used for mode decision.
[0130] In some embodiments, the set of prediction samples used for mode decision is subsampled. In some embodiments, the prediction samples within the current block are subsampled by a subsampling factor. For example, the subsampling factor in the width direction is equal to 1, 2, 4, or 8. As another example, the subsampling factor in the height direction is equal to 1, 2, 4, or 8.
[0131] In some embodiments, the value of the subsampling factor in the width direction is derived based on at least one of the block width or the block height. Alternatively or additionally, the value of the subsampling factor in the height direction is derived based on at least one of the block width or the block height. In some embodiments, if the size of a first block is larger than the size of a second block, a first subsampling factor that is larger than the second subsampling factor is used for the first block, and the second subsampling factor is used for the second block. In other words, a larger subsampling factor can be used for the larger block, and a smaller subsampling factor can be used for the smaller block.
[0132] In some embodiments, whether a set of prediction samples is subsampled is determined based on block information. For example, whether a set of prediction samples is subsampled is determined based on the number of samples in the block. As another example, whether a set of prediction samples is subsampled is determined based on the block width. In some embodiments, whether a set of prediction samples is subsampled is determined based on the block height. In some other embodiments, the subsampling method is based on at least one threshold.
[0133] In some embodiments, at least one of the first row, last row, first column, or last column of samples in the prediction block is not used for mode decision. For example, the prediction block is subsampled, or the prediction block is to be subsampled. As another example, a portion of the samples or the subsampled samples are used for mode decision.
[0134] In some embodiments, the gradient is determined based on a portion or sub-sampled samples. Alternatively or additionally, a histogram of at least one of the gradient, color, brightness or intensity is determined based on a portion or sub-sampled samples.
[0135] In some embodiments, an indication of whether to determine a set of prediction samples used for mode decision for a video unit and / or how to determine a set of prediction samples used for mode decision for a video unit is indicated at one of the following: sequence level, group of picture level, picture level, slice level, or slice group level. In some embodiments, an indication of whether to determine a set of prediction samples used for mode decision for a video unit and / or how to determine a set of prediction samples used for mode decision for a video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header. In some embodiments, an indication of whether and / or how to determine a set of prediction samples used for mode decision for a video unit is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or a region containing more than one sample or pixel.
[0136] In some embodiments, method 4000 further includes determining whether and / or how to determine a set of prediction samples to be used for mode decision of the video unit based on coded information of the video unit. The coded information may include at least one of the following: block size, color format, single-tree partitioning and / or dual-tree partitioning, color component, slice type, or picture type.
[0137] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: determining a set of prediction samples for mode decision for a video unit based on at least one of the following: codec information or a predefined rule; and generating a bitstream based on the set of prediction samples for conversion between a video unit and a bitstream of the video unit.
[0138] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided. The method includes: determining a set of prediction samples for mode decision of a video unit based on at least one of the following: codec information or a predefined rule, for conversion between a video unit and a bitstream of the video unit; generating a bitstream based on the set of prediction samples; and storing the bitstream in a non-transitory computer-readable medium.
[0139] The embodiments of the present disclosure may be described according to the following items, the features of which may be combined in any reasonable way.
[0140] Item 1. A method for video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, applying codec information of a second block to a codec process of a current block associated with the video unit, wherein a position of the second block is restricted based on a predefined rule; and performing the conversion based on the codec process.
[0141] Item 2. A method according to Item 1, wherein the encoding and decoding process of the current block includes at least one of the following: mode decision, motion candidate derivation, motion list generation, block vector candidate derivation, block vector list generation, intra-frame mode candidate derivation, intra-frame luminance most probable mode (MPM) list generation, intra-frame chroma block vector candidate derivation under dual tree, intra-frame chroma mode candidate derivation under dual tree, or block-level overlapping sub-block based motion compensation (OBMC) on / off decision.
[0142] Clause 3. The method of clause 1, wherein the second block is a reference block for the current block.
[0143] Item 4. The method of Item 1, wherein the second block is a first reference block of a second reference block of the current block.
[0144] Item 5. The method of Item 1, wherein the second block is a luma block that is co-located with the current chroma block, or wherein the second block is a luma block that is not co-located with the current chroma block.
[0145] Clause 6. The method of clause 1, wherein the second block does not exceed an effective search range.
[0146] Item 7. A method according to Item 6, wherein the effective search range is predefined, or wherein the effective search range is based on a codec tree unit (CTU) size or CTU information, or wherein the effective search range is based on a virtual pipeline data unit (VPDU) size or VPDU information, or wherein the effective search range is based on a slice size or slice information, or wherein the effective search range is based on a sub-picture size or sub-picture information.
[0147] Item 8. A method according to item 1, wherein if the second block is a reference block derived from a motion vector, the requirement for the location of the reference block is based on the location of the current block in one of the following: a CTU, a CTU row, a slice, or a sub-picture.
[0148] Item 9. The method according to Item 8, wherein the reference block is located no further than a co-located CTU and a certain number of sample columns adjacent to the right of the co-located CTU.
[0149] Item 10. The method of item 9, wherein the co-located CTU is in a reference picture and is co-located with the current CTU.
[0150] Item 11. The method of Item 9, wherein the number of sample columns comprises 3 sample columns.
[0151] Item 12. The method of Item 8, wherein the reference block is located no further than a co-located CTU and a CTU adjacent to the co-located CTU to the right.
[0152] Item 13. The method of Item 8, wherein the reference block is located no further than a co-located CTU row.
[0153] Item 14. The method of Item 8, wherein the reference block is located no more than a certain number of sample rows above the co-located CTU.
[0154] Item 15. The method of Item 14, wherein the number of sample rows comprises 3 sample rows.
[0155] Item 16. The method of Item 8, wherein the reference block is located no further than a co-located sub-picture.
[0156] Item 17. The method of Item 8, wherein the reference block is located no further than the co-located slice.
[0157] Item 18. A method according to item 1, wherein if the second block is a first reference block of a second reference block of the current block, the requirement for the position of the second reference block is based on the position of the current block in one of the following: a CTU, a CTU row, a slice or a sub-picture.
[0158] Item 19. The method of Item 18, wherein the second reference block is located no further than a co-located CTU and a certain number of sample columns adjacent to the right of the co-located CTU.
[0159] Item 20. The method of Item 19, wherein the number of sample columns comprises 3 sample columns.
[0160] Item 21. The method of Item 18, wherein the second reference block is located no further than a co-located CTU and a CTU adjacent to the co-located CTU to the right.
[0161] Item 22. The method of Item 18, wherein the second reference block is located no further than a co-located CTU row.
[0162] Item 23. The method of Item 18, wherein the second reference block is located no more than a certain number of sample rows above the co-located CTU.
[0163] Item 24. The method of Item 23, wherein the number of sample rows comprises 3 sample rows.
[0164] Item 25. The method of Item 18, wherein the second reference block is located no further than a co-located CTU and a CTU adjacent to the co-located CTU on the left.
[0165] Item 26. The method of Item 18, wherein the second reference block is located no further than a co-located CTU and a CTU adjacent to the co-located CTU on the left and a CTU adjacent to the co-located CTU on the right.
[0166] Item 27. The method of Item 18, wherein the second reference block is located no further than a co-located sub-picture.
[0167] Item 28. The method of Item 18, wherein the second reference block is located no further than the co-located slice.
[0168] Item 29. The method of Item 1, wherein if the second block is a luma block of a current chroma block, the luma block is positioned no further than a co-located luma codec unit (CU).
[0169] Item 30. The method of Item 1, wherein if the second block is a luma block of a current chroma block, the luma block is located no further than a current luma codec tree unit (CTU).
[0170] Item 31. The method of Item 1, wherein if the second block is a luma block of the current chroma block, the luma block is positioned no further than the current luma CTU and a CTU adjacent to the right of the current luma CTU.
[0171] Item 32. The method of Item 1, wherein if the second block is a luma block of a current chroma block, the luma block is positioned no further than a current luma CTU row.
[0172] Item 33. The method of Item 1, wherein the second block is deemed unavailable if the second block exceeds a valid range or the second block is outside a required position range.
[0173] Item 34. The method according to Item 33, wherein predefined codec information is used for the coding process of the current block.
[0174] Item 35. The method of Item 33, wherein the codec information of the second block is not used.
[0175] Item 36. A method according to any one of items 1 to 35, wherein an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is indicated at one of the following: sequence level, picture group level, picture level, slice level or slice group level.
[0176] Item 37. A method according to any one of Items 1 to 35, wherein an indication of whether the coding information of the second block is applied to the coding process of the current block and / or how the coding information of the second block is applied to the coding process of the current block is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.
[0177] Item 38. A method according to any one of Items 1 to 35, wherein an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or an area containing more than one sample or pixel.
[0178] Item 39. The method according to any one of Items 1 to 35 further includes: determining whether to apply the coding information of the second block to the coding process of the current block and / or how to apply the coding information of the second block to the coding process of the current block based on the coded information of the video unit, the coded information including at least one of the following: block size, color format, single tree partitioning and / or double tree partitioning, color component, slice type or picture type.
[0179] Item 40. A method of video processing, comprising: determining, for conversion between a video unit of a video and a bitstream of the video unit, a set of prediction samples to be used for mode decision of the video unit based on at least one of: codec information or a predefined rule; and performing the conversion based on the set of prediction samples.
[0180] Item 41. A method according to Item 40, wherein the number of the prediction samples in the set of prediction samples is determined based on at least one of the following: the codec information or the predefined rule, and / or which prediction samples are in the set of prediction samples are determined based on at least one of the following: the codec information or the predefined rule.
[0181] Item 42. A method according to Item 40, wherein the mode decision includes at least one of the following: gradient-based overlapped sub-block based motion compensation (OBMC) on / off determination, decoder-side intra-frame mode derivation (DIMD)-based OBMC on / off determination, gradient-based transform kernel determination, DIMD-based transform kernel determination, gradient-based intra-frame mode derivation for primary transform, DIMD-based intra-frame mode derivation for primary transform, gradient-based intra-frame mode derivation for secondary transform, DIMD-based intra-frame mode derivation for secondary transform, gradient-based intra-frame mode derivation for separable transform, DIMD-based intra-frame mode derivation for separable transform, gradient-based intra-frame mode derivation for non-separable transform, DIMD-based intra-frame mode derivation for non-separable transform, low-frequency non-separable transform (LFNST) kernel derivation for codec mode, or non-separable primary transform (NSPT) kernel derivation for intra-frame coding and decoding for codec mode.
[0182] Item 43. The method of Item 42, wherein the coding mode is a matrix-weighted intra prediction (MIP) mode.
[0183] Item 44. The method of Item 40, wherein less than all prediction samples within the current block are used for the mode decision.
[0184] Item 45. The method of Item 40, wherein a portion of the prediction samples within the current block are used for mode decision.
[0185] Item 46. The method of Item 40, wherein the set of prediction samples used for the mode decision is subsampled.
[0186] Item 47. The method of Item 40, wherein the prediction samples within the current block are subsampled by a subsampling factor.
[0187] Item 48. The method of Item 47, wherein the subsampling factor in the width direction is equal to 1 or 2 or 4 or 8.
[0188] Item 49. A method according to Item 47, wherein the subsampling factor in the height direction is equal to 1 or 2 or 4 or 8.
[0189] Item 50. A method according to Item 47, wherein the value of the subsampling factor in the width direction is derived based on at least one of the block width or the block height, and / or wherein the value of the subsampling factor in the height direction is derived based on at least one of the block width or the block height.
[0190] Item 51. A method according to Item 50, wherein if the size of the first block is larger than the size of the second block, a first subsampling factor larger than a second subsampling factor is used for the first block, and the second subsampling factor is used for the second block.
[0191] Item 52. The method of Item 40, wherein whether the set of prediction samples are subsampled is determined based on block information.
[0192] Item 53. The method of Item 52, wherein whether the set of prediction samples is subsampled is determined based on the number of samples in a block.
[0193] Item 54. The method of Item 52, wherein whether the set of prediction samples is subsampled is determined based on a block width.
[0194] Item 55. The method of Item 52, wherein whether the set of prediction samples is subsampled is determined based on a block height.
[0195] Item 56. The method of Item 52, wherein the subsampling method is based on at least one threshold.
[0196] Item 57. The method of Item 40, wherein at least one of a first row, a last row, a first column, or a last column of samples in the prediction block is not used for the mode decision.
[0197] Item 58. The method of Item 57, wherein the prediction block is subsampled, or wherein the prediction block is to be subsampled.
[0198] Item 59. The method of Item 40, wherein partial samples or subsampled samples are used for the mode decision.
[0199] Item 60. A method according to Item 40, wherein the gradient is determined based on a portion of the samples or subsampled samples, and / or wherein a histogram of at least one of the gradient, color, brightness or intensity is determined based on a portion of the samples or subsampled samples.
[0200] Item 61. A method according to any one of Items 1 to 60, wherein an indication of whether to determine the set of prediction samples used for the mode decision of the video unit and / or how to determine the set of prediction samples used for the mode decision of the video unit is indicated at one of the following: sequence level, picture group level, picture level, slice level or slice group level.
[0201] Item 62. A method according to any one of Items 1 to 60, wherein whether to determine the set of prediction samples used for the mode decision of the video unit and / or how to determine the set of prediction samples used for the mode decision of the video unit is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header or a slice group header.
[0202] Item 63. A method according to any one of Items 1 to 60, wherein an indication of whether to determine the set of prediction samples used for the mode decision of the video unit and / or how to determine the set of prediction samples used for the mode decision of the video unit is included in one of the following: a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec unit (CU), a virtual pipeline data unit (VPDU), a codec tree unit (CTU), a CTU row, a slice, a slice, a sub-picture, or a region containing more than one sample or pixel.
[0203] Item 64. The method according to any one of Items 1 to 60 further includes: determining whether to determine the set of prediction samples used for the mode decision of the video unit and / or how to determine the set of prediction samples used for the mode decision of the video unit based on the encoded and decoded information of the video unit, the encoded and decoded information including at least one of the following: block size, color format, single tree partitioning and / or double tree partitioning, color component, slice type or picture type.
[0204] Item 65. A method according to any one of Items 1 to 64, wherein the converting comprises encoding the video unit into the bitstream.
[0205] Item 66. A method according to any one of Items 1 to 64, wherein the converting comprises decoding the video unit from the bitstream.
[0206] Item 67. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 39 or any one of items 40 to 66.
[0207] Item 68. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform a method according to any one of Items 1 to 39 or any one of Items 40 to 66.
[0208] Item 69. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; and generating the bitstream based on the codec process.
[0209] Item 70. A method for storing a bitstream of a video, comprising: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a position of the second block is restricted based on a predefined rule; generating the bitstream based on the codec process; and storing the bitstream in a non-transitory computer-readable medium.
[0210] Item 71. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: determining, for conversion between a video unit of a video and a bitstream of the video unit, a set of prediction samples to be used for mode decision of the video unit based on at least one of: codec information or a predefined rule; and generating the bitstream based on the set of prediction samples.
[0211] Item 72. A method for storing a bitstream of a video, comprising: determining, for conversion between a video unit of the video and a bitstream of the video unit, a set of prediction samples to be used for mode decision of the video unit based on at least one of: codec information or a predefined rule; generating the bitstream based on the set of prediction samples; and storing the bitstream in a non-transitory computer-readable medium. Example device
[0212] Figure 41A block diagram of a computing device 4100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4100 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0213] It should be understood that Figure 41 The computing device 4100 shown in FIG. 4 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0214] like Figure 41 As shown, computing device 4100 comprises a general computing device 4100. Computing device 4100 may include at least one or more processors or processing units 4110, memory 4120, storage unit 4130, one or more communication units 4140, one or more input devices 4150, and one or more output devices 4160.
[0215] In some embodiments, the computing device 4100 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, and including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4100 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0216] The processing unit 4110 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 4120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of the computing device 4100. The processing unit 4110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0217] The computing device 4100 typically includes various computer storage media. Such media can be any media accessible by the computing device 4100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4120 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 4130 can be any removable or non-removable medium and can include machine-readable media, such as a memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 4100.
[0218] The computing device 4100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 41 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0219] The communication unit 4140 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 4100 can be implemented by a single computing cluster or multiple computing machines communicating via a communication connection. Thus, the computing device 4100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0220] Input device 4150 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4160 may be one or more of various output devices, such as a display, speaker, printer, etc. With the aid of communication unit 4140, computing device 4100 may also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 4100 may also communicate with one or more devices that enable a user to interact with computing device 4100, or any device that enables computing device 4100 to communicate with one or more other computing devices (e.g., a network card, modem, etc.), if necessary. Such communication may be performed via an input / output (I / O) interface (not shown).
[0221] In some embodiments, some or all components of the computing device 4100 may not be integrated into a single device, but may instead be arranged in a cloud computing architecture. In a cloud computing architecture, components may be provided remotely and may work together to implement the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to be aware of the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides applications over a wide area network that can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they appear to be a single access point for users. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, they may be provided from a traditional server or installed directly or otherwise on a client device.
[0222] In embodiments of the present disclosure, the computing device 4100 may be used to implement video encoding / decoding. The memory 4120 may include one or more video encoding / decoding modules 4125 having one or more program instructions. These modules can be accessed and executed by the processing unit 4110 to perform the functions of the various embodiments described herein.
[0223] In an example embodiment performing video encoding, an input device 4150 may receive video data as input 4170 to be encoded. The video data may be processed, for example, by a video codec module 4125 to generate an encoded bitstream. The encoded bitstream may be provided as output 4180 via an output device 4160.
[0224] In an example embodiment performing video decoding, an input device 4150 may receive an encoded bitstream as input 4170. The encoded bitstream may be processed, for example, by a video codec module 4125 to generate decoded video data. The decoded video data may be provided as output 4180 via an output device 4160.
[0225] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such changes are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A video processing method, comprising: For conversion between a video unit of a video and a bitstream of the video unit, applying codec information of a second block to a codec process of a current block associated with the video unit, wherein a position of the second block is restricted based on a predefined rule; and The conversion is performed based on the encoding and decoding process.
2. The method according to claim 1, wherein the encoding and decoding process of the current block comprises at least one of the following: Model decision, Motion candidate derivation, Exercise list generation, Block vector candidate derivation, Block vector list generation, Intra mode candidate derivation, Intra-frame luminance most probable mode (MPM) list generation, Intra-frame chroma block vector candidate derivation under dual tree, Intra chroma mode candidate derivation under dual tree, or Block-level overlapped sub-block based motion compensation (OBMC) on / off decision. The method of claim 1 , wherein the second block is a reference block of the current block. The method of claim 1 , wherein the second block is a first reference block of a second reference block of the current block.
5. The method of claim 1 , wherein the second block is a luma block co-located with the current chroma block, or The second block is a luminance block that is not co-located with the current chrominance block. The method according to claim 1 , wherein the second block does not exceed an effective search range.
7. The method according to claim 6, wherein the effective search range is predefined, or wherein the effective search range is based on a codec tree unit (CTU) size or CTU information, or wherein the effective search range is based on a virtual pipeline data unit (VPDU) size or VPDU information, or wherein the effective search range is based on the slice size or slice information, or The effective search range is based on the sub-picture size or sub-picture information.
8. The method of claim 1 , wherein if the second block is a reference block derived from a motion vector, the requirement for the location of the reference block is based on the location of the current block of one of the following: CTU, CTU line, slices, or Sub-picture.
9. The method according to claim 8, wherein the location of the reference block does not exceed the co-located CTU and a certain number of sample columns adjacent to the right of the co-located CTU.
10. The method of claim 9, wherein the co-located CTU is in a reference picture and is co-located with the current CTU. The method according to claim 9 , wherein the certain number of sample point columns comprises 3 sample point columns.
12. The method according to claim 8, wherein the location of the reference block does not exceed the co-located CTU and the CTU adjacent to the right of the co-located CTU. The method according to claim 8 , wherein the location of the reference block does not exceed a co-located CTU row. The method according to claim 8 , wherein the reference block is located no more than a certain number of sample rows above the co-located CTU. The method of claim 14 , wherein the certain number of sample rows comprises three sample rows. The method according to claim 8 , wherein the reference block is located within a co-located sub-picture. The method according to claim 8 , wherein the location of the reference block does not exceed the co-located slice.
18. The method of claim 1, wherein if the second block is a first reference block of a second reference block of the current block, the requirement for the position of the second reference block is based on a position of the current block where one of: CTU, CTU line, slices, or Sub-picture.
19. The method according to claim 18, wherein the location of the second reference block does not exceed the co-located CTU and a certain number of sample columns adjacent to the right side of the co-located CTU.
20. The method according to claim 19, wherein the certain number of sample point columns comprises 3 sample point columns.
21. The method according to claim 18, wherein a location of the second reference block does not exceed a co-located CTU and a CTU adjacent to the co-located CTU on the right side thereof.
22. The method of claim 18, wherein the second reference block is located no further than a co-located CTU row.
23. The method of claim 18, wherein the second reference block is located no more than a certain number of sample rows above the co-located CTU. The method of claim 23 , wherein the number of sample rows comprises three sample rows.
25. The method according to claim 18, wherein a location of the second reference block does not exceed a co-located CTU and a CTU adjacent to the co-located CTU on the left.
26. The method according to claim 18, wherein a position of the second reference block does not exceed a co-located CTU and a CTU adjacent to the co-located CTU on the left and a CTU adjacent to the co-located CTU on the right. The method according to claim 18 , wherein the second reference block is located no further than a co-located sub-picture. The method according to claim 18 , wherein the location of the second reference block does not exceed the co-located slice.
29. The method of claim 1, wherein if the second block is a luma block of a current chroma block, the luma block is located no further than a co-located luma codec unit (CU).
30. The method of claim 1, wherein if the second block is a luma block of a current chroma block, the luma block is located no further than a current luma codec tree unit (CTU).
31. The method according to claim 1, wherein if the second block is a luma block of a current chroma block, a position of the luma block does not exceed a current luma CTU and a CTU adjacent to the right of the current luma CTU.
32. The method of claim 1, wherein if the second block is a luma block of a current chroma block, the luma block is located no further than a current luma CTU row.
33. The method of claim 1, wherein the second block is deemed unavailable if the second block exceeds a valid range or the second block is outside a required position range. The method according to claim 33 , wherein predefined codec information is used for the coding process of the current block. The method of claim 33 , wherein the codec information of the second block is not used.
36. The method according to any one of claims 1 to 35, wherein an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is indicated at one of the following: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
37. The method according to any one of claims 1 to 35, wherein an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is indicated in one of the following: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
38. The method according to any one of claims 1 to 35, wherein an indication of whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block is included in one of the following: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Virtual Pipeline Data Unit (VPDU), Codec Tree Unit (CTU), CTU line, strips, piece, sub-image, or An area containing more than one sample or pixel.
39. The method according to any one of claims 1 to 35, further comprising: determining whether to apply the codec information of the second block to the codec process of the current block and / or how to apply the codec information of the second block to the codec process of the current block based on the codec information of the video unit, the codec information comprising at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color component, Strip type, or Image type.
40. A method for video processing, comprising: For conversion between a video unit of a video and a bitstream of the video unit, determining a set of prediction samples used for mode decision of the video unit based on at least one of: codec information or a predefined rule; and The conversion is performed based on the set of prediction samples.
41. The method according to claim 40, wherein the number of the prediction samples in the set of prediction samples is determined based on at least one of: the codec information or the predefined rule, and / or Which prediction sample points are in the set of prediction sample points is determined based on at least one of the following: the coding information or the predefined rule.
42. The method of claim 40, wherein the mode decision comprises at least one of: Gradient-based overlapped sub-block motion compensation (OBMC) on / off determination, OBMC on / off determination based on decoder-side intra mode derivation (DIMD), Gradient-based transformation kernel determination, DIMD-based transformation kernel determination, Gradient-based intra-mode derivation for the main transform, DIMD-based intra mode derivation for main transform, Gradient-based intra-mode derivation for secondary transforms, DIMD-based intra mode derivation for secondary transform, Gradient-based intra-mode derivation for separable transforms, DIMD-based intra mode derivation for separable transforms, Gradient-based intra-mode derivation for non-separable transforms, DIMD-based intra mode derivation for inseparable transforms, Low Frequency Non-separable Transform (LFNST) kernel derivation for codec mode, or Non-separable primary transform (NSPT) kernel derivation for intra coding for codec mode.
43. The method of claim 42, wherein the coding mode is a matrix-weighted intra prediction (MIP) mode.
44. The method of claim 40, wherein less than all prediction samples within the current block are used for the mode decision.
45. The method of claim 40, wherein a portion of prediction samples within the current block are used for mode decision.
46. The method of claim 40, wherein the set of prediction samples used for the mode decision is subsampled.
47. The method of claim 40, wherein the prediction samples within the current block are subsampled by a subsampling factor. The method according to claim 47 , wherein the subsampling factor in width direction is equal to 1 or 2 or 4 or 8. The method according to claim 47 , wherein the subsampling factor in the height direction is equal to 1 or 2 or 4 or 8.
50. The method of claim 47, wherein the value of the subsampling factor in the width direction is derived based on at least one of a block width or a block height, and / or The value of the subsampling factor in the height direction is derived based on at least one of the block width or the block height.
51. The method of claim 50, wherein if a size of a first block is larger than a size of a second block, a first subsampling factor larger than a second subsampling factor is used for the first block, and the second subsampling factor is used for the second block.
52. The method of claim 40, wherein whether the set of prediction samples is subsampled is determined based on block information.
53. The method of claim 52, wherein whether the set of prediction samples is subsampled is determined based on the number of samples in a block.
54. The method of claim 52, wherein whether the set of prediction samples is subsampled is determined based on a block width.
55. The method of claim 52, wherein whether the set of prediction samples is subsampled is determined based on block height.
56. The method of claim 52, wherein the subsampling method is based on at least one threshold.
57. The method of claim 40, wherein at least one of a first row, a last row, a first column, or a last column of samples in a prediction block is not used for the mode decision.
58. The method of claim 57, wherein the prediction block is subsampled, or The prediction block will be subsampled.
59. The method of claim 40, wherein partial samples or subsampled samples are used for the mode decision.
60. The method of claim 40, wherein the gradient is determined based on a portion of the samples or subsampled samples, and / or A histogram of at least one of gradient, color, brightness or intensity is determined based on a portion of the samples or subsampled samples.
61. The method of any of claims 40-60, wherein an indication of whether and / or how to determine the set of prediction samples used for the mode decision for the video unit is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
62. The method of any of claims 40-60, wherein an indication of whether and / or how to determine the set of prediction samples used for the mode decision for the video unit is indicated in one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependency Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
63. The method of any of claims 40-60, wherein an indication of whether and / or how to determine the set of prediction samples used for the mode decision for the video unit is included in one of: Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Unit (CU), Virtual Pipeline Data Unit (VPDU), Codec Tree Unit (CTU), CTU line, strips, piece, sub-image, or An area containing more than one sample or pixel.
64. The method of any one of claims 40-60, further comprising: Determining whether to determine the set of prediction samples used for the mode decision of the video unit and / or how to determine the set of prediction samples used for the mode decision of the video unit based on coded information of the video unit, the coded information comprising at least one of the following: Block size, Color format, Single-tree partitioning and / or dual-tree partitioning, Color component, Strip type, or Image type.
65. The method of any one of claims 1-64, wherein the converting comprises encoding the video unit into the bitstream.
66. The method of any one of claims 1-64, wherein the converting comprises decoding the video unit from the bitstream.
67. An apparatus for video processing, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-66.
68. A non-transitory computer-readable storage medium storing instructions for causing a processor to perform the method according to any one of claims 1-66.
69. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a location of the second block is restricted based on a predefined rule; as well as The bitstream is generated based on the encoding and decoding process.
70. A method for storing a bitstream of a video, comprising: applying codec information of a second block to a codec process of a current block associated with a video unit of the video, wherein a location of the second block is restricted based on a predefined rule; generating the bitstream based on the encoding and decoding process; as well as The bitstream is stored in a non-transitory computer-readable medium.
71. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: For conversion between a video unit of a video and a bitstream of the video unit, determining a set of prediction samples used for mode decision of the video unit based on at least one of: codec information or a predefined rule; and The bitstream is generated based on the set of prediction samples.
72. A method for storing a bitstream of a video, comprising: For conversion between a video unit of a video and a bitstream of the video unit, determining a set of prediction samples used for mode decision of the video unit based on at least one of: codec information or a predefined rule; generating the bitstream based on the set of prediction samples; as well as The bitstream is stored in a non-transitory computer-readable medium.